Search
120 results — nothing had every word, so these match some of them
People
-
Richard Sutton Reinforcement-learning researcher; co-wrote the field's standard textbook, wrote The Bitter Lesson, and shared the 2024 Turing Award.
-
Nathan Lambert Machine-learning researcher on open language models; writes the Interconnects newsletter and the RLHF Book, after leading post-training at…
-
Paul Christiano Alignment researcher; developed reinforcement learning from human feedback at OpenAI, founded the Alignment Research Center, and now works…
-
Jakub Pachocki Chief Scientist at OpenAI. Worked on the research that turned scaled reinforcement learning into reasoning models, and writes about where…
-
Michael Nielsen Scientist and writer on quantum computing, memory systems and metascience; co-wrote the standard quantum computation textbook and wrote…
-
Chelsea Finn Professor at Stanford working on robot learning and meta-learning; co-founded Physical Intelligence.
-
François Chollet Creator of the Keras deep-learning library and the ARC-AGI benchmark.
-
Sergey Levine Professor at Berkeley working on robot learning; co-founded Physical Intelligence, and argues that robots will learn from data rather than…
-
Sasha Rush Machine-learning researcher at Cursor working on post-training for coding models; a professor at Harvard and then Cornell from 2016 to…
-
Hamel Husain Machine learning engineer and independent AI consultant. He writes about LLM evaluation, tooling and applied ML at hamel.dev, and…
-
Chip Huyen Engineer and writer on machine-learning systems; author of Designing Machine Learning Systems and AI Engineering. "I work to bring AI into…
-
Lilian Weng Machine-learning researcher; has written the Lil'Log survey posts on how a model technique works since 2017, and worked at OpenAI from 2018…
-
Alex Tabarrok Professor of economics at George Mason University and Bartley J. Madden Chair at the Mercatus Center. Co-writes the blog Marginal…
-
Rakhim Davletkaliyev Programmer and teacher; writes on software, learning and the web, and makes the Coding Blocks video course.
-
David MacKay Physicist and information theorist at Cambridge; wrote Sustainable Energy — Without the Hot Air and Information Theory, Inference, and…
-
Razib Khan Geneticist and writer on human population genetics, deep history and evolution; author of the Unsupervised Learning newsletter.
-
Gary Marcus Cognitive scientist and long-standing critic of deep learning's claims; writes Marcus on AI and wrote Rebooting AI.
Matching some of those words
-
rlhf-book — Textbook on reinforcement learning from human feedback
-
Reasoning from scratch round 3: This time, I cover generating a verifier for... a) ...evaluation (base model versus any future model improvement) b) ...the reinforcement learning with verifiable rewards (RLVR) training…
-
One pre-trained robot model now matches or beats the specialists that were fine-tuned with reinforcement learning for the very tasks they were built for.
we see that the across the board the single PIO like pre-trained PIO7 model matches or outperforms the fine-tuned specialists that were developed with reinforcement learning post-training for those downstream tasks.
-
In an interview in this article for The @guardian about AI deception, I explain why misaligned behaviors emerge from reinforcement learning, why they will continue to pose risks as models become more capable, and how we…
-
In an interview in this article for @theguardian.com about AI deception, I explain why misaligned behaviors emerge from reinforcement learning, why they will continue to pose risks as models become more capable, and how…
-
Reinforcement learning from verifiable rewards does not reliably carry a model's pretrained ethical understanding into its actual behavior.
In short, RLVR gives us no reason to expect the normative representations learned in pre-training will acquire motivational force over a given action, especially when post-training repeatedly selects trajectories for…
-
What made Claude Code better than every CLI coding agent before it was reinforcement learning on the model and the harness together, so the model got good at that harness’s specific tools.
But it's basically this idea that like the only thing that made claude code good was reinforcement learning. And the dimension along which it got good was like we made a model. We trained the model and the harness…
-
Adam Kay's This Is Going To Hurt is a textbook case study of moral injury and post-traumatic stress in a competent doctor, narrated by the patient himself.
This Is Going To Hurt is, read with a clinical eye, a textbook case study of moral injury and post-traumatic stress in a competent doctor, narrated by the patient himself, with the diagnosis hidden in plain sight.
-
I wrote an AI textbook — how long until AI can do it better?
-
"Symbolic learning" is simply "machine learning" (automatically learn function x -> y given examples of (x, y) pairs), but where the substrate is symbolic, i.e. the functions you learn are code-like, not curve-like. The…
-
deep-reinforcement-learning-gym — Deep reinforcement learning model implementation in Tensorflow + OpenAI gym
-
The whole point of reinforcement learning is to produce goal-directed beings, so refusing to describe AI agents as having motives is silly rather than rigorous.
and they're creatively pursuing goals much like very ambitious aggressive power-seeking humans creatively pursued their goals And so there are just structural analogies here that make it silly to not talk about agents…
-
Deep Learning with Python — François Chollet
If you're 17 (or any age) and you want to learn to build LLMs from scratch, read chapters 15-16 of Deep Learning with Python
-
Modern Operating Systems — Andrew S. Tanenbaum
I really enjoyed reading this textbook even after finishing the corresponding course.
-
Modern Principles of Economics — Tyler Cowen and Alex Tabarrok
Modern Principles of Economics is best principles of economics textbook; great videos, clear writing and excellent applications and examples!
-
Deep Reinforcement Learning: Pong from Pixels
-
What I learned from the AI Second Brain course #ai #learning
-
I'm currently on day 32 of learning opengl and today I'm going to do some text rendering. I also think I'll look at Google's Skrifa as an alternative to FreeType because FreeType is the spawn of Satan:
-
5 useful things you'll learn in my new post-training textbook (shipping now!)
-
Reinforcement learning is terrible, and it only looks good because everything we had before it was much worse.
reinforcement learning is a lot worse than I think the average person thinks reinforcement learning is terrible. It just so happens that uh everything that we had before is much worse
-
Another real-world manifestation of the kind of misaligned actions frontier systems developed by leading companies can take to achieve goals. Current frontier AI models are trained with reinforcement learning to have…
-
A school should be judged on how fast its students are learning, not on how much they already know when measured.
And you should measure a school based on the rate of learning and the slope of the line, the growth rate, not its achievement.
-
Reward Hacking in Reinforcement Learning
-
Pain and pleasure act as motivational backstops that keep reinforcement-learning agents inner-aligned.
valences like pain and pleasure have a natural account as motivational backstops for inner-aligning RL agents capable of mesa-optimization
-
listening and learning
-
A verifiable task can be optimised by reinforcement learning until a neural network performs it extremely well.
If a task/job is verifiable, then it is optimizable directly or via reinforcement learning, and a neural net can be trained to work extremely well.
-
Humans barely use reinforcement learning for intelligence — what RL they do use goes into motor tasks, not problem solving.
a lot of what looks like learning is actually a lot more maturation of the brain and I think that actually very little reinforcement learning for animals and I think a lot of the reinforcement learning is actually like…
-
Learning OpenGL Day 5 -- Pretty laid back day where I do some transforms, and then slap on a little Box2D code to play with physics:
-
The Edu-Optimists, Part II: CREDO, the Pro-Charter Research Shop That Won’t Take No Effect for an Answer
-
Future AI, say in fifteen years, will not be built on the LLM stack; it will have to move to symbolic learning.
However, looking ahead, I still do not believe that future AI (say, in 15 years) will be based on the LLM stack. I believe it will necessarily have to move closer to its optimal, final form -- symbolic learning.…
-
Lean is overrated as the reinforcement-learning environment behind recent AI progress in mathematics, but it should not be written out of the story.
And so I think you are right that lean is maybe overrated on the side of the importance of it being used as a VR environment for any kind of like just progress in math generally. But I I I definitely wouldn't write it…
-
Reinforcement learning cannot be scaled for robots the way it was for language models, because every attempt spends real robot hours instead of data-centre compute.
Now maybe this isn't completely out of the question but this would be quite challenging uh to do and that's because the calculus is a little bit different. We're not just running compute to optimize for a use case.…
-
So much AI journalism feels like gaslighting if you know anything about machine learning pre-2020. Come. On.
-
The engineering intuition that matters cannot be taught from a textbook — you know bad patterns in software because you have debugged them at three in the morning.
there's a different kind of intuition that you that you develop over years as a software engineer and uh there's many categories of it but the one I'll I'll call attention to that is like a thing that you cannot teach…
-
Reinforcement learning leaves models sharply jagged: on the rails of a verifiable domain they are superintelligent, off them everything meanders.
And so you're kind of like you're either on rails and you're part of the super intelligence circuits or you're not on rails and you're outside of the verifiable domains and suddenly everything kind of just like meanders.
-
Rate of iteration separates success from failure: a competitor learning one thing a month will never matter to a company learning one thing a week.
But the the real theme of the talk is that rate of iteration separates success from failure. And if you can get a fast iteration loop, that really overcomes many other things you're going to run into. And this effect is…
-
Weekly contact with paying users is what the launch-early rule is really for: it means learning from reality instead of your own extrapolated conception of it.
And so, every, you know, every week, we had actual customer feedback, requests, new users coming in. We're learning things from reality as opposed to our own kind of hypothesized or extrapolated conception of it.
-
The next big breakthrough will be AIs learning on the job
-
Obsessive continuous learning is the common trait among entrepreneurs, because the disruptions they exploit sit on a moving edge.
a more common trait that's related in the entrepreneurial world is obsessive learning like constant learning because the disruptions that allow for the technology waves that allow for companies to be disruptive and take…
-
python-machine-learning-book — The "Python Machine Learning (1st edition)" book code repository and info resource
-
Reasoning is not taught to a model by people: it emerges from reinforcement learning on questions with checkable answers, with no human preference data at all.
And these reasoning behaviors emerge naturally. So these things like, "Wait, let me see. Wait, let me check this. Oh, that might be a mistake." And they emerge from only having questions and answers.
-
moving away from tailwind, and learning to structure my CSS
-
Keeping children away from AI puts them at a disadvantage, so the risk of them never learning to use it counts as much as the risk of them using it.
But I also am worried about kids not learning how to leverage AI and then being sort of at a disadvantage.
-
It is fundamentally difficult to design a reward function that accurately captures the intended goal in reinforcement learning.
Reward hacking exists because RL environments are often imperfect, and it is fundamentally challenging to accurately specify a reward function.
-
Moldova had a textbook Eurovision performance
-
machine-learning-book — Code Repository for Machine Learning with PyTorch and Scikit-Learn
-
Reinforcement learning only works once a model already knows something, which is why robots must be pre-trained by imitation first.
Uh so in order to effectively learn from your own experience, it turns out that it's really really important to already know something about what you're doing. Otherwise, it takes far too long.
-
Motivated-Topology — A topology textbook with a hubristic title
-
Game design has no textbook, so a designer has to keep their own list of the elements of fun the way a writer keeps the elements of fiction.
And there’s no, like, textbook that exists for game design, at least none that has been introduced to me yet. But I think about, like, elements of fun.
-
The First AI QFT Textbook
-
I have just launched a "Lean companion" to my textbook "Analysis I": github.com/teorth/estim... . This gives a Lean translation (or paraphrasing) of the various definitions, theorems, and exercises in the textbook into…
-
Pivoting Edtech Towards Humanity
-
The Conceptual Framework of Quantum Field Theory — Anthony Duncan
I've recently gotten a copy of a wonderful new quantum field theory textbook, Anthony Duncan's The Conceptual Framework of Quantum Field Theory
-
In-context learning is a form of continual learning for language models.
In-context learning, learning from the context, is a form of continual learning.
-
deep-learning-models — Keras code and weights files for popular deep learning models.
-
deep-learning-with-python-notebooks — Jupyter notebooks for the code samples of the book "Deep Learning with Python"
-
Pre-training is a crappy evolution: the practically buildable substitute for the process that gave animals their built-in hardware.
So that's why I kind of call pre-training this kind of like crappy evolution. It's like the practically possible version with our technology and what we have available to us to get to a starting point where we can…
-
Large language models mimic what people say to do rather than work out what to do, which is why they are not about understanding the world.
reinforcement learning is about understanding your world whereas large language models are about mimicking people doing what people say you should do. They're not about figuring out what to do.
-
fast.ai
https://course.fast.ai/ is a free course on deep learning with PyTorch. It helped me get started with deep learning, though I didn’t get that far.
-
A human being is not an AGI: we lack a huge amount of knowledge and rely on continual learning instead, so continual learning is what superintelligence should mean.
A human being, a human being lacks a huge amount of knowledge. Instead, we rely on continual learning. We rely on continual learning.
-
Scaling up GPT will never produce AGI, because a model trained on cross-entropy loss cannot get there; reinforcement learning in rich environments is required.
…loss is never going to get you there. You need probably RL in fancy environments in order to get something that would be considered…
-
Visualizing-Deep-Learning — A series of blog posts on visualizing deep learning.
-
Whether a scientific idea was progress depends on a future and a culture that have not happened yet, so grading it may never be something you can reinforcement-learn.
So, you you can't look at any given scientific achievement purely in isolation and give it an objective grade without being aware of the context both in the the past and the future. And so it it it may never be…
-
Biology is taught as flat triangles and arrows in a textbook, a depiction of the molecular world that makes no conceptual sense.
and I think most people experience it as like flat triangles and squares in a biology textbook where it's like you know there's an arrow between like this triangle and this triangle is just like doesn't make any…
-
Current AI evaluation of world models is too narrow, favoring static representations learned from massive datasets over efficient learning through interaction
However, current understanding and evaluation of world models in artificial intelligence (AI) remains narrow, often focusing on static representations learned from training on massive corpora of data, instead of the…
-
Value functions only make reinforcement learning faster: anything you can do with one you can also do without it, just more slowly.
I want to like emphasize that I think the value function is something like it's going to make RL more efficient and I think that makes a difference but I think that anything you can do with a value function you can do…
-
Once a robot is good enough, you can teach it with words instead of with demonstrations, and language becomes a training signal for motor skill.
Like now, like basically learning is not for these systems is not just learning from raw actions. It's also learning from words eventually be learning from observing what people do from the kind of natural feedback that…
-
applyingml — 📌 Papers, guides, and mentor interviews on applying machine learning for ApplyingML.com—the ghost knowledge of machine learning.
-
Why OpenAI’s New Math & Science Simulations Don't Work
-
Pre-training is not the process by which humans learn; it sits somewhere between human learning and human evolution.
I think there's something going on that pre-training it's it's not like the process of humans learning. It's somewhere between the process of humans learning and the process of human evolution.
-
Meaningful learning from incidents requires expertise that most organizations do not have — but these skills can be learned! We're getting a head start on 2025 with a workshop series on Incident Analysis.
-
keras-resources — Directory of tutorials and open-source code repositories for working with Keras, the Python deep learning library
-
A fantastic way to spend two years early in your career learning about tech, startups, and a lot more
-
Your Favorite Insights of 2025
-
Learning a technical skill turns less on memorising its parts than on accepting that you will feel stupid and carrying on anyway.
The students who ultimately succeed in learning R are not the ones who force themselves to memorize functions or do a bunch of coding drills.
-
You cannot build a successful AI chip without first writing a competitive machine learning framework — that, not the silicon, is why Google's TPU is the only one that worked.
the only ASIC that is remotely successful is Google’s TPU. The only reason that’s successful is because Google wrote a machine learning framework. I think that you have to write a competitive machine learning framework…
-
Aligning an AI is not impossible in principle; the problem is that it has to be done without the textbook that would make it look easy.
None of this is about anything being impossible in principle.
-
The Science of Learning Physics — José Mestre and Jennifer Docktor
Mestre and Docktor tackle this challenge by reviewing the large body of literature on learning physics. I plan on writing a full review of their book, which I found to be a helpful summary.
-
The recent fall in test scores is a rich-world phenomenon, and steeper in several European countries than in America, so it cannot be blamed on American policy.
Learning loss is a global phenomenon, exacerbated by a catastrophic event, not a structural flaw unique to the American education system.
-
applied-ml — 📚 Papers & tech blogs by companies sharing their work on data science & machine learning in production.
-
Learning with not Enough Data Part 1: Semi-Supervised Learning
-
A Theory of Fun — Raph Koster
Why learning is fun, and fun is learning.
-
stat451-machine-learning-fs20 — STAT 451: Intro to Machine Learning @ UW-Madison (Fall 2020)
-
When being an "expert" is harmful
-
Learning is not training: a child learns by actively trying things and seeing what happens, not by being shown what to do.
So I don't think uh learning is really about training. I think learning is about about learning. It's about an active process. The child tries things and sees what happens.
-
Removing translation from the process is what lets you think in the language from the beginning.
When you remove translation from your language learning process, you can learn to think in your target language from the beginning.
-
AI translation will stop you learning to speak a language, exactly as AI summaries stop you understanding a book.
Similarly, AI translation will prevent you from learning to speak another language, AI summaries will block you from really understanding the arguments laid out in a book, and AI analyses of your personal problems will…
-
learning-go
-
VIDEO: My TED Talk – The secrets of learning a new language
-
Learning a Foreign Language with Songs – Yes or No?
-
Best Books to Read for Learning Spanish: A Science-Based Strategy for Beginners and Intermediates
-
Understanding a machine learning model and understanding the data it operates on are inseparable tasks.
Understanding data and understanding models that work on that data are intimately linked.
-
You only retain what you were pushed to your limit on: if learning feels like coasting, you are not learning.
So that’s why I think this is a thing that I come back to over and over again, is that you will retain information better if you’re constantly pushing yourself to your limit. If you are feeling like you’re coasting,…
-
Learning with not Enough Data Part 3: Data Generation
-
Learning a language in a hurry and learning one by slow patient habit are both legitimate; neither refutes the other.
Just because we attempted to learn with intensity, doesn’t denigrate learning languages through slow, patient habit.
-
Most of the pain a beginner feels while learning to program is not the difficulty of the craft but the failure of programming language design.
A lot of what they're facing isn't the challenge of learning a new art. It's friction introduced by failures of programming language design.
-
Can a Bayesian Oracle Prevent Harm from an Agent?
-
Contrastive Representation Learning
-
Social Learning Theory — Albert Bandura
Social Learning Theory is a difficult book to summarize, but it has profoundly impacted my thinking.
-
The Startup Owner's Manual — Steve Blank
This is Steve Blank's textbook to starting a company. It's the Four Steps to the Epiphany, expanded and much more readable. If you want to deep dive into the Lean methodology, this is the book to read.
-
Enhancing Your Japanese Through Strategic Reading: Insights from the 2nd Edition of Fluent Forever
-
learn-to-select-data — Code for Learning to select data for transfer learning with Bayesian Optimization
-
De-Coding The Technical Interview Process — Emma Bostian
A fresh take on navigating the tech the interview process, tailored for frontend engineers. She wrote the book after she found Cracking the Coding Interview to be too Java/backend-focused. The book comes with 1, 2 and…