Related posts
Nathan Lambert GitHub
rlhf-book — Textbook on reinforcement learning from human feedback
The the words it uses that this site has seen least often elsewhere. Posts are matched on those words alone — nothing here is a summary of this one.
Top people
Showing Profile →
Hiding
Hiding
13 September
1 September
-
Their words
and they're creatively pursuing goals much like very ambitious aggressive power-seeking humans creatively pursued their goals And so there are just structural analogies here that make it silly to not talk about agents as having motives and goals.
12 August
From one piece Chelsea Finn: This is the State of the Art in Robotics 2 beliefs, in the piece's order there
-
Their words
Now maybe this isn't completely out of the question but this would be quite challenging uh to do and that's because the calculus is a little bit different. We're not just running compute to optimize for a use case. We're actually running the robot in the real world and using the hardware and attempting the task in the real world.
-
Their words
we see that the across the board the single PIO like pre-trained PIO7 model matches or outperforms the fine-tuned specialists that were developed with reinforcement learning post-training for those downstream tasks.
11 August
-
Their words
In short, RLVR gives us no reason to expect the normative representations learned in pre-training will acquire motivational force over a given action, especially when post-training repeatedly selects trajectories for terminal task success.
5 August
15 July
-
Their words
But it's basically this idea that like the only thing that made claude code good was reinforcement learning. And the dimension along which it got good was like we made a model. We trained the model and the harness together. And so the model got really good at calling the specific tools in that harness.
30 June
-
Their words
And so I think you are right that lean is maybe overrated on the side of the importance of it being used as a VR environment for any kind of like just progress in math generally. But I I I definitely wouldn't write it out of the story.
27 May
-
korrents.com
Pain and pleasure act as motivational backstops that keep reinforcement-learning agents inner-aligned.Their words
valences like pain and pleasure have a natural account as motivational backstops for inner-aligning RL agents capable of mesa-optimization
20 March
-
Their words
So, you you can't look at any given scientific achievement purely in isolation and give it an objective grade without being aware of the context both in the the past and the future. And so it it it may never be something that you can just reinforcement learn the same way that that you can for much sort of more localized problems.
-
Their words
And so you're kind of like you're either on rails and you're part of the super intelligence circuits or you're not on rails and you're outside of the verifiable domains and suddenly everything kind of just like meanders.
25 November 2025
-
Their words
I want to like emphasize that I think the value function is something like it's going to make RL more efficient and I think that makes a difference but I think that anything you can do with a value function you can do without just more slowly.
17 November 2025
-
korrents.com
A verifiable task can be optimised by reinforcement learning until a neural network performs it extremely well.Their words
If a task/job is verifiable, then it is optimizable directly or via reinforcement learning, and a neural net can be trained to work extremely well.
17 October 2025
From one piece Andrej Karpathy — “We’re summoning ghosts, not building animals” 3 beliefs, in the piece's order there
-
Their words
a lot of what looks like learning is actually a lot more maturation of the brain and I think that actually very little reinforcement learning for animals and I think a lot of the reinforcement learning is actually like more like motor tasks. It's not intelligence tasks. So I actually kind of think humans don't actually like really use RL roughly speaking is what I would say.
-
korrents.com
Reinforcement learning is terrible, and it only looks good because everything we had before it was much worse.Their words
reinforcement learning is a lot worse than I think the average person thinks reinforcement learning is terrible. It just so happens that uh everything that we had before is much worse
-
Their words
So that's why I kind of call pre-training this kind of like crappy evolution. It's like the practically possible version with our technology and what we have available to us to get to a starting point where we can actually do things like reinforcement learning and so on.
26 September 2025
-
Their words
reinforcement learning is about understanding your world whereas large language models are about mimicking people doing what people say you should do. They're not about figuring out what to do.
12 September 2025
-
Their words
Uh so in order to effectively learn from your own experience, it turns out that it's really really important to already know something about what you're doing. Otherwise, it takes far too long.
23 March 2025
-
Their words
And it struck me how the most well-known brands have stood for one clear thing. Like they have a clear position. And so in order for superhuman to be memorable, I believed that we needed to occupy a clear position that was unique and which was available and which reinforced our product strategy.
3 February 2025
-
Their words
And these reasoning behaviors emerge naturally. So these things like, "Wait, let me see. Wait, let me check this. Oh, that might be a mistake." And they emerge from only having questions and answers.
28 November 2024
-
Their words
Reward hacking exists because RL environments are often imperfect, and it is fundamentally challenging to accurately specify a reward function.
Reward Hacking in Reinforcement Learninglilianweng.github.io
29 June 2023
-
Their words
When I talked about will GPT12 be AGI, my answer is no. Of course not. I mean, cross-entropy loss is never going to get you there. You need probably RL in fancy environments in order to get something that would be considered AGI-like.
24 March 2023
31 May 2016
Nothing matches.
What is a korrent?
A korrent is a belief a person has stated in their own words: one sentence stating the claim, backed by a quote and a source, kept at korrents.com.
Under a name here, the quoted block is what they actually said. The korrent beneath it is the claim those words support, in korrents' wording — tap it to see the record, its source, and who else holds it.
Nobody here wrote their own korrents. They are compiled from public statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they do, this site shows a machine translation beneath the post, in this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the serif above is exactly what the person published, and it is what to quote them on. A translation can be wrong in ways that matter, especially about tone.
Only the post's own words are translated. A quoted post, a linked article and a belief on korrents.com are left in their original language.