-
Their words
This Is Going To Hurt is, read with a clinical eye, a textbook case study of moral injury and post-traumatic stress in a competent doctor, narrated by the patient himself, with the diagnosis hidden in plain sight.
Related posts
Nathan Lambert GitHub
rlhf-book — Textbook on reinforcement learning from human feedback
The the words it uses that this site has seen least often elsewhere. Posts are matched on those words alone — nothing here is a summary of this one.
Top people
Showing Profile →
Hiding
Hiding
18 September
13 September
4 September
-
Recommendstheir ownrcmnd.app
Modern Principles of EconomicsTheir words
Modern Principles of Economics is best principles of economics textbook; great videos, clear writing and excellent applications and examples!
1 September
-
Their words
and they're creatively pursuing goals much like very ambitious aggressive power-seeking humans creatively pursued their goals And so there are just structural analogies here that make it silly to not talk about agents as having motives and goals.
16 August
-
Readrcmnd.app
Modern Operating SystemsTheir words
I really enjoyed reading this textbook even after finishing the corresponding course.
12 August
From one piece Chelsea Finn: This is the State of the Art in Robotics 2 beliefs, in the piece's order there
-
Their words
Now maybe this isn't completely out of the question but this would be quite challenging uh to do and that's because the calculus is a little bit different. We're not just running compute to optimize for a use case. We're actually running the robot in the real world and using the hardware and attempting the task in the real world.
-
Their words
we see that the across the board the single PIO like pre-trained PIO7 model matches or outperforms the fine-tuned specialists that were developed with reinforcement learning post-training for those downstream tasks.
11 August
-
Their words
In short, RLVR gives us no reason to expect the normative representations learned in pre-training will acquire motivational force over a given action, especially when post-training repeatedly selects trajectories for terminal task success.
10 August
5 August
15 July
From one piece Context engineering with Dex Horthy 2 beliefs, in the piece's order there
-
Their words
But it's basically this idea that like the only thing that made claude code good was reinforcement learning. And the dimension along which it got good was like we made a model. We trained the model and the harness together. And so the model got really good at calling the specific tools in that harness.
-
Their words
there's a different kind of intuition that you that you develop over years as a software engineer and uh there's many categories of it but the one I'll I'll call attention to that is like a thing that you cannot teach you cannot do you cannot learn in a textbook. The only way to learn it is like I know bad patterns in software because I have debugged them at three in the morning.
30 June
-
Their words
And so I think you are right that lean is maybe overrated on the side of the importance of it being used as a VR environment for any kind of like just progress in math generally. But I I I definitely wouldn't write it out of the story.
7 June
27 May
-
korrents.com
Pain and pleasure act as motivational backstops that keep reinforcement-learning agents inner-aligned.Their words
valences like pain and pleasure have a natural account as motivational backstops for inner-aligning RL agents capable of mesa-optimization
20 March
-
Their words
So, you you can't look at any given scientific achievement purely in isolation and give it an objective grade without being aware of the context both in the the past and the future. And so it it it may never be something that you can just reinforcement learn the same way that that you can for much sort of more localized problems.
-
Their words
And so you're kind of like you're either on rails and you're part of the super intelligence circuits or you're not on rails and you're outside of the verifiable domains and suddenly everything kind of just like meanders.
11 March
-
Their words
And there’s no, like, textbook that exists for game design, at least none that has been introduced to me yet. But I think about, like, elements of fun.
29 January
-
Their words
and I think most people experience it as like flat triangles and squares in a biology textbook where it's like you know there's an arrow between like this triangle and this triangle is just like doesn't make any conceptual sense and is like confusing and annoying
25 January
-
Their words
So what usually h what often happens is you raise prices and signups don't change. When I say signups, I mean the like signups per month, you know, the rate at sign or signups go up. This happens all the time.
25 November 2025
-
Their words
I want to like emphasize that I think the value function is something like it's going to make RL more efficient and I think that makes a difference but I think that anything you can do with a value function you can do without just more slowly.
17 November 2025
-
korrents.com
A verifiable task can be optimised by reinforcement learning until a neural network performs it extremely well.Their words
If a task/job is verifiable, then it is optimizable directly or via reinforcement learning, and a neural net can be trained to work extremely well.
17 October 2025
From one piece Andrej Karpathy — “We’re summoning ghosts, not building animals” 3 beliefs, in the piece's order there
-
Their words
a lot of what looks like learning is actually a lot more maturation of the brain and I think that actually very little reinforcement learning for animals and I think a lot of the reinforcement learning is actually like more like motor tasks. It's not intelligence tasks. So I actually kind of think humans don't actually like really use RL roughly speaking is what I would say.
-
korrents.com
Reinforcement learning is terrible, and it only looks good because everything we had before it was much worse.Their words
reinforcement learning is a lot worse than I think the average person thinks reinforcement learning is terrible. It just so happens that uh everything that we had before is much worse
-
Their words
So that's why I kind of call pre-training this kind of like crappy evolution. It's like the practically possible version with our technology and what we have available to us to get to a starting point where we can actually do things like reinforcement learning and so on.
26 September 2025
-
Their words
reinforcement learning is about understanding your world whereas large language models are about mimicking people doing what people say you should do. They're not about figuring out what to do.
12 September 2025
-
Their words
Uh so in order to effectively learn from your own experience, it turns out that it's really really important to already know something about what you're doing. Otherwise, it takes far too long.
31 May 2025
23 March 2025
-
Their words
And it struck me how the most well-known brands have stood for one clear thing. Like they have a clear position. And so in order for superhuman to be memorable, I believed that we needed to occupy a clear position that was unique and which was available and which reinforced our product strategy.
3 February 2025
From one piece DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters | Lex Fridman Podcast #459 3 beliefs, in the piece's order there
-
Their words
And humans are actually very good at reading or judging between two things versus... This goes back to the core of what RLHF and preference tuning is that it's hard to generate a good answer for a lot of problems, but it's easy to see which one is better.
-
Their words
And these reasoning behaviors emerge naturally. So these things like, "Wait, let me see. Wait, let me check this. Oh, that might be a mistake." And they emerge from only having questions and answers.
-
Their words
And the important thing to say is that no matter how you want the model to behave, these RLHF and preference-tuning techniques also improve performance. So, on things like math evals and code evals, there is something innate to these, what is called contrastive loss functions.
28 November 2024
-
Their words
Reward hacking exists because RL environments are often imperfect, and it is fundamentally challenging to accurately specify a reward function.
Reward Hacking in Reinforcement Learninglilianweng.github.io
30 September 2024
-
korrents.com
The Inca stones were not fitted by patient pecking with hammer stones; they were fused with acid.Their words
The common archeological wisdom that you’d find out of a textbook is that they just kept pecking away at it with hammer stones and setting them and resetting them until they were perfect, which has to be bullshit, that there is no way that they just were that meticulous. I mean, everybody’s got a hammerstone. I personally think it’s acids.
29 June 2023
-
Their words
When I talked about will GPT12 be AGI, my answer is no. Of course not. I mean, cross-entropy loss is never going to get you there. You need probably RL in fancy environments in order to get something that would be considered AGI-like.
9 May 2023
24 March 2023
10 June 2022
-
Their words
None of this is about anything being impossible in principle.
28 October 2020
-
Recommendsaffiliate linkrcmnd.app
The Startup Owner's ManualTheir words
This is Steve Blank's textbook to starting a company. It's the Four Steps to the Epiphany, expanded and much more readable. If you want to deep dive into the Lean methodology, this is the book to read.
31 May 2016
13 January 2016
Nothing matches.
What is a korrent?
A korrent is a belief a person has stated in their own words: one sentence stating the claim, backed by a quote and a source, kept at korrents.com.
Under a name here, the quoted block is what they actually said. The korrent beneath it is the claim those words support, in korrents' wording — tap it to see the record, its source, and who else holds it.
Nobody here wrote their own korrents. They are compiled from public statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they do, this site shows a machine translation beneath the post, in this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the serif above is exactly what the person published, and it is what to quote them on. A translation can be wrong in ways that matter, especially about tone.
Only the post's own words are translated. A quoted post, a linked article and a belief on korrents.com are left in their original language.