-
korrents.com
The fundamental challenge of alignment is generalization: holding values in situations the training never covered.Their words
The fundamental challenge of AI alignment is generalization.
Related posts
Lilian Weng GitHub
The the words it uses that this site has seen least often elsewhere. Posts are matched on those words alone — nothing here is a summary of this one.
Top people
Showing Profile →
Hiding
Hiding
6 September
12 August
-
Their words
we find that the performance on held out tasks decreases dramatically. Whereas if we um just take out a random 20% of the data that's less diverse than the most diverse subset, the performance um only decreases a little bit. And so this suggests that actually having really diverse data plays an important role in enabling it to generalize to new tasks.
13 February
From one piece Dario Amodei — “We are near the end of the exponential” 2 beliefs, in the piece's order there
-
Their words
the goal is not to teach the model every possible skill within RL just as we don't do that within pre-training, right? Within pre-training, we're not trying to expose the model to, you know, every every possible you know, way that words could be put together, right? You know, we're it's it's rather that the model trains on a lot of things and then and then it reaches generalization across pre-training, right?
-
korrents.com
Models already generalise substantially from tasks that can be verified to tasks that cannot.Their words
We already see substantial generalization from things that that verify to things that don't verify. We're already seeing that.
12 September 2025
-
Their words
Yeah. So there's a subtlety here. Emerging capabilities don't just come from the fact that internet data has a lot of stuff in it. They also come from the fact that generalization once it reaches a certain level becomes compositional.
31 August 2024
From one piece Why A.I. Isn't Going to Make Art 2 beliefs, in the piece's order there
-
korrents.com
Art is what results from making an enormous number of choices, and a prompt contains almost none of them.Their words
But let me offer a generalization: art is something that results from making a lot of choices.
-
korrents.com
Writing that deserves a reader's attention is always the product of effort by the person who wrote it.Their words
Let me offer another generalization: any writing that deserves your attention as a reader is the result of effort expended by the person who wrote it.
20 December 2023
8 September 2022
5 November 2019
-
Their words
We argue that solely measuring skill at any given task falls short of measuring intelligence, because skill is heavily modulated by prior knowledge and experience: unlimited priors or unlimited training data allow experimenters to "buy" arbitrary levels of skills for a system, in a way that masks the system's own generalization power.
Nothing matches.
What is a korrent?
A korrent is a belief a person has stated in their own words: one sentence stating the claim, backed by a quote and a source, kept at korrents.com.
Under a name here, the quoted block is what they actually said. The korrent beneath it is the claim those words support, in korrents' wording — tap it to see the record, its source, and who else holds it.
Nobody here wrote their own korrents. They are compiled from public statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they do, this site shows a machine translation beneath the post, in this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the serif above is exactly what the person published, and it is what to quote them on. A translation can be wrong in ways that matter, especially about tone.
Only the post's own words are translated. A quoted post, a linked article and a belief on korrents.com are left in their original language.