when AIs are extremely extremely capable my view is that those AIs will be harder to align than current systems. So for current systems, we have this feedback loop where we basically like we create an AI. We do some evaluations on it. We see that it has some kind of messed up behavior that we can kind of quickly understand. Then we like can like go look in training and be like, "Oh, the these training environments led to this problematic behavior. Let's like tweak that training data. Let's introduce some additional training data to like correct this other issue and then move forward from there." But in a regime where the AIs are extremely situationally aware, very very very very capable and um you know uh we don't necessarily understand what they're doing, this feedback loop breaks down.
This is the number one unsolved problem in AI. It's not the tech. We're making great progress on the technical alignment problem. But we haven't made jack progress on the human alignment problem
Today’s best AI systems are good enough that they’re now inside the fuzzy conceptual cloud of “AGI-ish”: that is, they’ve surpassed some people’s definitions of AGI, while falling well short of others’.
But the goal we need to achieve is so much easier: we just need to build a model that’s as good as us at alignment research, and that we trust more than ourselves to do this research well because it’s sufficiently aligned.
What’s ironic about all these AI safety people is they’re going to build the exact thing they fear. We need to have one model that we control and align. This is the only way you end up paper clipped. There’s no way you end up paper clipped if everybody has an AI.
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.