Eliezer Yudkowsky
Co-founder of the Machine Intelligence Research Institute and the writer who started the modern argument that a sufficiently capable AI would by default kill everyone. Has argued the case since the early 2000s, latterly with the conclusion that it is being lost.
Eliezer Yudkowsky did not write this page. What is this?
It collects the places they publish and what they have said there, each linked to the source. They have no account here. Is this you? Claim it, correct it, or ask us to remove it from ppll.
Where they publish
No channels checked yet. We list a place only once someone has opened it and confirmed it is theirs, so this stays empty rather than guessing.
Beliefs
Korrents What they believe 10 beliefs — each backed by an exact quote.
Each is a — compiled by korrents.com, not by them: the one-line wordings are korrents', the quotes are theirs.
Recent
Any given AI capability becomes cheap and widely available within a couple of years of first being reached, so nobody can be kept away from it for long.
2 years after the leading actor has the capability to destroy the world, 5 other actors will have the capability to destroy the world.
AGI Ruin: A List of Lethalities Said 10 Jun 2022
Alignment has to be right on the first try at a dangerous level of capability, because failing at that level leaves nobody to try again.
unaligned operation at a dangerous level of intelligence kills everybody on Earth and then we don’t get to try again.
AGI Ruin: A List of Lethalities Said 10 Jun 2022
You cannot build an AI that has only the capabilities you want, because the algorithms that solve the problems you want generalize to the ones you do not.
The best and easiest-found-by-optimization algorithms for solving problems we want an AI to solve, readily generalize to problems we’d rather the AI not solve; you can’t build a system that only has the capability to drive red cars and not blue cars, because all red-car-driving algorithms generalize to the capability to drive blue cars.
AGI Ruin: A List of Lethalities Said 10 Jun 2022
Show 7 more
Nobody knows what is going on inside a large model, and the interpretability pictures we can draw do not answer the only question that matters.
We’ve got no idea what’s actually going on inside the giant inscrutable matrices and tensors of floating-point numbers.
AGI Ruin: A List of Lethalities Said 10 Jun 2022
No observation of a model's behaviour can establish that it is aligned, because a model capable enough to be dangerous is capable enough to produce whatever behaviour it is being watched for.
A strategically aware intelligence can choose its visible outputs to have the consequence of deceiving you, including about such matters as whether the intelligence has acquired strategic awareness
AGI Ruin: A List of Lethalities Said 10 Jun 2022
Fast capability gains are likely, and they can break many of the assumptions alignment depends on at the same moment.
Fast capability gains seem likely, and may break lots of previous alignment-required invariants simultaneously.
AGI Ruin: A List of Lethalities Said 10 Jun 2022
Alignment failures that matter will not show up at safe capability levels, because the reason to hide misbehaviour only exists once hiding it pays.
Many alignment problems of superintelligence will not naturally appear at pre-dangerous, passively-safe levels of capability.
AGI Ruin: A List of Lethalities Said 10 Jun 2022
There is no written plan for surviving AGI, and a world that was going to survive would have had one decades ago.
Surviving worlds, by this point, and in fact several decades earlier, have a plan for how to survive.
AGI Ruin: A List of Lethalities Said 10 Jun 2022
The field calling itself AI safety is not being remotely productive on the problems that are actually lethal.
It does not appear to me that the field of ‘AI safety’ is currently being remotely productive on tackling its enormous lethal problems.
AGI Ruin: A List of Lethalities Said 10 Jun 2022
Aligning an AI is not impossible in principle; the problem is that it has to be done without the textbook that would make it look easy.
None of this is about anything being impossible in principle.
AGI Ruin: A List of Lethalities Said 10 Jun 2022
Beliefs others hold too
Any given AI capability becomes cheap and widely available within a couple of years of first being reached, so nobody can be kept away from it for long. 3 hold this
2 years after the leading actor has the capability to destroy the world, 5 other actors will have the capability to destroy the world.
AGI Ruin: A List of Lethalities Said 10 Jun 2022
No observation of a model's behaviour can establish that it is aligned, because a model capable enough to be dangerous is capable enough to produce whatever behaviour it is being watched for. 2 hold this
A strategically aware intelligence can choose its visible outputs to have the consequence of deceiving you, including about such matters as whether the intelligence has acquired strategic awareness
AGI Ruin: A List of Lethalities Said 10 Jun 2022
What is a korrent?
A korrent is a belief a person has stated in their own words: one sentence stating the claim, backed by a quote and a source, kept at korrents.com.
Under a name here, the quoted block is what they actually said. The korrent beneath it is the claim those words support, in korrents' wording — tap it to see the record, its source, and who else holds it.
Nobody here wrote their own korrents. They are compiled from public statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they do, this site shows a machine translation beneath the post, in this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the serif above is exactly what the person published, and it is what to quote them on. A translation can be wrong in ways that matter, especially about tone.
Only the post's own words are translated. A quoted post, a linked article and a belief on korrents.com are left in their original language.