Buck Shlegeris
CEO of Redwood Research, where he works on AI control: making sure a powerful model cannot cause a catastrophe even if it is misaligned and trying to.
Buck Shlegeris did not write this page. What is this?
It collects the places they publish and what they have said there, each linked to the source. They have no account here. Is this you? Claim it, correct it, or ask us to remove it from ppll.
Where they publish
No channels checked yet. We list a place only once someone has opened it and confirmed it is theirs, so this stays empty rather than guessing.
Beliefs
Korrents What they believe 8 beliefs — each backed by an exact quote.
Each is a — compiled by korrents.com, not by them: the one-line wordings are korrents', the quotes are theirs.
Recent
Labs should make powerful models safe to deploy even if those models are misaligned and actively trying to subvert the safeguards.
That is, labs should make sure that the safety measures they apply to their powerful models prevent unacceptably bad outcomes, even if the AIs are misaligned and intentionally try to subvert those safety measures.
The case for ensuring that powerful AIs are controlled Said 7 May 2024
Controlling a scheming model needs no research breakthrough, which is what makes it the tractable half of the problem today.
We think no fundamental research breakthroughs are required for labs to implement safety measures that meet our standard for AI control for early transformatively useful AIs; we think that meeting our standard would substantially reduce the risks posed by intentional subversion.
The case for ensuring that powerful AIs are controlled Said 7 May 2024
Catching an AI trying to cause harm should count as a win, because being caught changes the situation more than the attempt cost.
So when evaluating control, we should count catching an AI red-handed as a win condition.
The case for ensuring that powerful AIs are controlled Said 7 May 2024
Show 5 more
Whether a model is controlled can be settled with capability evaluations, which makes control far easier to check than alignment.
Because evaluating control just requires evaluating capabilities, it's far easier to robustly evaluate than alignment.
The case for ensuring that powerful AIs are controlled Said 7 May 2024
Control is a fundamentally limited approach that stops working on sufficiently powerful models, so alignment is still needed in the end.
AI control (with only black-box techniques) seems like a fundamentally limited approach.
The case for ensuring that powerful AIs are controlled Said 7 May 2024
A control evaluation that reports under one per cent risk should be read as several per cent, because the evaluation can itself fail.
In particular, suppose that our control evaluations directly estimate a less than 1% chance of catastrophe if our untrusted AIs are scheming, then the actual risk conditional on scheming is probably more like 5% to 20% due to the potential for failures in the evaluation.
The case for ensuring that powerful AIs are controlled Said 7 May 2024
Risk from a scheming model should be managed the way insider security risk is: assume some of them are compromised and stay safe anyway.
We're advocating that companies handle risk from scheming models in a similar way–striving to ensure that they'll be safe even if their alignment efforts fail to prevent models from scheming.
The case for ensuring that powerful AIs are controlled Said 7 May 2024
No observation of a model's behaviour can establish that it is aligned, because a model capable enough to be dangerous is capable enough to produce whatever behaviour it is being watched for.
The basic problem with evaluating alignment is that no matter what behaviors you observe, you have to worry that your model is just acting that way in order to make you think that it is aligned.
The case for ensuring that powerful AIs are controlled Said 7 May 2024
Beliefs others hold too
No observation of a model's behaviour can establish that it is aligned, because a model capable enough to be dangerous is capable enough to produce whatever behaviour it is being watched for. 2 hold this
The basic problem with evaluating alignment is that no matter what behaviors you observe, you have to worry that your model is just acting that way in order to make you think that it is aligned.
The case for ensuring that powerful AIs are controlled Said 7 May 2024
What is a korrent?
A korrent is a belief a person has stated in their own words: one sentence stating the claim, backed by a quote and a source, kept at korrents.com.
Under a name here, the quoted block is what they actually said. The korrent beneath it is the claim those words support, in korrents' wording — tap it to see the record, its source, and who else holds it.
Nobody here wrote their own korrents. They are compiled from public statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they do, this site shows a machine translation beneath the post, in this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the serif above is exactly what the person published, and it is what to quote them on. A translation can be wrong in ways that matter, especially about tone.
Only the post's own words are translated. A quoted post, a linked article and a belief on korrents.com are left in their original language.