From one piece The case for ensuring that powerful AIs are controlled 5 beliefs, in the piece's order there
-
Their words
The basic problem with evaluating alignment is that no matter what behaviors you observe, you have to worry that your model is just acting that way in order to make you think that it is aligned.
-
Their words
We're advocating that companies handle risk from scheming models in a similar way–striving to ensure that they'll be safe even if their alignment efforts fail to prevent models from scheming.
-
Their words
AI control (with only black-box techniques) seems like a fundamentally limited approach.
+ 2 more
-
Their words
Because evaluating control just requires evaluating capabilities, it's far easier to robustly evaluate than alignment.
-
Their words
That is, labs should make sure that the safety measures they apply to their powerful models prevent unacceptably bad outcomes, even if the AIs are misaligned and intentionally try to subvert those safety measures.