From one piece Alignment is not solved 5 beliefs, in the piece's order there
-
Their words
In fact, making an evil version of Claude that’s just as smart and agentic would be pretty easy.
-
korrents.com
Recursive self-improvement has already started, because AI research itself is now being automated.Their words
We are starting to automate AI research and the recursive self-improvement process has begun.
-
Their words
But the goal we need to achieve is so much easier: we just need to build a model that’s as good as us at alignment research, and that we trust more than ourselves to do this research well because it’s sufficiently aligned.
+ 2 more
-
korrents.com
Simple training interventions turned out to be very effective at steering models towards aligned behaviour.Their words
But the most important lesson is that simple interventions are very effective at steering the model towards more aligned behavior.
-
Their words
This is the hard problem of alignment we need to solve in order to succeed at building superintelligence, and to this day it is an unsolved problem.