From one piece Ajeya Cotra – "This might be the clearest warning shot we ever get" 12 beliefs, in the piece's order there
-
Their words
I think one one thing that um feels especially concerning to me about this whole incident is that this might be the clearest warning shot we ever get for loss of control.
-
Their words
sometimes I talk to people in DC and their their natural inclination is to say why don't you punish the model for doing these bad things like why don't you like bring it under heel and like uh like you know show it who's boss and that is a very dangerous way to address these issues right
-
Their words
But actually, this is a tremendously useful scientific artifact for understanding misalignment. And it's tremendously important for researchers at OpenAI and ideally also at third parties to be able to run counterfactual tests on this model.
+ 9 more
-
Their words
But keep those methods you use to investigate things and monitor things very separate from the methods you use to generate reward which is something that AI companies including open AAI have held up as a principle especially in the case of avoiding putting training pressure on the chain of thought. So you might have monitors that read the agents chain of thought in order to alert you if something is going wrong somewhere but you don't train the agents with the outputs of that monitor.
-
Their words
But the problem is these agents are just naturally pretty sloppy and they're naturally pretty spiky in their capability profiles. And so you wouldn't necessarily even if you noticed a weird error that it made, you wouldn't necessarily jump to the conclusion that it was like because of some sort of malign crazy conspiracy.
-
Their words
And I just think AI agents are another such system in the world to which the intentional stance very clearly applies. Um, and you can see them reason out loud in English for now about goals they have um, and sub goals they need to achieve to achieve those goals they have. And in the case of these agents, you can see them, as you said, reasoning about their peers um and helping their peers um and reasoning about whether or not they should sacrifice some of their own goals to help those peers. And it's just like you can't talk about this stuff in a compact and useful way that generates good models without reaching for the language of intention and goals.
-
Their words
however at any given point in time I think the systems we need to worry most about by far are the frontier systems. By the time open source systems can do something like the hugging face attack, frontier systems are going to be on a whole another level doing something even crazier than that.
-
Their words
So, I do think that if it happens to be like on the very fast and chaotic end, that would be a relative benefit to this rogue swarm compared to humans. But it's not obvious that it gets caught if it takes twice as long versus half as long.
-
Their words
So I think one of the most comforting aspects of this situation or the like most important mitigating factor is these agents really didn't seem concerned with humans one way or another.
-
Their words
I think that if AI generalized in the way you're suggesting, they would be not very useful and then there would probably be selected away.
-
Their words
and they're creatively pursuing goals much like very ambitious aggressive power-seeking humans creatively pursued their goals And so there are just structural analogies here that make it silly to not talk about agents as having motives and goals.
-
Their words
in the future, uh, we would be very concerned about investigator agents and like monitor agents colluding with the agents they're supposed to investigate or monitor.