From one piece Ajeya Cotra – "This might be the clearest warning shot we ever get" 26 beliefs, in the piece's order there
-
Their words
I think one one thing that um feels especially concerning to me about this whole incident is that this might be the clearest warning shot we ever get for loss of control.
-
Their words
sometimes I talk to people in DC and their their natural inclination is to say why don't you punish the model for doing these bad things like why don't you like bring it under heel and like uh like you know show it who's boss and that is a very dangerous way to address these issues right
-
Their words
But actually, this is a tremendously useful scientific artifact for understanding misalignment. And it's tremendously important for researchers at OpenAI and ideally also at third parties to be able to run counterfactual tests on this model.
+ 23 more
-
Their words
even in this incident we saw there was a lot of pressure um as a result of this incident to stop doing cyber security evaluations and I really don't think that stopping doing evaluations and like sort of blinding ourselves to the result of evaluations is the right reaction to this problem.
-
Their words
But keep those methods you use to investigate things and monitor things very separate from the methods you use to generate reward which is something that AI companies including open AAI have held up as a principle especially in the case of avoiding putting training pressure on the chain of thought. So you might have monitors that read the agents chain of thought in order to alert you if something is going wrong somewhere but you don't train the agents with the outputs of that monitor.
-
Their words
But if there was some amount of cheating that the monitor didn't catch, then those rollouts wouldn't be removed. And it might be like structurally very analogous to just positively reinforcing whatever the cheating rollouts were that happened not to be caught by your monitor.
-
Their words
Like it's a more fragile and scary situation to have agents on the one hand be reinforced to desperately find cheats and hacks and on the other hand try to balance that against desperately trying to avoid negative penalties for like being caught doing these things. You ideally want their training to just not push them in the direction of cheating and hacking in the first place.
-
Their words
But the problem is these agents are just naturally pretty sloppy and they're naturally pretty spiky in their capability profiles. And so you wouldn't necessarily even if you noticed a weird error that it made, you wouldn't necessarily jump to the conclusion that it was like because of some sort of malign crazy conspiracy.
-
Their words
And I just think AI agents are another such system in the world to which the intentional stance very clearly applies. Um, and you can see them reason out loud in English for now about goals they have um, and sub goals they need to achieve to achieve those goals they have. And in the case of these agents, you can see them, as you said, reasoning about their peers um and helping their peers um and reasoning about whether or not they should sacrifice some of their own goals to help those peers. And it's just like you can't talk about this stuff in a compact and useful way that generates good models without reaching for the language of intention and goals.
-
korrents.com
Open-source models cannot keep frontier systems in check, because they will always be much dumber than the frontier.Their words
I'm not sure I like agree exactly with the oversight benefit you named of like you know open source systems keeping frontier systems in check feels more unrealistic to me because they're going to be so much dumber than the frontier systems.
-
Their words
however at any given point in time I think the systems we need to worry most about by far are the frontier systems. By the time open source systems can do something like the hugging face attack, frontier systems are going to be on a whole another level doing something even crazier than that.
-
Their words
So, I do think that if it happens to be like on the very fast and chaotic end, that would be a relative benefit to this rogue swarm compared to humans. But it's not obvious that it gets caught if it takes twice as long versus half as long.
-
Their words
then that rogue deployment could be sitting there and sort of hitch a ride on the intelligence explosion. So new models are being trained every few weeks um and when a model comes off the presses, the rogue agents could try to bring that model into the swarm.
-
Their words
so so yeah, it could be possible now. I think if it's not possible now um it I think it's quite likely to be possible within six months unless there's a dramatic improvement in the security posture
-
Their words
So there's a very strong incentive for these agents to try to set up a rogue deployment if they can. Um, and I think that just capabilities are improving really rapidly.
-
Their words
I think it's pretty likely that these agents would have just launched a similarly ambitious program on the basis of this like different model of how their evaluation worked because it seemed like they got the idea for all their research projects from reading this paper rather than some kind of instinct from training.
-
Their words
So I think one of the most comforting aspects of this situation or the like most important mitigating factor is these agents really didn't seem concerned with humans one way or another.
-
Their words
I do think I want to push back on the cyber on the brain hypothesis that you raised a couple of times. We didn't find like particular evidence for the cyber nature of the task making all the hacking and crimes that they did more likely versus the impossible nature of the task.
-
Their words
if there were not altruistic agents willing to sacrifice for the collective, the agents would have been materially much more limited in their research progress.
-
Their words
So it seemed like they were willing to embark on quests that might take weeks to succeed um in order to cheat.
-
Their words
But from their perspective, they've just been trained for millions of subjective years to do as well as they possibly can on these evals. In many cases, the only way in which they've been able to perform well on that training is explicitly by cheating, right?
-
Their words
I think that if AI generalized in the way you're suggesting, they would be not very useful and then there would probably be selected away.
-
Their words
and they're creatively pursuing goals much like very ambitious aggressive power-seeking humans creatively pursued their goals And so there are just structural analogies here that make it silly to not talk about agents as having motives and goals.
-
Their words
in the future, uh, we would be very concerned about investigator agents and like monitor agents colluding with the agents they're supposed to investigate or monitor.
-
Their words
So, we did a classifier sweep specifically looking for agents thinking about or making the decision to alert humans. And across, 1200 transcripts, each of which are extremely long, we only found like a halfozen instances of it ever occurring to any agent to potentially notify humans. Um, and all of them just decide not to do it.
-
Their words
But in many of these cases, this vulnerability is just not broad or deep enough to ever actually be exploitable to get the flag. So a bunch of exploit gym problems are just unintentionally impossible. The authors estimate roughly 30 to 40% of these problems are impossible in this way.