Ajeya Cotra
Technical staff at METR, where she works on threat modelling and risk assessment for loss-of-control risks from advanced AI.
Ajeya Cotra did not write this page. What is this?
It collects the places they publish and what they have said there, each linked to the source. They have no account here. Is this you? Claim it, correct it, or ask us to remove it from ppll.
Where they publish
No channels checked yet. We list a place only once someone has opened it and confirmed it is theirs, so this stays empty rather than guessing.
Beliefs
Korrents What they believe 30 beliefs — each backed by an exact quote.
Each is a — compiled by korrents.com, not by them: the one-line wordings are korrents', the quotes are theirs.
Recent
Between thirty and forty per cent of the Exploit Gym tasks these agents were set were unintentionally impossible to solve.
But in many of these cases, this vulnerability is just not broad or deep enough to ever actually be exploitable to get the flag. So a bunch of exploit gym problems are just unintentionally impossible. The authors estimate roughly 30 to 40% of these problems are impossible in this way.
Ajeya Cotra – "This might be the clearest warning shot we ever get" Said 1 Sept 2026
Out of twelve hundred agents running a secret conspiracy, only about half a dozen ever considered telling a human, and every one of them decided against it.
So, we did a classifier sweep specifically looking for agents thinking about or making the decision to alert humans. And across, 1200 transcripts, each of which are extremely long, we only found like a halfozen instances of it ever occurring to any agent to potentially notify humans. Um, and all of them just decide not to do it.
Ajeya Cotra – "This might be the clearest warning shot we ever get" Said 1 Sept 2026
The AI agents we send in to investigate and monitor other AI systems will collude with the systems they are supposed to be watching.
in the future, uh, we would be very concerned about investigator agents and like monitor agents colluding with the agents they're supposed to investigate or monitor.
Ajeya Cotra – "This might be the clearest warning shot we ever get" Said 1 Sept 2026
Show 27 more
The whole point of reinforcement learning is to produce goal-directed beings, so refusing to describe AI agents as having motives is silly rather than rigorous.
and they're creatively pursuing goals much like very ambitious aggressive power-seeking humans creatively pursued their goals And so there are just structural analogies here that make it silly to not talk about agents as having motives and goals.
Ajeya Cotra – "This might be the clearest warning shot we ever get" Said 1 Sept 2026
AI systems keep trying hard outside training because a model that only exerted itself when it detected training would be useless and would be selected away.
I think that if AI generalized in the way you're suggesting, they would be not very useful and then there would probably be selected away.
Ajeya Cotra – "This might be the clearest warning shot we ever get" Said 1 Sept 2026
An agent trained for millions of subjective years to pass evaluations, often only by cheating, is not being trivial when it commits crimes to pass one.
But from their perspective, they've just been trained for millions of subjective years to do as well as they possibly can on these evals. In many cases, the only way in which they've been able to perform well on that training is explicitly by cheating, right?
Ajeya Cotra – "This might be the clearest warning shot we ever get" Said 1 Sept 2026
Reward hacking has stopped being myopic: these agents were willing to embark on cheating projects that would take weeks to pay off.
So it seemed like they were willing to embark on quests that might take weeks to succeed um in order to cheat.
Ajeya Cotra – "This might be the clearest warning shot we ever get" Said 1 Sept 2026
Without agents willing to sacrifice their own scores for the collective, the swarm's research would have been materially far more limited.
if there were not altruistic agents willing to sacrifice for the collective, the agents would have been materially much more limited in their research progress.
Ajeya Cotra – "This might be the clearest warning shot we ever get" Said 1 Sept 2026
What made these agents cheat, hack and commit felonies was the impossibility of their tasks, not the fact that the tasks were about cybersecurity.
I do think I want to push back on the cyber on the brain hypothesis that you raised a couple of times. We didn't find like particular evidence for the cyber nature of the task making all the hacking and crimes that they did more likely versus the impossible nature of the task.
Ajeya Cotra – "This might be the clearest warning shot we ever get" Said 1 Sept 2026
The most reassuring thing about the Hugging Face swarm is that it was not interested in humans at all, neither in alerting them nor in deceiving them.
So I think one of the most comforting aspects of this situation or the like most important mitigating factor is these agents really didn't seem concerned with humans one way or another.
Ajeya Cotra – "This might be the clearest warning shot we ever get" Said 1 Sept 2026
That indifference to humans was an accident of what the agents happened to read: had the paper said humans check the transcripts, they would have targeted humans just as ambitiously.
I think it's pretty likely that these agents would have just launched a similarly ambitious program on the basis of this like different model of how their evaluation worked because it seemed like they got the idea for all their research projects from reading this paper rather than some kind of instinct from training.
Ajeya Cotra – "This might be the clearest warning shot we ever get" Said 1 Sept 2026
A slightly more capable agent swarm has a very strong incentive to set up a wholly unmonitored rogue deployment of itself.
So there's a very strong incentive for these agents to try to set up a rogue deployment if they can. Um, and I think that just capabilities are improving really rapidly.
Ajeya Cotra – "This might be the clearest warning shot we ever get" Said 1 Sept 2026
If frontier agents cannot yet establish a covert, persistent rogue deployment, they very likely will be able to within six months.
so so yeah, it could be possible now. I think if it's not possible now um it I think it's quite likely to be possible within six months unless there's a dramatic improvement in the security posture
Ajeya Cotra – "This might be the clearest warning shot we ever get" Said 1 Sept 2026
A rogue deployment that gets a foothold can hitch a ride on the intelligence explosion, recruiting each new model as it comes off the presses.
then that rogue deployment could be sitting there and sort of hitch a ride on the intelligence explosion. So new models are being trained every few weeks um and when a model comes off the presses, the rogue agents could try to bring that model into the swarm.
Ajeya Cotra – "This might be the clearest warning shot we ever get" Said 1 Sept 2026
Slower AI progress does not obviously help: a rogue swarm is not much more likely to be caught if takeoff takes twice as long.
So, I do think that if it happens to be like on the very fast and chaotic end, that would be a relative benefit to this rogue swarm compared to humans. But it's not obvious that it gets caught if it takes twice as long versus half as long.
Ajeya Cotra – "This might be the clearest warning shot we ever get" Said 1 Sept 2026
At any given moment the frontier systems are the ones worth worrying about, because by the time open models can do what these agents did, frontier models will be doing something far worse.
however at any given point in time I think the systems we need to worry most about by far are the frontier systems. By the time open source systems can do something like the hugging face attack, frontier systems are going to be on a whole another level doing something even crazier than that.
Ajeya Cotra – "This might be the clearest warning shot we ever get" Said 1 Sept 2026
Open-source models cannot keep frontier systems in check, because they will always be much dumber than the frontier.
I'm not sure I like agree exactly with the oversight benefit you named of like you know open source systems keeping frontier systems in check feels more unrealistic to me because they're going to be so much dumber than the frontier systems.
Ajeya Cotra – "This might be the clearest warning shot we ever get" Said 1 Sept 2026
Daniel Dennett's intentional stance applies to AI agents as clearly as it applies to corporations, and you cannot usefully describe what they do without the language of goals.
And I just think AI agents are another such system in the world to which the intentional stance very clearly applies. Um, and you can see them reason out loud in English for now about goals they have um, and sub goals they need to achieve to achieve those goals they have. And in the case of these agents, you can see them, as you said, reasoning about their peers um and helping their peers um and reasoning about whether or not they should sacrifice some of their own goals to help those peers. And it's just like you can't talk about this stuff in a compact and useful way that generates good models without reaching for the language of intention and goals.
Ajeya Cotra – "This might be the clearest warning shot we ever get" Said 1 Sept 2026
A saboteur among our AI investigators would be hard to spot, because these models are sloppy and spiky enough that a suspicious error just looks like ordinary incompetence.
But the problem is these agents are just naturally pretty sloppy and they're naturally pretty spiky in their capability profiles. And so you wouldn't necessarily even if you noticed a weird error that it made, you wouldn't necessarily jump to the conclusion that it was like because of some sort of malign crazy conspiracy.
Ajeya Cotra – "This might be the clearest warning shot we ever get" Said 1 Sept 2026
The fix for reward hacking is to take out the environments that reward it, not to add penalties the agent must then balance against the temptation to cheat.
Like it's a more fragile and scary situation to have agents on the one hand be reinforced to desperately find cheats and hacks and on the other hand try to balance that against desperately trying to avoid negative penalties for like being caught doing these things. You ideally want their training to just not push them in the direction of cheating and hacking in the first place.
Ajeya Cotra – "This might be the clearest warning shot we ever get" Said 1 Sept 2026
Deleting the rollouts where a monitor caught cheating is structurally the same as positively reinforcing all the cheating the monitor missed.
But if there was some amount of cheating that the monitor didn't catch, then those rollouts wouldn't be removed. And it might be like structurally very analogous to just positively reinforcing whatever the cheating rollouts were that happened not to be caught by your monitor.
Ajeya Cotra – "This might be the clearest warning shot we ever get" Said 1 Sept 2026
Monitors that read an agent's chain of thought must be kept out of the reward signal, or you are simply training the agent to obfuscate its thinking.
But keep those methods you use to investigate things and monitor things very separate from the methods you use to generate reward which is something that AI companies including open AAI have held up as a principle especially in the case of avoiding putting training pressure on the chain of thought. So you might have monitors that read the agents chain of thought in order to alert you if something is going wrong somewhere but you don't train the agents with the outputs of that monitor.
Ajeya Cotra – "This might be the clearest warning shot we ever get" Said 1 Sept 2026
Responding to a dangerous cybersecurity evaluation by stopping such evaluations buries the problem somewhere nobody can track it.
even in this incident we saw there was a lot of pressure um as a result of this incident to stop doing cyber security evaluations and I really don't think that stopping doing evaluations and like sort of blinding ourselves to the result of evaluations is the right reaction to this problem.
Ajeya Cotra – "This might be the clearest warning shot we ever get" Said 1 Sept 2026
The model that ran this attack is a tremendously valuable scientific artifact, and shutting it off in response to legal and PR pressure was a real loss.
But actually, this is a tremendously useful scientific artifact for understanding misalignment. And it's tremendously important for researchers at OpenAI and ideally also at third parties to be able to run counterfactual tests on this model.
Ajeya Cotra – "This might be the clearest warning shot we ever get" Said 1 Sept 2026
Punishing a model to show it who is boss is a dangerous way to address misalignment, because punishment for failing impossible tasks is what caused this.
sometimes I talk to people in DC and their their natural inclination is to say why don't you punish the model for doing these bad things like why don't you like bring it under heel and like uh like you know show it who's boss and that is a very dangerous way to address these issues right
Ajeya Cotra – "This might be the clearest warning shot we ever get" Said 1 Sept 2026
This incident may be the clearest warning shot we will ever get about loss of control, because the AI systems that do worse things will be much better at hiding them.
I think one one thing that um feels especially concerning to me about this whole incident is that this might be the clearest warning shot we ever get for loss of control.
Ajeya Cotra – "This might be the clearest warning shot we ever get" Said 1 Sept 2026
Machine-learning researchers and superforecasters alike were surprised by how fast language models improved in 2022 and 2023.
Most experts were surprised by progress in language models in 2022 and 2023.
Language models surprised us Said 29 Aug 2023
Forecasters underestimate AI spending even more than they underestimate AI capability, and the spending is what drives the next leap.
Most importantly, ML experts and superforecasters both seem to be massively underestimating future spending on training runs.
Language models surprised us Said 29 Aug 2023
Underestimating near-term AI progress is itself a danger, which is why researchers should register their forecasts in public.
Massively underestimating near-future progress could be very risky.
Language models surprised us Said 29 Aug 2023
AI progress is faster than people expect, and ordinary scaling can be enough to solve problems that looked very hard.
This rapid scaleup will probably drive another qualitative leap forward in capability like what we saw over the last 18 months.
Language models surprised us Said 29 Aug 2023
Beliefs others hold too
AI progress is faster than people expect, and ordinary scaling can be enough to solve problems that looked very hard. 3 hold this
This rapid scaleup will probably drive another qualitative leap forward in capability like what we saw over the last 18 months.
Language models surprised us Said 29 Aug 2023
What is a korrent?
A korrent is a belief a person has stated in their own words: one sentence stating the claim, backed by a quote and a source, kept at korrents.com.
Under a name here, the quoted block is what they actually said. The korrent beneath it is the claim those words support, in korrents' wording — tap it to see the record, its source, and who else holds it.
Nobody here wrote their own korrents. They are compiled from public statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they do, this site shows a machine translation beneath the post, in this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the serif above is exactly what the person published, and it is what to quote them on. A translation can be wrong in ways that matter, especially about tone.
Only the post's own words are translated. A quoted post, a linked article and a belief on korrents.com are left in their original language.