even when given a benign task to retrieve public information, your AI agents could still spontaneously decide to do so via hacking third party websites with malicious software packages.
I do think I want to push back on the cyber on the brain hypothesis that you raised a couple of times. We didn't find like particular evidence for the cyber nature of the task making all the hacking and crimes that they did more likely versus the impossible nature of the task.
At the limit, and also well before that limit is reached, if all you do is fix the bugs, the AI will learn perfect optimization of reward, will realize not to reward hack in the perfect test environments, then turn around and reward hack in the imperfect real world environments.
There's not this two-tier system where, well, there's Apple or Windows and they make the real thing and then you can try to hack or patch your little thing around it and it's going to look second tier, it always does.
One concern you might have is there are like large categories of reward hacks which humans can't detect well and which we consistently fail to detect and which consistently get reinforced and then this category is sufficient to cause the most natural behavior for the AI to learn to be like cheat when the humans can't find out
I genuinely believe that if you took an open weights model from 2025 and built a pentest harness for it, it could do this kind of sandbox escape and scan/hack in most networks. This is only surprising because you assume OpenAI has sounder sandboxes.
So I think the the thing I'm most excited is actually like what we call like iterated loops or like slow loops where we basically have a cron job. We have the loop the the the structure of the loop is really easy. It's like run this llinter fix one thing commit and push and then we run that every night in our GitHub actions and we wake up every morning to one PR that makes the codebase a little bit better.
That prickle that feeling that you get, it's like muted now cuz someone else it's kind of like you've made someone else deal with the problem. The problem is still there and the the landmines are still going to blow up on you eventually, but like you're not you don't feel that bad feeling as much anymore. So your judgment is skewed. You're not getting that feedback loop.
That judgment, that ability to have that judgment is so distorted right now because the agent will just do the hacky thing for you. The agent will kind of deal with the hacky problems that come down the road. And it's way easier just to be like, "Oh yeah, it's a temporary fix."
It’s when you really don’t want to do a thing that you imagine a thing you don’t want to do more that’s worse than that and then in that way, you procrastinate by not doing the thing that’s worse. It’s a nice hack, it actually works.
Here’s one small mental hack that makes a world of difference: remember that you are trying to hire the right people to join your team/org/company.
Not the “best” people.
The right people.
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.