even when given a benign task to retrieve public information, your AI agents could still spontaneously decide to do so via hacking third party websites with malicious software packages.
Nonetheless, ancient rock art is precious, because it may convey a not-so-detailed message that would nonetheless be amazing to understand, if someday we can.
the math community needs to adopt a version of the ethical standards of experimental science. If you are using AI agents, you can't just give a proof (formalized or not), but need to also provide a detailed explanation of how these agents were used to get the result.
I do think I want to push back on the cyber on the brain hypothesis that you raised a couple of times. We didn't find like particular evidence for the cyber nature of the task making all the hacking and crimes that they did more likely versus the impossible nature of the task.
At the limit, and also well before that limit is reached, if all you do is fix the bugs, the AI will learn perfect optimization of reward, will realize not to reward hack in the perfect test environments, then turn around and reward hack in the imperfect real world environments.
There's not this two-tier system where, well, there's Apple or Windows and they make the real thing and then you can try to hack or patch your little thing around it and it's going to look second tier, it always does.
What you would want to do is have a detailed highdimensional readout of okay your your gallbladder and your liver and your kidney and your heart and changes all over your body and look at that and ask is there is there some kind of association between specific emotion states and changes some kind of complex change in your body and I would bet that there is but we just don't have those data yet
And so basically everything that we can verify reasonably well with some feedback loop, the AIS are doing pretty well on. And that's sufficient to make AR and D go quite fast and to continue. But there's some parts of of developing uh aligned and safe AIs that are more subtle, hard to check, depend on, you know, detailed in the weeds things.
One concern you might have is there are like large categories of reward hacks which humans can't detect well and which we consistently fail to detect and which consistently get reinforced and then this category is sufficient to cause the most natural behavior for the AI to learn to be like cheat when the humans can't find out
Um, and to give you an example of a a a use of a coding agent that works extremely well is you can ask today's models to translate software from one computer language to another very effectively because in that case you actually have a incredibly detailed specification.
I genuinely believe that if you took an open weights model from 2025 and built a pentest harness for it, it could do this kind of sandbox escape and scan/hack in most networks. This is only surprising because you assume OpenAI has sounder sandboxes.
To change course in democratic nations will require a major change in public opinion as a signal to political leaders that the current course is not acceptable.
That prickle that feeling that you get, it's like muted now cuz someone else it's kind of like you've made someone else deal with the problem. The problem is still there and the the landmines are still going to blow up on you eventually, but like you're not you don't feel that bad feeling as much anymore. So your judgment is skewed. You're not getting that feedback loop.
That judgment, that ability to have that judgment is so distorted right now because the agent will just do the hacky thing for you. The agent will kind of deal with the hacky problems that come down the road. And it's way easier just to be like, "Oh yeah, it's a temporary fix."
We need to accept that at best, we will just barely avoid some of the worst case scenarios (e.g., an AI-enabled biological weapon that kills billions, a rogue AI takeover, or stable global totalitarianism enabled by AI), given the current pace of AI capabilities relative to the pace of governance.
and we'll gradually layer all this stuff everywhere and there will be fewer and fewer people who understand it and that there will be a sort of this like scenario of a gradual loss of control and understanding of what's happening that to me seems most likely outcome of how all this stuff will go down.
It’s when you really don’t want to do a thing that you imagine a thing you don’t want to do more that’s worse than that and then in that way, you procrastinate by not doing the thing that’s worse. It’s a nice hack, it actually works.
You can't calculate every photon in the scene. You need really detailed approximations, and that's the field of computer graphics. It's about increasingly effective approximations of the laws of physics, which are just totally intractable.
By trying to use misuse as a fig leaf for their real concerns, they end up sounding much less credible than if they just tried to argue for what they actually meant.
And for that reason, there's physical constraints to things like AGI, like recursive improvement to kill us all type stuff. For the physical reasons and for how humans have figured things out before, I'm not too worried about AI takeover.
Here’s one small mental hack that makes a world of difference: remember that you are trying to hire the right people to join your team/org/company.
Not the “best” people.
The right people.
Josh Hardman at Psychedelic Alpha provided a detailed live account of the advisory committee meeting, which I found very helpful in developing a more granular sense of how the meeting unfolded without having to watch it myself.
The proportion of profiteers seems highly correlated with the number of VERITAS plaques on campus and the high-mindedness of the organization's mission statement.
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.