1000+ scientists at frontier AI companies are speaking out to warn that the current commercial race leads to unacceptable security risks. I agree with their call for an international effort to develop technical and governance guardrails to ensure a safer way forward.
This is the first time a debate over whether a model ‘was AGI’ felt non-silly. I do not think it is AGI, and I would warn against the dangers of using that label prematurely, but I would not laugh at you for disagreeing.
The safety guardrails in a way make it more dangerous not less I think because after rejecting you it becomes kind of a pedantic child that will just reject whatever If in that moment your server will go down and it feels it's related it could just reject fixing it
But, simply just having your loops build everything without having some guardrails around the blast radius without having guardrails around how you think about quality, I think is a recipe for disaster.
the most powerful actors for whom this is the biggest concern if these guard rails or the constitution or whatever is getting in the way that will just get steamrolled and so the constitution will only be you know hitting the everyday man rather than hitting governments.
In a world of AI with agents operating across multiple systems, wanting source of truth data, the importance of having preferred paved paths that get the most of the benefits and produce some guardrails so we can make sure we're doing good work, common infrastructure, common paved paths, solving problems once with a core set of capabilities becomes more important.
So often with new software paradigms, what looks like inevitability turns out to be just design failure that can be solved with the right guardrails or affordances or system instructions.
We're now doing it in a much heavier way because we find that these like kind of bory enterprisey uh patterns end up being pretty useful because you have a bunch of idiots on your team now. The coding agents are a bunch of idiots and they are going to work 24/7 and they're going to like ship a lot of stuff. So you need way more guardrails than you used to.
The thing that I find interesting is that's not novel. This has been the thing we've always been trying to do forever. How do we get a junior engineer to ship code safely without breaking stuff? Right? How do we make patterns in the codebase? How do we make tests? Like it's it's all the old stuff that we've always wanted to do.
You cannot just reverse that, right? It's just not working. So, I do think that we need to build out the whole guardrails for the reversibility of actions because that actually where things get really really scary.
And so, um I'm seeing across at least a handful of companies that explicit exec sponsorship makes a huge difference in not just using them, but trying new things and feeling safe to fail within, you know, kind of guardrails.
That's why I warn in my security documentation, don't use cheap models. Don't use Haiku or a local model. Even though I, I very much love the idea that this thing could completely run local. If you use a, a very weak local model, they are very gullible. It's very easy to, to prompt inject them.
These systems will be absolutely central to the economy, technology, and national security, and will be capable of so much autonomy that I consider it basically unacceptable for humanity to be totally ignorant of how they work.
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.