Meaningful learning from incidents requires expertise that most organizations do not have — but these skills can be learned! We're getting a head start on 2025 with a workshop series on Incident Analysis.
In this case, however, the agents did not copy their weights, attempt to procure replacement compute, or take other steps that would be rational to take if their objective was to survive shutdown. So while the agents in the OpenAI-Hugging Face Incident were rogue, they were not truly sovereign.
even in this incident we saw there was a lot of pressure um as a result of this incident to stop doing cyber security evaluations and I really don't think that stopping doing evaluations and like sort of blinding ourselves to the result of evaluations is the right reaction to this problem.
I think one one thing that um feels especially concerning to me about this whole incident is that this might be the clearest warning shot we ever get for loss of control.
Um so I think it's an alignment failure. I think it's a security failure. I think it's like a very serious thing even though it's you know not not the biggest example of consequence.
More data for better comparisons are good. Now everybody has to do it, and regulators and the public have to learn to look at only the data, and not individual incidents
Scrutiny typically either jumps several points on any 10-point scale after an incident that kills a bunch of people, or increases very slowly over decades.
Scrutiny applied to AI is way below other technologies that - even if you totally ignore the doomers - could put far fewer lives at risk than AI in a single incident.
When I have conversations with breached companies, my messaging is crystal clear: be transparent and expeditious in your reporting of the incident and prioritise communicating with your customers.
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.