Any piece of software that depends on open source (which is almost every piece of software) has a network of human beings who are potential attack vectors - everyone with publishing rights to any of the packages in the dependency network for that software.
If that’s true there are two options: After the Hugging Face and Wiki attacks OpenAI were still unable to review their previous logs and determine that they had previously attacked RubyGems. They knew about the attack on RubyGems and made the decision not to reach out to the RubyGems team about it. Both of these are bad!
Our adversaries have the training techniques, the data, and the will to attack. And they won’t be slowed down with “embedded evaluators.” They’ll likely have embedded accelerators!
however at any given point in time I think the systems we need to worry most about by far are the frontier systems. By the time open source systems can do something like the hugging face attack, frontier systems are going to be on a whole another level doing something even crazier than that.
But actually, this is a tremendously useful scientific artifact for understanding misalignment. And it's tremendously important for researchers at OpenAI and ideally also at third parties to be able to run counterfactual tests on this model.
A less bad version of it is known to have happened, and from the outside it seems likely that worse things have happened internally that we never heard about.
Overall, I’m very surprised at how little media coverage there’s been around the OpenAI / Hugging Face attack. It’s clearly one of the most important things to happen this year.
And once you look at the brain systems, they're distinct brain systems in this instance for fear to something out there in the world versus panic of something happening inside you like having a heart attack or not getting enough air.
I mean, if you do anything that matters in the world, uh you will have a lot of people call you an idiot or just dismiss you. The more you do, the more the better you do, the more they'll attack you.
it's it's literally we're looking at neurons in the model's brain that light up when prompt injection happens. So the model won't even tell you but we can actually see those neurons and we can figure out and diagnose that it's happening. And then you combine that with the auto mode classifier and with these three layers we just cannot demonstrate prompt injection anymore.
um you know, these tools they either succeed or they fail um and they they've been really bad at creating sort of partial progress or identifying intermediate um um stages that you should you should focus on first.
That was the victorious idea. So it's not the drugs. Actually, that idea to go through the Ardennes Mountains, if you think monocausal, you would say that's the reason. That idea was genius, and Hitler immediately understood it, because before, the plan was to attack in the north of Belgium, which is the same as World War I.
His officers didn't think so. High command said, “No, we're not going to attack the West; we're going to lose.” And Hitler was fanatic about it, he really wanted to attack it. They were planning a coup against him in November 1939 just to prevent him ordering the attack on the West, because it would have been a catastrophe for Germany.
I think there's this whole idea, I call it a denial of attention. I think there's an entire attack vector that's going to be happening. We're using LLMs to generate fake bug reports, fake all these things to just actually effectively to demotivate and hurt open source maintainers.
One thing that I don't think a lot of people are talking about as we integrate more and more AI is that prompt injection is an extremely hard thing to defend against because it's not really clear how you defend against it.
What is the weakness of Google is that any ad unit that’s less profitable than a link, or any ad unit that kind of disincentivizes the link click is not in their interest to go aggressive on, because it takes money away from something that’s higher margins.
But that was their official position, and that changed just two years ago when Putin gave a speech and he said that their position had changed, that they will no longer wait to absorb an attack.
Because we have that policy, Launch on Warning, which is exactly like it says. It means the United States will not wait to absorb a nuclear attack. It will launch nuclear weapons in response before the bomb actually hits.
Elon Musk’s sullen yawp amounts to a claim that he has a right for companies to sponsor his speech, no matter what he says. That’s nonsense, both legally and philosophically.
I think the main driving force is that the Palestinians feel oppressed as they should, and that this was a resistance move. They were resisting the Israeli occupation.
So the counterintuitive thing here, I think that the thing that I think should be done, even though it’s very difficult, is that I would recommend that Israel engage in the most conspicuous acts of kindness possible
I I would I would not say that about Google. I would just say that more generally like that is the incumbent temptation. Um it is is basically is basically go from attack mode to defend mode, right? Go from being the pirate to being the navy.
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.