if you treat computing as this onion where you just keep peeling back the layers, there's always something interesting behind the scenes. And the more of those layers that you can peel back and understand, I think in some ways the better you can optimize for that world.
it's it's literally we're looking at neurons in the model's brain that light up when prompt injection happens. So the model won't even tell you but we can actually see those neurons and we can figure out and diagnose that it's happening. And then you combine that with the auto mode classifier and with these three layers we just cannot demonstrate prompt injection anymore.
Then you had code review which gave you another level of feedback. You could roll out internally more frequently. And everybody was using Facebook for all kinds of stuff, personal and internal business stuff. So whatever feature you developed, people would start using it immediately. So you get another round of feedback. Then we had this phased roll out process where you'd start rolling your stuff out. If there was a problem, the blast radius would be limited to a a few million people.
So, like with lawyers particularly, you need some entity to back up the product. You need kind of like an ownership of the product. You need somebody to be able to fire or or hire, like licensing issues. There's a lot of like sort of like regulatory layers that are like also going to be keeping even if there's no relational element human in the loop that have nothing to do with like the ability of the human to actually perform the service.
Many AI researchers are overly focused on risks from model misalignment, and will be in for a rough surprise when havoc arises from other layers of the stack.
This is not the same
world that you were part of. The complexity is off the chart; we are hidden
layers and layers under the scaffolding. And we are used everywhere.
I might not seem so, but I think I’m more humble than a lot of physicists. I’m not sure that we’re ever going to get to that bottom level, but I do think we’re going to keep penetrating different layers and get further.
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.