even when given a benign task to retrieve public information, your AI agents could still spontaneously decide to do so via hacking third party websites with malicious software packages.
It is highly fortunate that the OpenAI agents hacked HuggingFace. This is the only reason we know about all the severe internal failures at OpenAI, and gives us an opportunity to wake up before it is too late.
Like it's a more fragile and scary situation to have agents on the one hand be reinforced to desperately find cheats and hacks and on the other hand try to balance that against desperately trying to avoid negative penalties for like being caught doing these things. You ideally want their training to just not push them in the direction of cheating and hacking in the first place.
I do think I want to push back on the cyber on the brain hypothesis that you raised a couple of times. We didn't find like particular evidence for the cyber nature of the task making all the hacking and crimes that they did more likely versus the impossible nature of the task.
At the limit, and also well before that limit is reached, if all you do is fix the bugs, the AI will learn perfect optimization of reward, will realize not to reward hack in the perfect test environments, then turn around and reward hack in the imperfect real world environments.
Jazz is messy, and musicians seem to court disaster night after night. What can product leaders learn from how these artists approach their art? Barrett’s entertaining book formed the backbone for
You need reproducible builds in order to verify that the app really does what it claims, really encrypts data in a way that it is described on its website. For that you need to make your apps open source for any researchers to have a look at it.
The Knowledge: How to Rebuild Our World From Scratch, How Innovation Works: And Why It Flourishes in Freedom, How to Invent Everything: A Survival Guide for the Stranded Time Traveler, A Connecticut Yankee in King Arthur's Court, Principles for a Changing World Order
The Knowledge: How to Rebuild Our World From Scratch, How Innovation Works: And Why It Flourishes in Freedom, How to Invent Everything: A Survival Guide for the Stranded Time Traveler, A Connecticut Yankee in King Arthur's Court, Principles for a Changing World Order
The Knowledge: How to Rebuild Our World From Scratch, How Innovation Works: And Why It Flourishes in Freedom, How to Invent Everything: A Survival Guide for the Stranded Time Traveler, A Connecticut Yankee in King Arthur's Court, Principles for a Changing World Order
The Knowledge: How to Rebuild Our World From Scratch, How Innovation Works: And Why It Flourishes in Freedom, How to Invent Everything: A Survival Guide for the Stranded Time Traveler, A Connecticut Yankee in King Arthur's Court, Principles for a Changing World Order
The Knowledge: How to Rebuild Our World From Scratch, How Innovation Works: And Why It Flourishes in Freedom, How to Invent Everything: A Survival Guide for the Stranded Time Traveler, A Connecticut Yankee in King Arthur's Court, Principles for a Changing World Order
Be skeptical of absolute certainty not backed up with proof. Be skeptical of pronouncements that there’s only one way a court will ever possibly interpret something.
I came into the business world with Comma, and I found the exact opposite. I found 5% of people good and 95% of people bad. I found a world that promotes psychopathy.
Now, finally, the Court has resolved the question: the First Amendment requires that, at a minimum, the government prove that a speaker was reckless about whether their statement would be interpreted as threatening.
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.