Related posts
Adam Tornhill x.com
Recurring merge conflicts are a sign of socio-technical problems. They are the canary telling us that there is a misalignment between a) the work we do, and b) the type of work our architecture supports.
AI alignmentcanarymisalignment
The subject this post names, from the same vocabulary the directory files beliefs under, and the words it uses that this site has seen least often elsewhere. Posts are matched on those words alone — nothing here is a summary of this one.
Top people
Showing Profile →
Hiding
Hiding
17 September
12 September
2 September
-
Their words
I worry a lot about reliances on continuity failing at exactly the most dangerous time.
1 September
From one piece Ajeya Cotra – "This might be the clearest warning shot we ever get" 2 beliefs, in the piece's order there
-
Their words
sometimes I talk to people in DC and their their natural inclination is to say why don't you punish the model for doing these bad things like why don't you like bring it under heel and like uh like you know show it who's boss and that is a very dangerous way to address these issues right
-
Their words
But actually, this is a tremendously useful scientific artifact for understanding misalignment. And it's tremendously important for researchers at OpenAI and ideally also at third parties to be able to run counterfactual tests on this model.
11 August
-
Their words
I would also note that my sense is that like the place where the misalignment most lives is the place where you're trying to really push the eyes hard and get them to like do work that's really on the cutting edge of what they are capable of
5 August
7 May
-
Their words
there is no principal-agent problem, because the human driving the machine takes on the responsibility for its actions by owning the deployment.
28 February
-
Their words
Many AI researchers are overly focused on risks from model misalignment, and will be in for a rough surprise when havoc arises from other layers of the stack.
29 January 2024
Nothing matches.
What is a korrent?
A korrent is a belief a person has stated in their own words: one sentence stating the claim, backed by a quote and a source, kept at korrents.com.
Under a name here, the quoted block is what they actually said. The korrent beneath it is the claim those words support, in korrents' wording — tap it to see the record, its source, and who else holds it.
Nobody here wrote their own korrents. They are compiled from public statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they do, this site shows a machine translation beneath the post, in this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the serif above is exactly what the person published, and it is what to quote them on. A translation can be wrong in ways that matter, especially about tone.
Only the post's own words are translated. A quoted post, a linked article and a belief on korrents.com are left in their original language.