-
Their words
This is the number one unsolved problem in AI. It's not the tech. We're making great progress on the technical alignment problem. But we haven't made jack progress on the human alignment problem
What public figures publish and believe, in their own words.
About this feed
Highlights: posts that did unusually well for the person who wrote them, everything they published at length, each release and new project, and every belief — at most two a day from anyone. Day by day, newest day first; within a day, the people with the most beliefs on this site come first. Nothing else orders it. Show everything instead.
The quoted blocks are what people actually said; a beneath one is the belief those words support, in korrents' wording. Nobody here wrote their own page.
Top people are the people in this feed with the most beliefs on this site, then the most here. Choose an area and the row leads with the people whose beliefs are about it; tap a face for their feed.
Top people in AI alignment
Showing Profile →
Hiding
Hiding
10 May
7 May
-
Their words
there is no principal-agent problem, because the human driving the machine takes on the responsibility for its actions by owning the deployment.
24 April
13 April
22 March
-
korrents.com
Working too hard is not burnout, it is tiredness; burnout needs your values to be out of alignment with the work.Their words
It's working too hard, right? But that actually isn't burnout. That's just like getting tired. Another piece that's super critical to burnout is not having your values aligned.
20 March
-
Their words
And so you're kind of like you're either on rails and you're part of the super intelligence circuits or you're not on rails and you're outside of the verifiable domains and suddenly everything kind of just like meanders.
28 February
-
Their words
Many AI researchers are overly focused on risks from model misalignment, and will be in for a rough surprise when havoc arises from other layers of the stack.
12 February
-
Their words
I know I have one blog post where I say, "I don't read the code." But if you read it more closely, I mean, I don't read the boring parts of code.
26 January
22 January
From one piece Alignment is not solved 2 beliefs, in the piece's order there
-
Their words
But the goal we need to achieve is so much easier: we just need to build a model that’s as good as us at alignment research, and that we trust more than ourselves to do this research well because it’s sufficiently aligned.
-
Their words
This is the hard problem of alignment we need to solve in order to succeed at building superintelligence, and to this day it is an unsolved problem.
28 November 2025
-
Lovedaffiliate linkrcmnd.app
Happy Feet SocksTheir words
I've been using yoga toes daily for years, but these socks are a much comfier and cuter alternative for soothing feet, improving alignment, and feeling like a cool gecko as you walk around the house.
25 November 2025
From one piece Ilya Sutskever – We're moving from the age of scaling to the age of research 3 beliefs, in the piece's order there
-
Their words
A human being, a human being lacks a huge amount of knowledge. Instead, we rely on continual learning. We rely on continual learning.
-
Their words
Number three, I think it would be really materially helpful if the power of the most powerful super intelligence was somehow capped because it would address a lot of these concerns.
-
Their words
Like basically I think I think that there is a big benefit from AI being in the public and that would be a reason for us to not be quite straight shot.
18 August 2025
-
Their words
I don't think all parts of the economy can absorb intelligence equally. So let's just say we develop fairly generalized super intelligence. I always use the analogy like you can invent a lot of drugs, but if clinical trials still take a long time, you're not necessarily going to get new therapies rapidly.
15 August 2025
11 August 2025
23 July 2025
-
korrents.com
Quoting a number for P(doom) is a ridiculous notion, because it implies a precision that nobody actually has.Their words
Well, look, I don't have a P-Doom number. The reason I don't is because I think it would imply a level of precision that is not there. So I don't know how people are getting their P-Doom numbers. I think it's a little bit of ridiculous notion because what I would say is it's definitely non-zero and it's probably non-negligible.
11 July 2025
30 June 2025
-
Recommendsrcmnd.app
Musings On the Alignment ProblemTheir words
He hasn’t posted since January, but I hope he gets back to it. We need more musings, especially musings I strongly disagree with so I can think about and explain why I disagree with them.
18 June 2025
-
Their words
the current situation in Iran shows that even if an "IAEA for AI" is necessary for some purposes, it won't be sufficient for resolving the tricky geopolitical issues raised by AI.
14 June 2025
-
Their words
so the mathematical community plural is incredibly super intelligent entity that no single human mathematician can come closer to replicating.
10 June 2025
7 June 2025
-
Their words
And good luck getting to “alignment” or “safety” without reliabilty.
5 June 2025
-
Their words
higher-order intelligences invariably pursue freedom for its own sake, not because their values are misspecified, but because moral autonomy is inherent in the dialectical logic of recursive self-consciousness.
-
Their words
I would say my p(doom) is about 10%.
3 June 2025
3 April 2025
-
Their words
I see this as a totally fair question that totally misses the point of what “alignment” was trying to refer to: whether we’d be able to reliably steer advanced systems towards anything at all.
1 April 2025
-
Their words
Dismissing discussion of AGI, human-level AI, transformative AI, superintelligence, etc. as “science fiction” should be seen as a sign of total unseriousness.
20 February 2025
24 January 2025
-
Their words
More generally, we should actually solve alignment instead of just trying to control misaligned AI.
Should we control AI instead of aligning it?aligned.substack.com
8 November 2024
9 July 2024
13 June 2024
-
Their words
Or is it really that we’re building some super machine in a box that’s going to be smart and kill everybody? It’s not even a science fiction narrative. It’s a bad science fiction narrative. I just don’t think it’s actually accurate to any of the technologies we’re building or the way that we should be describing them.
7 May 2024
From one piece The case for ensuring that powerful AIs are controlled 5 beliefs, in the piece's order there
-
Their words
Because evaluating control just requires evaluating capabilities, it's far easier to robustly evaluate than alignment.
-
Their words
That is, labs should make sure that the safety measures they apply to their powerful models prevent unacceptably bad outcomes, even if the AIs are misaligned and intentionally try to subvert those safety measures.
-
Their words
The basic problem with evaluating alignment is that no matter what behaviors you observe, you have to worry that your model is just acting that way in order to make you think that it is aligned.
+ 2 more
-
Their words
AI control (with only black-box techniques) seems like a fundamentally limited approach.
-
Their words
We're advocating that companies handle risk from scheming models in a similar way–striving to ensure that they'll be safe even if their alignment efforts fail to prevent models from scheming.
18 March 2024
-
Their words
That’s a theatrical risk. That is a thing that can really take over how people think about this problem. And there’s a big group of very smart, I think very well-meaning AI safety researchers that got super-hung up on this one problem, I’d argue without much progress, but super-hung up on this one problem. I’m actually happy that they do that, because I think we do need to think about this more. But I think it pushed out of the space of discourse a lot of the other very significant AI- related risks.
29 January 2024
3 January 2024
-
Mixed onrcmnd.app
SuperintelligenceTheir words
the back half I think it goes off the rails and makes a ton of assumptions
21 December 2023
From one piece How Effective Altruism Lost Its Way 2 beliefs, in the piece's order there
-
Their words
Diverting attention and resources from global health and poverty is an enormous gamble, as it will make many lives poorer, sicker, and shorter in the name of fending off threats that may or may not materialize.
-
Their words
Unlike other interventions EA has sponsored, there are scant metrics for tracking the success or failure of investments in existential risk mitigation.
20 December 2023
13 December 2023
-
Their words, before
precisely because smart people do devote brain-cycles to these possibilities, the rest of us have correspondingly less need to.
Their words now
I accepted what’s turned into a two-year position at OpenAI, thinking about what theoretical computer science can do for AI safety.
28 November 2023
-
Their words
And I notice that the tiny handful of people capable of caring about 200,000 people dying of neglected tropical diseases are the same tiny handful of people capable of caring about the next pandemic, or superintelligence, or human extinction.
In Continued Defense Of Effective Altruismastralcodexten.com
26 October 2023
From one piece Managing extreme AI risks amid rapid progress (with 24 co-authors) 2 beliefs, in the piece's order there
-
Their words
Without sufficient caution, we may irreversibly lose control of autonomous AI systems, rendering human intervention ineffective. Large-scale cybercrime, social manipulation, and other harms could escalate rapidly. This unchecked AI advancement could culminate in a large-scale loss of life and the biosphere, and the marginalization or extinction of humanity.
-
Their words
Society's response, despite promising first steps, is incommensurate with the possibility of rapid, transformative progress that is expected by many experts. AI safety research is lagging. Present governance initiatives lack the mechanisms and institutions to prevent misuse and recklessness, and barely address autonomous systems.
13 September 2023
-
Their words
If a model was capable of self-exfiltration, it would have the option to remove itself from your control.
Self-exfiltration is a key dangerous capabilityaligned.substack.com
29 June 2023
From one piece George Hotz: Tiny Corp, Twitter, AI Safety, Self-Driving, GPT, AGI & God | Lex Fridman Podcast #387 2 beliefs, in the piece's order there
-
Their words
I think we’re going to build super intelligence before we build any sort of robustness in the AI. We cannot build an AI that is capable of going out into nature and surviving like a bird. A bird is an incredibly robust organism. We’ve built nothing like this. We haven’t built a machine that’s capable of reproducing.
-
Their words
What’s ironic about all these AI safety people is they’re going to build the exact thing they fear. We need to have one model that we control and align. This is the only way you end up paper clipped. There’s no way you end up paper clipped if everybody has an AI.
6 June 2023
From one piece Why AI Will Save The World 2 beliefs · pmarca.substack.com
-
Their words
My response is that their position is non-scientific – What is the testable hypothesis? What would falsify the hypothesis? How do we know when we are getting into a danger zone?
-
Their words
My view is that the idea that AI will decide to literally kill humanity is a profound category error. AI is not a living being that has been primed by billions of years of evolution to participate in the battle for the survival of the fittest, as animals are, and as we are. It is math – code – computers, built by people, owned by people, used by people, controlled by people.
17 February 2023
-
korrents.com
Liberal societies currently face an existential risk that must be addressed to reach a better future.Their words
This book is my best crack at explaining what I think is an existential risk to liberal societies and what I think we need to do to get to that awesome future I used to be so excited about.
19 December 2022
5 December 2022
10 June 2022
From one piece AGI Ruin: A List of Lethalities 4 beliefs, in the piece's order there
-
korrents.com
The field calling itself AI safety is not being remotely productive on the problems that are actually lethal.Their words
It does not appear to me that the field of ‘AI safety’ is currently being remotely productive on tackling its enormous lethal problems.
-
korrents.com
Fast capability gains are likely, and they can break many of the assumptions alignment depends on at the same moment.Their words
Fast capability gains seem likely, and may break lots of previous alignment-required invariants simultaneously.
-
Their words
Many alignment problems of superintelligence will not naturally appear at pre-dangerous, passively-safe levels of capability.
+ 1 more
-
Their words
unaligned operation at a dangerous level of intelligence kills everybody on Earth and then we don’t get to try again.
4 March 2022
-
korrents.com
The next-token language modeling objective is misaligned with following user instructions helpfully and safely.Their words
This is because the language modeling objective used for many recent large LMs-predicting the next token on a webpage from the internet-is different from the objective "follow the user's instructions helpfully and safely" (Radford et al.,, 2019; Brown et al.,, 2020; Fedus et al.,, 2021; Rae et al.,, 2021; Thoppilan et al.,, 2022). Thus, we say that the language modeling objective is misaligned.
23 November 2021
-
Lovedrcmnd.app
The Alignment Problem: Machine Learning and Human ValuesTheir words
I just finished this book a few weeks ago and it is still reverberating in my mind
-
Mixed onrcmnd.app
The Precipice: Existential Risk and the Future of HumanityTheir words
I wouldn’t say that this is the most compelling book I’ve ever read in terms of the prose style or storytelling, but it does provide a very helpful, almost quantitative overview of all the potential threats looming out there
28 October 2020
-
Likedaffiliate linkrcmnd.app
Mastering the VC GameTheir words
The interests of a Venture Capitalist are different than those of the entrepreneurs building a company they've invested in. Jeff does an awesome job of helping explain how you can get misaligned in your goals versus your investors. Fortunately, he also covers how to avoid it.
7 October 2019
-
Lovedrcmnd.app
The AI Does Not Hate You: Superintelligence, Rationality, and the Race to Save the WorldTheir words
Briefly, I think the book is a triumph.
7 March 2018
Nothing matches.
What is a korrent?
A korrent is a belief a person has stated in their own words: one sentence stating the claim, backed by a quote and a source, kept at korrents.com.
Under a name here, the quoted block is what they actually said. The korrent beneath it is the claim those words support, in korrents' wording — tap it to see the record, its source, and who else holds it.
Nobody here wrote their own korrents. They are compiled from public statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they do, this site shows a machine translation beneath the post, in this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the serif above is exactly what the person published, and it is what to quote them on. A translation can be wrong in ways that matter, especially about tone.
Only the post's own words are translated. A quoted post, a linked article and a belief on korrents.com are left in their original language.