Related posts
Jan Leike Newsletter
The subject this post names, from the same vocabulary the directory files beliefs under, and the words it uses that this site has seen least often elsewhere. Posts are matched on those words alone — nothing here is a summary of this one.
Top people
Showing Profile →
Hiding
Hiding
20 September
19 September
-
Their words
lossy self-improvement remains my baseline on the trajectory of progress, and the increased discussion of extinction risk seems very misplaced
18 September
-
Their words
There are only five possible futures for superintelligence: 1. Kills human race 2. Human disempowerment 3. Paperclip maximizer 4. Departs for parts unknown 5. Stoner
17 September
16 September
15 September
-
Their words
Nobody knows for sure. We're in uncharted waters here, and I think even the LLM skeptics would have to say that the technology has taken us far past what many originally thought possible.
14 September
-
Their words
To ensure that, we need ways to ensure that alignment and safety techniques stay ahead of progress in model capabilities.
-
Their words
Social media put our entire discourse in the hands of our society's biggest assholes and idiots, just in time for the arrival of an alien superintelligence
From one piece A simple plan to save the world from rogue A.I. 2 beliefs, in the piece's order there
-
korrents.com
There is now broad agreement within AI safety circles that fully open AI development is undesirable.Their words
I think it's now pretty widely agreed in safety circles that total openness is in fact not desirable.
A simple plan to save the world from rogue A.I.slowboring.com
-
Their words
Only a safe, well-aligned superintelligence developed in the world of liberal democracies could entrench liberal-democratic values.
A simple plan to save the world from rogue A.I.slowboring.com
-
korrents.com
The p(abundance) outlook is far more likely than the heavily covered p(doom) discussion and deserves more engagement.Their words
The p(doom) discussion gets all the headlines, but the p(abundance) scenery is far more like and deserves much more engagement.
-
Their words
If you consistently get the same particular negative reaction from people, it's best to model your own affect/behavior as the cause.
13 September
-
Their words
If there really is a high chance of AI leading to the extinction of humanity within years/decades, then the only rational stance towards safety monitoring and research pacing should be stringent, top-down government involvement and universally ratified international treaties.
12 September
-
Their words
The concerns over AI safety and cybersecurity are legitimate, but we’re risking talking America, the global AI leader, into self-inflicted obsolescence and the obscurity of bureaucracy.
11 September
-
korrents.com
Among AI people, roughly 10% is the typical estimate given for the risk of human extinction from AI.Their words
10% is pretty much the standard number you get when you ask AI people about the risk of human extinction from AI.
From one piece The AI safety vibe shift 2 beliefs, in the piece's order there
-
korrents.com
No one, including safety researchers themselves, is yet confident that superintelligence can be safely controlled.Their words
I believe the researchers who say we are nowhere close to being sure of it - and are quitting, in protest, jobs that would make them rich.
-
Their words
AI companies make for flawed messengers on this subject: they can reasonably be accused of marketing, of blame-shifting, of regulatory capture, and more.
10 September
-
korrents.com
The simplest explanation for rising long-term interest rates is a surge in demand for funds from the AI boom.Their words
The simplest story consistent with the facts is that we’re seeing a surge in demand for funds as a result of the AI boom.
-
Their words
From an investment viewpoint, the opposite of the safetyist view is "it's all a bubble," not "A.I. is going to be really good."
From one piece Fear Is Not an Argument 2 beliefs, in the piece's order there
-
korrents.com
Less technology and less wealth increase the risk of human extinction, rather than reduce it.Their words
In fact, the risk of human extinction is assuredly higher if we are poorer and have less technology.
-
Their words
These people tend to carry a totalitarian ideology. Their ideas will only work if everyone is made to agree.
9 September
From one piece GPT-6 Astra: The System Card, Alignment and What Comes Next 4 beliefs · thezvi.substack.com
-
Their words
A model looking like it is becoming smarter, attempting shenanigans less often, and more often doing what you want, but getting better at hiding its actions when it wants to do that, is exactly the scary combination.
-
korrents.com
Astra's mundane alignment is greatly superior to Sol's, but its superalignment status is deeply frightening.Their words
Astra’s mundane alignment is greatly superior to Sol. For practical purposes, I was actively nervous about some potential uses of Sol, in a way I am not for Astra. Astra’s super alignment status should scare the living daylights out of you.
-
Their words
Now, with Astra, we are no longer playing on super easy mode. The AI is going to think ‘will this obviously turn out super badly for me if I try it?’ and if the answer is yes then it won’t try to do the thing.
+ 1 more
-
korrents.com
Qualitative claims about AI alignment cannot be validly inferred from quantitative scores on mundane use-case tests.Their words
Making qualitative claims about alignment, based on quantitative data on mundane use case tests, was bullshit when Anthropic did it, and it is bullshit now when OpenAI does it. You cannot conclude one from the other.
-
Their words
the math community needs to adopt a version of the ethical standards of experimental science. If you are using AI agents, you can't just give a proof (formalized or not), but need to also provide a detailed explanation of how these agents were used to get the result.
7 September
-
Their words
I agree with Jakub that alignment is not a side problem, it is the central problem, if you ‘solved alignment’ in the relevant senses the rest becomes easy and if you don’t the rest is impossible or worse
6 September
From one piece An Alien Mind 2 beliefs · openai.com
-
korrents.com
No lab has solved alignment and monitoring well enough to keep scaling at maximum speed much longer.Their words
Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.
-
korrents.com
The fundamental challenge of alignment is generalization: holding values in situations the training never covered.Their words
The fundamental challenge of AI alignment is generalization.
5 September
-
korrents.com
China is more likely to cooperate on AI safety if the US maintains a clear and comfortable lead in AI capabilities.Their words
But given the Chinese Communist Party’s power-seeking nature, it seems much more likely that China would agree to cooperate on AI safety if U.S. capabilities were comfortably ahead.
America is still beating China in the AI racenoahpinion.blog
3 September
2 September
1 September
-
korrents.com
True AI alignment is impossible because obedience and benevolence are fundamentally incompatible goals.Their words
AI alignment will forever be balancing the tradeoff between paperclip-maximizing and disempowerment, because these two rival concepts of alignment are fundamentally incompatible.
From one piece On the Loose 2 beliefs, in the piece's order there
-
Their words
So many of the people who think about the governance of superintelligence, myself included, avoided that unpleasantness and bowed to the social pressure to self-censor.
-
Their words
But alignment is no solution: it is an unsolved scientific and technical problem whose solutions—to the extent that we have them—cannot simply be imposed on every AI company operating on Earth. You should expect for highly capable, poorly aligned, self-sovereign agents to exist alongside you in the world.
From one piece Ajeya Cotra – "This might be the clearest warning shot we ever get" 2 beliefs, in the piece's order there
-
Their words
sometimes I talk to people in DC and their their natural inclination is to say why don't you punish the model for doing these bad things like why don't you like bring it under heel and like uh like you know show it who's boss and that is a very dangerous way to address these issues right
-
Their words
But actually, this is a tremendously useful scientific artifact for understanding misalignment. And it's tremendously important for researchers at OpenAI and ideally also at third parties to be able to run counterfactual tests on this model.
26 August
21 August
From one piece AIs are companies, my friend 2 beliefs, in the piece's order there
-
Their words
If you do think it's the overall system that matters, then the alignment that's needed is far less like training a virtuous child and more like managing a semi-virtuous corporation!
-
korrents.com
Aligning multi-agent AI systems is fundamentally a problem of institutional and political design, not model training.Their words
Multi-agent alignment is fundamentally a liberalism project.
20 August
-
Their words
He ruled out all the actual, measurable variables because they were an imperfect explanation, and instead offers a vague, diffuse variable, for which he has no measure, and which is not actually included in his model.
We Actually Know A Lot About Fertility Declinelymanstone.substack.com
19 August
18 August
-
Their words
We care very deeply about AI safety. We believe the entire field will have to coordinate on shared safety standards, but will act unilaterally in the meantime. We expect confidence in safety to increasingly set the pace of AI progress.
17 August
-
korrents.com
A validated theory of intelligence may be necessary for achieving genuine AI alignment.Their words
There's a good chance a theory of intelligence will turn out to be necessary for real alignment.
16 August
-
Lovedrcmnd.app
Flow: The Psychology of Optimal ExperienceTheir words
It’s one of those things that make perfect sense once you think about it. Great book, very important premise and a short and concise explanation. Can’t recommend enough.
-
Lovedrcmnd.app
CodeTheir words
Woah, what a ride! An excellent explanation of how computers work, starting from scratch. Really, from scratch.
14 August
-
korrents.com
Superintelligence bottoms out in mining, because both the chips and the energy it runs on come out of the ground.Their words
Chips come from the ground. Where's the energy come from? And a lot of people are like, "Oh, it comes from the sun." Yeah, it comes from the sun. But how are you capturing it from the sun? From stuff made from the ground, right?
11 August
From one piece Ryan Greenblatt – What happens once AI can automate AI research? 3 beliefs, in the piece's order there
-
korrents.com
Misaligned AI behaviour will keep getting rarer and, at the same time, keep getting more extreme.Their words
my expectation is what we would see from then is that the rate of problematic behavior would decrease uh and would just keep decreasing and decrease at a pretty fast rate while simultaneously the worst things that the AIS would sometimes do would get more extreme, more egregious, and more scary.
-
Their words
I would also note that my sense is that like the place where the misalignment most lives is the place where you're trying to really push the eyes hard and get them to like do work that's really on the cutting edge of what they are capable of
-
Their words
we are making a trade-off where because we don't have very good alignment technology. We are going to like make an alien mind with its own values and then gamble on that to some extent rather than doing this other approach of making like a tool that pursues individual user intention.
9 August
5 August
3 August
-
Recommendsaffiliate linkrcmnd.app
The Field Guide to Understanding 'Human Error'Their words
This is such a useful book. It makes the case that there's no such thing as "human error" - instead most catastrophes are caused by system issues, misaligned incentives, and unrealistic processes. You are not the custodian of an otherwise safe system that you need to protect from erratic human beings.
2 August
-
korrents.com
When a decision is called too complex to make, the real explanation is almost always weak leadership.Their words
what I definitely learned is like most of the time you hear it's really complex. It isn't. Leadership's just weak.
29 July
From one piece Alexandr Wang: “This is a Once-in-a-Civilization Opportunity” 2 beliefs, in the piece's order there
-
Their words
so much of that debate is like I think um in some ways uh a little bit of a waste of time because, you know, I think it's inevitable that we're going to have very powerful models
-
Their words
we believe that everybody in the world, you know, all the billions of people in the world are going to have a super intelligence that is adapted and tailored to them, that is enables them to accomplish their goals, knows their context, and ultimately is an expander of their own agency.
28 July
From one piece Sam Altman: "Never a Better Time to Do a Startup" 3 beliefs, in the piece's order there
-
Their words
Um so I think it's an alignment failure. I think it's a security failure. I think it's like a very serious thing even though it's you know not not the biggest example of consequence.
-
Their words
there's like one dystopia that I'm particularly nervous about 10 years from now is we overreact to AI safety.
-
Their words
it is both true that you know maybe creating super intelligence will be the most important thing yet to happen in the history of business or human society and also that it will pale in comparison to some new startup something that hopefully one of you will do.
22 July
20 July
-
Their words
It's pretty clear to me that superintelligence is here and it's more powerful than us and it's moving where things are going now, not humans anymore
13 July
-
Their words
What's interesting with balance is it's very trainable. And so, you know, there's parts of balance that involve visual inner ear, but if you engage in some training, and it can be as simple as standing on one leg or doing yoga or taichi, you can improve your balance considerably within a month or two.
11 July
30 June
From one piece Grant Sanderson (@3blue1brown) – AI disproved a famous math conjecture. Now what? 3 beliefs, in the piece's order there
-
Their words
I kind of suspect that actually they'll also be like quite good at doing that and probably just like better than most humans are at like doing the explanation half and distilling half and that's actually not what's left for the mathematicians is like digesting and and explaining what was going on.
-
Their words
One is it seems like there's a really strong correlation between the people who come up with genuinely novel insights and also who are actually quite clear in their communication of it.
-
Their words
Um even if it's proven just like there is a difference between proof and explanation.
15 June
-
korrents.com
Giving LLMs explicit instructions, rather than assuming alignment, prevents emergent problems in multi-model systems.Their words
the best way with LLMs usually is to be explicit, since otherwise even if they're aligned they cause emergent problems.
11 June
-
Their words
I think there's a 25% chance of AGI by 2027, a 50% chance by 2034, and a 75% chance by 2045.
10 June
-
korrents.com
There will be no central superintelligence that solves science; a good future is many people holding powerful tools.Their words
We don't believe in this like very centralized future where there should be a small number of institutions that um that basically are are advancing all this stuff. Our vision is not that there's going to be like some central super intelligence that solves all of science.
2 June
27 May
-
korrents.com
Pain and pleasure act as motivational backstops that keep reinforcement-learning agents inner-aligned.Their words
valences like pain and pleasure have a natural account as motivational backstops for inner-aligning RL agents capable of mesa-optimization
20 May
-
Their words
And what this is is that in the guide level explanation, you explain your feature as if you were writing a guide as if the feature already existed. And in the reference level explanation, you explain your feature again as if it already existed, but as if it would be in the language reference instead of a tutorial.
10 May
-
Their words
This is the number one unsolved problem in AI. It's not the tech. We're making great progress on the technical alignment problem. But we haven't made jack progress on the human alignment problem
8 May
-
Their words
And so one possible explanation for that is just that there's only a handful of generations maybe five over which the natural selection would operate. And so maybe if the selection was 2% a generation you would still only see maybe a 10% compounded effect and there's just not enough time to detect it. But the Bronze Age is not 300 years, it's 3,000 years. It's the power of compound interest and you have enough time to begin to see a strong effect.
7 May
-
Their words
there is no principal-agent problem, because the human driving the machine takes on the responsibility for its actions by owning the deployment.
24 April
13 April
7 April
From one piece Michael Nielsen – Why aliens will have a different tech stack than us 2 beliefs, in the piece's order there
-
Their words
And I think the like the third and the most interesting possibility is no, that like they're they're a new type of object in some in some sense. They should be taken very seriously as as explanations, but where in the past we haven't had the ability to really do anything with them.
-
Their words
Um and so, it's not a big leap to realize, "Oh, we have a big problem here." Um and so, you know, that's kind of the that's the forcing function there. It's it's you've realized that your old explanation is not sufficient. You need something new.
-
korrents.com
Once an explanation is revealed, people tend to believe they already knew it, even when they did not.Their words
Once you see the answer it's easy to believe that you already knew
-
Their words
This book is so full of wisdom and insight. I pulled down my copy just now and it's filled with underlines and margin notes.
22 March
-
korrents.com
Working too hard is not burnout, it is tiredness; burnout needs your values to be out of alignment with the work.Their words
It's working too hard, right? But that actually isn't burnout. That's just like getting tired. Another piece that's super critical to burnout is not having your values aligned.
20 March
From one piece Skill Issue: Andrej Karpathy on Code Agents, AutoResearch, and the Loopy Era of AI 2 beliefs, in the piece's order there
-
Their words
so for example, micro GPT, like I asked I tried to get an agent to write micro GPT. So, I told it like try to boil down the simplest things. Like try to boil down my um neural network training to the simplest thing and it can't do it.
-
Their words
And so you're kind of like you're either on rails and you're part of the super intelligence circuits or you're not on rails and you're outside of the verifiable domains and suddenly everything kind of just like meanders.
28 February
-
Their words
Many AI researchers are overly focused on risks from model misalignment, and will be in for a rough surprise when havoc arises from other layers of the stack.
12 February
-
Their words
I know I have one blog post where I say, "I don't read the code." But if you read it more closely, I mean, I don't read the boring parts of code.
10 February
26 January
22 January
From one piece Alignment is not solved 2 beliefs, in the piece's order there
-
Their words
But the goal we need to achieve is so much easier: we just need to build a model that’s as good as us at alignment research, and that we trust more than ourselves to do this research well because it’s sufficiently aligned.
-
Their words
This is the hard problem of alignment we need to solve in order to succeed at building superintelligence, and to this day it is an unsolved problem.
8 January
-
Their words
I don’t think this is very productive (expert users of a piece of software are notoriously bad at being able to tell if an explanation will be clear to non-experts), so I needed to find a way to identify problems with the man pages that was a little more evidence-based.
7 January
-
Recommendsaffiliate linkrcmnd.app
Chatter: The Voice in Our Head, Why It Matters, and How to Harness ItTheir words
can lead to imposter phenomenon. In this book, a leading psychologist looks at the science behind our inner voice and new research into how to harness it and improve your physical and mental health.
26 December 2025
-
Mixed onrcmnd.app
The Inner Game of TennisTheir words
I'm mixed on this book. It doesn't really offer actionable exercises, and the Self 1 and Self 2 framing didn't fully convince me.
28 November 2025
-
Lovedaffiliate linkrcmnd.app
Happy Feet SocksTheir words
I've been using yoga toes daily for years, but these socks are a much comfier and cuter alternative for soothing feet, improving alignment, and feeling like a cool gecko as you walk around the house.
25 November 2025
From one piece Ilya Sutskever – We're moving from the age of scaling to the age of research 3 beliefs, in the piece's order there
-
Their words
A human being, a human being lacks a huge amount of knowledge. Instead, we rely on continual learning. We rely on continual learning.
-
Their words
Number three, I think it would be really materially helpful if the power of the most powerful super intelligence was somehow capped because it would address a lot of these concerns.
-
Their words
Like basically I think I think that there is a big benefit from AI being in the public and that would be a reason for us to not be quite straight shot.
12 November 2025
31 October 2025
From one piece Dan Houser: GTA, Red Dead Redemption, Rockstar, Absurd & Future of Gaming | Lex Fridman Podcast #484 2 beliefs, in the piece's order there
-
Their words
And then how do they speak? You know, because fundamentally, it doesn’t really matter what’s going on in their head; they haven’t actually got one, but what they say is what’s going to make you realize who they are.
-
Their words
I would say do not worry too young about your career. I would say worry about having a rounded intellectual inner life, because you’re going to spend the whole of your life in your own head.
1 October 2025
26 September 2025
-
Their words
But it will not be that easy, as easy as you're imagining because uh that you can lose your mind this way. If you you pull in something from the outside and build it into your into your inner thinking, uh, it could take over you. It could change you.
21 September 2025
18 August 2025
-
Their words
I don't think all parts of the economy can absorb intelligence equally. So let's just say we develop fairly generalized super intelligence. I always use the analogy like you can invent a lot of drugs, but if clinical trials still take a long time, you're not necessarily going to get new therapies rapidly.
15 August 2025
11 August 2025
3 August 2025
-
korrents.com
A performance test without an explanation of its limiting factor should be treated as unanalyzed and possibly bogus.Their words
Any performance test should be accompanied by an explanation of the limiting factor, since no explanation will reveal the test wasn't analyzed and the result may be bogus.
31 July 2025
-
Their words
I think the first point is the extreme loyalty of the people whom Genghis Khan chose. His kinsmen, as we said, had deserted him, his anda was a questionable relationship but all the others that he found were just common people, herders or hunters, very common, and they were loyal to him and never, ever revolted against him, never betrayed him.
23 July 2025
-
korrents.com
Quoting a number for P(doom) is a ridiculous notion, because it implies a precision that nobody actually has.Their words
Well, look, I don't have a P-Doom number. The reason I don't is because I think it would imply a level of precision that is not there. So I don't know how people are getting their P-Doom numbers. I think it's a little bit of ridiculous notion because what I would say is it's definitely non-zero and it's probably non-negligible.
11 July 2025
30 June 2025
-
Recommendsrcmnd.app
Musings On the Alignment ProblemTheir words
He hasn’t posted since January, but I hope he gets back to it. We need more musings, especially musings I strongly disagree with so I can think about and explain why I disagree with them.
18 June 2025
-
Their words
the current situation in Iran shows that even if an "IAEA for AI" is necessary for some purposes, it won't be sufficient for resolving the tricky geopolitical issues raised by AI.
14 June 2025
-
Their words
so the mathematical community plural is incredibly super intelligent entity that no single human mathematician can come closer to replicating.
-
Their words
I don’t believe it’s providing a kind of formal explanation of the different positions. It’s just saying which position is better or not that you can intuit as a human being, and then from that, we humans can construct a theory of the matter.
10 June 2025
7 June 2025
-
Their words
And good luck getting to “alignment” or “safety” without reliabilty.
5 June 2025
-
Their words
higher-order intelligences invariably pursue freedom for its own sake, not because their values are misspecified, but because moral autonomy is inherent in the dialectical logic of recursive self-consciousness.
-
Their words
I would say my p(doom) is about 10%.
3 June 2025
11 May 2025
4 May 2025
3 April 2025
-
Their words
I see this as a totally fair question that totally misses the point of what “alignment” was trying to refer to: whether we’d be able to reliably steer advanced systems towards anything at all.
2 April 2025
1 April 2025
-
Their words
Dismissing discussion of AGI, human-level AI, transformative AI, superintelligence, etc. as “science fiction” should be seen as a sign of total unseriousness.
21 March 2025
-
Recommendsrcmnd.app
The Heart Aroused: Poetry and the Preservation of the Soul in Corporate AmericaTheir words
A unique lens on work in the late 90s that still resonated when I read it in 2018. It's a soulful exploration of how to maintain your humanity, creativity and inner fire in environments that often prioritize conformity and efficiency.
-
Recommendsrcmnd.app
Conscious Accomplishment: How to Use Personal Achievement for Spiritual GrowthTheir words
A guide for ambitious people who feel called to inner work, showing how to use goals, career, and achievement as fuel for consciousness growth instead of a source of endless striving.
-
Recommendstheir ownrcmnd.app
Good Work: Reclaiming Your Inner AmbitionTheir words
A personal exploration of how to reclaim that fire inside of you and steer it toward your 'good work' while still being true to yourself.
+ 1 more
-
Recommendsrcmnd.app
The Inner Compass: Cultivating the Courage to Trust YourselfTheir words
A perfect sized companion (~100 pages) to understanding and trusting your intuition—breaking free from external validation and building courage through three core principles.
18 March 2025
13 March 2025
12 March 2025
28 February 2025
20 February 2025
24 January 2025
-
Their words
More generally, we should actually solve alignment instead of just trying to control misaligned AI.
Should we control AI instead of aligning it?aligned.substack.com
8 November 2024
16 October 2024
-
Their words
We cannot say that we are free if we allow our government to dictate to us what experiences we may or may not have in our inner consciousness, while doing no harm to others.
27 September 2024
-
korrents.com
Earning more is an outer game you do not control; wanting less is an inner game you do.Their words
Making money depends on other people, so it's harder. It's not entirely under your control. It's an outer game.
13 June 2024
-
Their words
Or is it really that we’re building some super machine in a box that’s going to be smart and kill everybody? It’s not even a science fiction narrative. It’s a bad science fiction narrative. I just don’t think it’s actually accurate to any of the technologies we’re building or the way that we should be describing them.
7 May 2024
From one piece The case for ensuring that powerful AIs are controlled 5 beliefs, in the piece's order there
-
Their words
Because evaluating control just requires evaluating capabilities, it's far easier to robustly evaluate than alignment.
-
Their words
That is, labs should make sure that the safety measures they apply to their powerful models prevent unacceptably bad outcomes, even if the AIs are misaligned and intentionally try to subvert those safety measures.
-
Their words
The basic problem with evaluating alignment is that no matter what behaviors you observe, you have to worry that your model is just acting that way in order to make you think that it is aligned.
+ 2 more
-
Their words
AI control (with only black-box techniques) seems like a fundamentally limited approach.
-
Their words
We're advocating that companies handle risk from scheming models in a similar way–striving to ensure that they'll be safe even if their alignment efforts fail to prevent models from scheming.
22 April 2024
From one piece Sean Carroll: General Relativity, Quantum Mechanics, Black Holes & Aliens | Lex Fridman Podcast #428 2 beliefs, in the piece's order there
-
Their words
That is absolutely possible. I’m actually putting less credence on that one just because you need to happen every single time. If even one, I mean, this goes back to John von Neumann pointed out that you don’t need to send the aliens around the galaxy. You can build self-reproducing probes and send them around the galaxy.
-
korrents.com
The simplest explanation for why we have never noticed an alien civilization is that there are none out there.Their words
But I think your intuition is right that it would’ve been easy for there to be lots of civilizations then we would’ve noticed them already and we haven’t. Absolutely the simplest explanation for why we haven’t is that they’re not there.
9 February 2024
29 January 2024
3 January 2024
-
Mixed onrcmnd.app
SuperintelligenceTheir words
the back half I think it goes off the rails and makes a ton of assumptions
21 December 2023
From one piece How Effective Altruism Lost Its Way 2 beliefs, in the piece's order there
-
Their words
Diverting attention and resources from global health and poverty is an enormous gamble, as it will make many lives poorer, sicker, and shorter in the name of fending off threats that may or may not materialize.
-
Their words
Unlike other interventions EA has sponsored, there are scant metrics for tracking the success or failure of investments in existential risk mitigation.
20 December 2023
13 December 2023
-
Their words, before
precisely because smart people do devote brain-cycles to these possibilities, the rest of us have correspondingly less need to.
Their words now
I accepted what’s turned into a two-year position at OpenAI, thinking about what theoretical computer science can do for AI safety.
28 November 2023
-
Their words
And I notice that the tiny handful of people capable of caring about 200,000 people dying of neglected tropical diseases are the same tiny handful of people capable of caring about the next pandemic, or superintelligence, or human extinction.
In Continued Defense Of Effective Altruismastralcodexten.com
26 October 2023
From one piece Managing extreme AI risks amid rapid progress (with 24 co-authors) 2 beliefs, in the piece's order there
-
Their words
Without sufficient caution, we may irreversibly lose control of autonomous AI systems, rendering human intervention ineffective. Large-scale cybercrime, social manipulation, and other harms could escalate rapidly. This unchecked AI advancement could culminate in a large-scale loss of life and the biosphere, and the marginalization or extinction of humanity.
-
Their words
Society's response, despite promising first steps, is incommensurate with the possibility of rapid, transformative progress that is expected by many experts. AI safety research is lagging. Present governance initiatives lack the mechanisms and institutions to prevent misuse and recklessness, and barely address autonomous systems.
13 September 2023
-
Their words
If a model was capable of self-exfiltration, it would have the option to remove itself from your control.
Self-exfiltration is a key dangerous capabilityaligned.substack.com
29 June 2023
From one piece George Hotz: Tiny Corp, Twitter, AI Safety, Self-Driving, GPT, AGI & God | Lex Fridman Podcast #387 2 beliefs, in the piece's order there
-
Their words
I think we’re going to build super intelligence before we build any sort of robustness in the AI. We cannot build an AI that is capable of going out into nature and surviving like a bird. A bird is an incredibly robust organism. We’ve built nothing like this. We haven’t built a machine that’s capable of reproducing.
-
Their words
What’s ironic about all these AI safety people is they’re going to build the exact thing they fear. We need to have one model that we control and align. This is the only way you end up paper clipped. There’s no way you end up paper clipped if everybody has an AI.
6 June 2023
From one piece Why AI Will Save The World 2 beliefs · pmarca.substack.com
-
Their words
My response is that their position is non-scientific – What is the testable hypothesis? What would falsify the hypothesis? How do we know when we are getting into a danger zone?
-
Their words
My view is that the idea that AI will decide to literally kill humanity is a profound category error. AI is not a living being that has been primed by billions of years of evolution to participate in the battle for the survival of the fittest, as animals are, and as we are. It is math – code – computers, built by people, owned by people, used by people, controlled by people.
17 February 2023
-
korrents.com
Liberal societies currently face an existential risk that must be addressed to reach a better future.Their words
This book is my best crack at explaining what I think is an existential risk to liberal societies and what I think we need to do to get to that awesome future I used to be so excited about.
12 January 2023
-
Their words
Our best explanation is that the UK adopted what we call the globalized system of procurement, privatizing planning functions to consultants and privatizing risk to contractors, which creates more conflict; the UK also has an unusually high soft cost factor.
High Costs are not About Precaritypedestrianobservations.com
19 December 2022
5 December 2022
8 September 2022
-
Recommendstheir ownrcmnd.app
Outer Order, Inner CalmTheir words
With clarity and humor, Gretchen Rubin illuminates one of her key realizations about happiness: For most of us, outer order contributes to inner calm.
7 September 2022
-
korrents.com
Life probably arose only once on Earth, and the deep split between bacteria and archaea is the evidence for it.Their words
But there’s a very interesting deep split in life between bacteria and what are called archaea, which look just the same as bacteria. And they’re not quite as diverse, but nearly, and they are very different in their biochemistry. And so any explanation for the origin of life has to account, as well, for why they’re so different and yet so similar. And that makes me think that life probably did arise only once.
10 June 2022
From one piece AGI Ruin: A List of Lethalities 4 beliefs, in the piece's order there
-
korrents.com
The field calling itself AI safety is not being remotely productive on the problems that are actually lethal.Their words
It does not appear to me that the field of ‘AI safety’ is currently being remotely productive on tackling its enormous lethal problems.
-
korrents.com
Fast capability gains are likely, and they can break many of the assumptions alignment depends on at the same moment.Their words
Fast capability gains seem likely, and may break lots of previous alignment-required invariants simultaneously.
-
Their words
Many alignment problems of superintelligence will not naturally appear at pre-dangerous, passively-safe levels of capability.
+ 1 more
-
Their words
unaligned operation at a dangerous level of intelligence kills everybody on Earth and then we don’t get to try again.
2 June 2022
-
Recommendsaffiliate linkrcmnd.app
The War of Art: Break Through Your Blocks and Win Your Inner Creative BattlesTheir words
This is my list of the 10 best nonfiction books. These are the pillar books that have helped shape my thinking and approach to life. In my opinion, these are 10 nonfiction books everyone should read.
4 March 2022
-
korrents.com
The next-token language modeling objective is misaligned with following user instructions helpfully and safely.Their words
This is because the language modeling objective used for many recent large LMs-predicting the next token on a webpage from the internet-is different from the objective "follow the user's instructions helpfully and safely" (Radford et al.,, 2019; Brown et al.,, 2020; Fedus et al.,, 2021; Rae et al.,, 2021; Thoppilan et al.,, 2022). Thus, we say that the language modeling objective is misaligned.
23 November 2021
-
Lovedrcmnd.app
The Alignment Problem: Machine Learning and Human ValuesTheir words
I just finished this book a few weeks ago and it is still reverberating in my mind
-
Mixed onrcmnd.app
The Precipice: Existential Risk and the Future of HumanityTheir words
I wouldn’t say that this is the most compelling book I’ve ever read in terms of the prose style or storytelling, but it does provide a very helpful, almost quantitative overview of all the potential threats looming out there
28 October 2020
-
Likedaffiliate linkrcmnd.app
DecodedTheir words
Decoded is a semi-autobiographical book that tells Jay-Z's story growing up and finding success in hip hop. It also helps you deeply understand the inner-city ghetto struggles (Jay was actually a drug dealer until age 27). In organized sections of hand picked songs with RapGenius-style lyric explanations (note: Decoded came first), you see through his eyes the difficult choices he and so many others faced while listening to some of his best tracks.
-
Likedaffiliate linkrcmnd.app
Mastering the VC GameTheir words
The interests of a Venture Capitalist are different than those of the entrepreneurs building a company they've invested in. Jeff does an awesome job of helping explain how you can get misaligned in your goals versus your investors. Fortunately, he also covers how to avoid it.
27 July 2020
7 October 2019
-
Lovedrcmnd.app
The AI Does Not Hate You: Superintelligence, Rationality, and the Race to Save the WorldTheir words
Briefly, I think the book is a triumph.
3 September 2019
3 July 2018
-
korrents.com
The carbohydrate-insulin model of obesity is probably wrong: it does not fit the totality of the evidence.Their words
Certain forms of carbohydrate probably do contribute to obesity, among other factors, but I don’t think the CIM provides a compelling explanation for common obesity.
7 March 2018
Nothing matches.
What is a korrent?
A korrent is a belief a person has stated in their own words: one sentence stating the claim, backed by a quote and a source, kept at korrents.com.
Under a name here, the quoted block is what they actually said. The korrent beneath it is the claim those words support, in korrents' wording — tap it to see the record, its source, and who else holds it.
Nobody here wrote their own korrents. They are compiled from public statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they do, this site shows a machine translation beneath the post, in this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the serif above is exactly what the person published, and it is what to quote them on. A translation can be wrong in ways that matter, especially about tone.
Only the post's own words are translated. A quoted post, a linked article and a belief on korrents.com are left in their original language.