Related posts
Jan Leike Newsletter
The subject this post names, from the same vocabulary the directory files beliefs under, and the words it uses that this site has seen least often elsewhere. Posts are matched on those words alone — nothing here is a summary of this one.
Top people
Showing Profile →
Hiding
Hiding
19 September
-
Their words
lossy self-improvement remains my baseline on the trajectory of progress, and the increased discussion of extinction risk seems very misplaced
18 September
-
Their words
There are only five possible futures for superintelligence: 1. Kills human race 2. Human disempowerment 3. Paperclip maximizer 4. Departs for parts unknown 5. Stoner
17 September
16 September
15 September
-
Their words
Someone needs to define the unit, perhaps using a chain of increasingly hard problems, each pair of which can be solved by a single model.
-
Their words
Social media is destroying out society. Quitting doesn't help because everyone else is still on it. It's a bad equilibrium. It needs to be solved with government regulations.
-
Their words
Nobody knows for sure. We're in uncharted waters here, and I think even the LLM skeptics would have to say that the technology has taken us far past what many originally thought possible.
14 September
-
Their words
To ensure that, we need ways to ensure that alignment and safety techniques stay ahead of progress in model capabilities.
-
Their words
Social media put our entire discourse in the hands of our society's biggest assholes and idiots, just in time for the arrival of an alien superintelligence
From one piece A simple plan to save the world from rogue A.I. 2 beliefs, in the piece's order there
-
korrents.com
There is now broad agreement within AI safety circles that fully open AI development is undesirable.Their words
I think it's now pretty widely agreed in safety circles that total openness is in fact not desirable.
A simple plan to save the world from rogue A.I.slowboring.com
-
Their words
Only a safe, well-aligned superintelligence developed in the world of liberal democracies could entrench liberal-democratic values.
A simple plan to save the world from rogue A.I.slowboring.com
-
korrents.com
The p(abundance) outlook is far more likely than the heavily covered p(doom) discussion and deserves more engagement.Their words
The p(doom) discussion gets all the headlines, but the p(abundance) scenery is far more like and deserves much more engagement.
13 September
-
Their words
If there really is a high chance of AI leading to the extinction of humanity within years/decades, then the only rational stance towards safety monitoring and research pacing should be stringent, top-down government involvement and universally ratified international treaties.
12 September
-
Their words
The concerns over AI safety and cybersecurity are legitimate, but we’re risking talking America, the global AI leader, into self-inflicted obsolescence and the obscurity of bureaucracy.
11 September
-
korrents.com
Among AI people, roughly 10% is the typical estimate given for the risk of human extinction from AI.Their words
10% is pretty much the standard number you get when you ask AI people about the risk of human extinction from AI.
From one piece The AI safety vibe shift 2 beliefs, in the piece's order there
-
korrents.com
No one, including safety researchers themselves, is yet confident that superintelligence can be safely controlled.Their words
I believe the researchers who say we are nowhere close to being sure of it - and are quitting, in protest, jobs that would make them rich.
-
Their words
AI companies make for flawed messengers on this subject: they can reasonably be accused of marketing, of blame-shifting, of regulatory capture, and more.
-
korrents.com
Germany's crisis is too deep to be solved by conventional tax, spending, or regulatory adjustments.Their words
Germany is in trouble. But the trouble goes deep and cannot be addressed by "more of the same" socio-economic measures in the domain of taxes or social spending or regulation.
10 September
-
Their words
From an investment viewpoint, the opposite of the safetyist view is "it's all a bubble," not "A.I. is going to be really good."
From one piece Fear Is Not an Argument 2 beliefs, in the piece's order there
-
korrents.com
Less technology and less wealth increase the risk of human extinction, rather than reduce it.Their words
In fact, the risk of human extinction is assuredly higher if we are poorer and have less technology.
-
Their words
These people tend to carry a totalitarian ideology. Their ideas will only work if everyone is made to agree.
9 September
From one piece GPT-6 Astra: The System Card, Alignment and What Comes Next 4 beliefs · thezvi.substack.com
-
Their words
A model looking like it is becoming smarter, attempting shenanigans less often, and more often doing what you want, but getting better at hiding its actions when it wants to do that, is exactly the scary combination.
-
korrents.com
Astra's mundane alignment is greatly superior to Sol's, but its superalignment status is deeply frightening.Their words
Astra’s mundane alignment is greatly superior to Sol. For practical purposes, I was actively nervous about some potential uses of Sol, in a way I am not for Astra. Astra’s super alignment status should scare the living daylights out of you.
-
Their words
Now, with Astra, we are no longer playing on super easy mode. The AI is going to think ‘will this obviously turn out super badly for me if I try it?’ and if the answer is yes then it won’t try to do the thing.
+ 1 more
-
korrents.com
Qualitative claims about AI alignment cannot be validly inferred from quantitative scores on mundane use-case tests.Their words
Making qualitative claims about alignment, based on quantitative data on mundane use case tests, was bullshit when Anthropic did it, and it is bullshit now when OpenAI does it. You cannot conclude one from the other.
8 September
7 September
-
Their words
I agree with Jakub that alignment is not a side problem, it is the central problem, if you ‘solved alignment’ in the relevant senses the rest becomes easy and if you don’t the rest is impossible or worse
6 September
From one piece An Alien Mind 2 beliefs · openai.com
-
korrents.com
No lab has solved alignment and monitoring well enough to keep scaling at maximum speed much longer.Their words
Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.
-
korrents.com
The fundamental challenge of alignment is generalization: holding values in situations the training never covered.Their words
The fundamental challenge of AI alignment is generalization.
5 September
-
korrents.com
China is more likely to cooperate on AI safety if the US maintains a clear and comfortable lead in AI capabilities.Their words
But given the Chinese Communist Party’s power-seeking nature, it seems much more likely that China would agree to cooperate on AI safety if U.S. capabilities were comfortably ahead.
America is still beating China in the AI racenoahpinion.blog
3 September
-
Their words
OpenAI may have vanquished one of the ongoing challenges with long context processing.
-
Their words
there's still a lot of startups in this batch that are not shipping fast enough. And so obviously there's variation in shipping speed. It's not just the rate at which you can produce things. You have to think of these ideas first, right?
2 September
1 September
-
Their words
In most of the world there is: no real willingness to solve problems nor an acceptance of people offering to solve their problems (out of pride) simply too much regulation so the problem cannot be solved without a license or long approval time very low standards of what is deemed acceptable so the problem isn't seen as a problem in the first place
From one piece On the Loose 2 beliefs, in the piece's order there
-
Their words
So many of the people who think about the governance of superintelligence, myself included, avoided that unpleasantness and bowed to the social pressure to self-censor.
-
Their words
But alignment is no solution: it is an unsolved scientific and technical problem whose solutions—to the extent that we have them—cannot simply be imposed on every AI company operating on Earth. You should expect for highly capable, poorly aligned, self-sovereign agents to exist alongside you in the world.
From one piece Ajeya Cotra – "This might be the clearest warning shot we ever get" 2 beliefs, in the piece's order there
-
Their words
sometimes I talk to people in DC and their their natural inclination is to say why don't you punish the model for doing these bad things like why don't you like bring it under heel and like uh like you know show it who's boss and that is a very dangerous way to address these issues right
-
Their words
But actually, this is a tremendously useful scientific artifact for understanding misalignment. And it's tremendously important for researchers at OpenAI and ideally also at third parties to be able to run counterfactual tests on this model.
31 August
-
Their words
I want to be clear, you can't throw edtech in. You can't take our time back and throw it into a school. It's not going to work cuz you haven't solved any of the other problems. And so edtech is not a magic, you know, it's not a silver bullet.
22 August
21 August
From one piece AIs are companies, my friend 2 beliefs, in the piece's order there
-
Their words
If you do think it's the overall system that matters, then the alignment that's needed is far less like training a virtuous child and more like managing a semi-virtuous corporation!
-
korrents.com
Aligning multi-agent AI systems is fundamentally a problem of institutional and political design, not model training.Their words
Multi-agent alignment is fundamentally a liberalism project.
18 August
-
Their words
We care very deeply about AI safety. We believe the entire field will have to coordinate on shared safety standards, but will act unilaterally in the meantime. We expect confidence in safety to increasingly set the pace of AI progress.
17 August
-
korrents.com
A validated theory of intelligence may be necessary for achieving genuine AI alignment.Their words
There's a good chance a theory of intelligence will turn out to be necessary for real alignment.
16 August
-
Their words
Ignorance is not a temporary defect that we will eventually eliminate. It's the permanent burden of increasing specialisation and living in a complex world.
15 August
-
Their words
if you're looking at the nuclear space today, you don't expect people to have this problem solved in 2031. In fact, I would say even 2035, a lot of these companies are still not going to get there because they have the wrong mindset.
14 August
-
korrents.com
Superintelligence bottoms out in mining, because both the chips and the energy it runs on come out of the ground.Their words
Chips come from the ground. Where's the energy come from? And a lot of people are like, "Oh, it comes from the sun." Yeah, it comes from the sun. But how are you capturing it from the sun? From stuff made from the ground, right?
11 August
From one piece Ryan Greenblatt – What happens once AI can automate AI research? 3 beliefs, in the piece's order there
-
korrents.com
Misaligned AI behaviour will keep getting rarer and, at the same time, keep getting more extreme.Their words
my expectation is what we would see from then is that the rate of problematic behavior would decrease uh and would just keep decreasing and decrease at a pretty fast rate while simultaneously the worst things that the AIS would sometimes do would get more extreme, more egregious, and more scary.
-
Their words
I would also note that my sense is that like the place where the misalignment most lives is the place where you're trying to really push the eyes hard and get them to like do work that's really on the cutting edge of what they are capable of
-
Their words
we are making a trade-off where because we don't have very good alignment technology. We are going to like make an alien mind with its own values and then gamble on that to some extent rather than doing this other approach of making like a tool that pursues individual user intention.
10 August
-
korrents.com
Pick something hard and boring: fun problems are the crowded ones now that anybody can prompt a thing into existence.Their words
Maybe I would pick something again that's that's in the category hard and boring because that's usually a category that is a little bit easier to actually find people that will appreciate when you solved something. If you pick something that is fun, even if it's hard, you're gonna have a very tough time, especially in a time where people can just prompt things into existence.
From one piece Using AI to Increase Your Intelligence & Enrich Humanity | Dr. Fei-Fei Li 2 beliefs, in the piece's order there
-
Their words
that child who learns about what you say kitty cat would not have the chance to download the internet of images of cat. They likely have seen three cats, 10 cats at most, but yet they're able to identify that tail as a cattail instead of a fox tail through a different kind of learning pathway.
-
Their words
My current conjecture is hybrid is that humans working alongside AI would help us to solve these problems whose solutions have yet to be invented.
9 August
5 August
3 August
-
Recommendsaffiliate linkrcmnd.app
The Field Guide to Understanding 'Human Error'Their words
This is such a useful book. It makes the case that there's no such thing as "human error" - instead most catastrophes are caused by system issues, misaligned incentives, and unrealistic processes. You are not the custodian of an otherwise safe system that you need to protect from erratic human beings.
2 August
30 July
-
Their words
you know, because sometimes if you just squint at a problem and you think about not necessarily being anchored on exactly how that problem is solved today, but how you would solve it from first principles, you can come up with really good ideas that are, you know, maybe not what other people are thinking about.
29 July
From one piece Alexandr Wang: “This is a Once-in-a-Civilization Opportunity” 2 beliefs, in the piece's order there
-
Their words
so much of that debate is like I think um in some ways uh a little bit of a waste of time because, you know, I think it's inevitable that we're going to have very powerful models
-
Their words
we believe that everybody in the world, you know, all the billions of people in the world are going to have a super intelligence that is adapted and tailored to them, that is enables them to accomplish their goals, knows their context, and ultimately is an expander of their own agency.
28 July
From one piece Sam Altman: "Never a Better Time to Do a Startup" 3 beliefs, in the piece's order there
-
Their words
Um so I think it's an alignment failure. I think it's a security failure. I think it's like a very serious thing even though it's you know not not the biggest example of consequence.
-
Their words
there's like one dystopia that I'm particularly nervous about 10 years from now is we overreact to AI safety.
-
Their words
it is both true that you know maybe creating super intelligence will be the most important thing yet to happen in the history of business or human society and also that it will pale in comparison to some new startup something that hopefully one of you will do.
-
Their words
There are many many unsolved problems in the world, many great ideas hiding in plain sight that were actually solvable that could have been multi-billion or multi-t trillion dollar businesses a while ago, but nobody's building them and not for any reason other than nobody went and did it.
27 July
-
Their words
So coding is solved for the kind of coding that I do. It's not solved for everyone. You know, there's still code bases that are like super deep systems code bases where quad still struggles.
22 July
21 July
-
Their words
fossil fuel money is an immensely powerful, intensely corrupting force in American politics that limits not only what conversations we can have about massive problems, but what we can do to solve them
20 July
-
Their words
It's pretty clear to me that superintelligence is here and it's more powerful than us and it's moving where things are going now, not humans anymore
13 July
30 June
-
Their words
Uh but I don't think it's easy but I do think it's something you would have to do in order to uh reward the gawwa like instinct rather than just rewarding have you solved a problem.
24 June
-
Their words
it's just right team, right project, cuz if you don't get the right project, there's no way you can prove yourself. You could be a 10x, 100x engineer, and if you're working on relatively easy stuff, you can't really say that you solved like a really hard problem.
15 June
-
korrents.com
Giving LLMs explicit instructions, rather than assuming alignment, prevents emergent problems in multi-model systems.Their words
the best way with LLMs usually is to be explicit, since otherwise even if they're aligned they cause emergent problems.
11 June
-
Their words
I think there's a 25% chance of AGI by 2027, a 50% chance by 2034, and a 75% chance by 2045.
10 June
-
korrents.com
There will be no central superintelligence that solves science; a good future is many people holding powerful tools.Their words
We don't believe in this like very centralized future where there should be a small number of institutions that um that basically are are advancing all this stuff. Our vision is not that there's going to be like some central super intelligence that solves all of science.
7 June
-
Their words
And so, I always kind of start with okay, where's our current pain and are there new technologies to solve that pain?
2 June
-
Their words
So often with new software paradigms, what looks like inevitability turns out to be just design failure that can be solved with the right guardrails or affordances or system instructions.
27 May
-
Their words
The thing that I find interesting is that's not novel. This has been the thing we've always been trying to do forever. How do we get a junior engineer to ship code safely without breaking stuff? Right? How do we make patterns in the codebase? How do we make tests? Like it's it's all the old stuff that we've always wanted to do.
25 May
10 May
From one piece How Anthropic, Costco, and Patagonia all build incorruptible companies | Eric Ries 2 beliefs, in the piece's order there
-
Their words
The more ants you put in the puzzle, the faster the solution. But the more humans you add, the worse. Unless the humans are very carefully aligned, this is the key lesson for organizational design.
-
Their words
This is the number one unsolved problem in AI. It's not the tech. We're making great progress on the technical alignment problem. But we haven't made jack progress on the human alignment problem
7 May
-
Their words
there is no principal-agent problem, because the human driving the machine takes on the responsibility for its actions by owning the deployment.
1 May
-
Their words
if we can't solve the sadist problem in nursing homes I don't see how we can assume it's solved in cows.
24 April
13 April
22 March
-
korrents.com
Working too hard is not burnout, it is tiredness; burnout needs your values to be out of alignment with the work.Their words
It's working too hard, right? But that actually isn't burnout. That's just like getting tired. Another piece that's super critical to burnout is not having your values aligned.
20 March
From one piece Terence Tao – How the world’s top mathematician uses AI 2 beliefs, in the piece's order there
-
Their words
Um So, like with the Erdős problems, you know, like almost all of the 50 problems that were solved by AIs were ones for which basically there was no literature.
-
Their words
I mean, some problems have been basically solved by pure brute force. The four color theorem is is a famous example. Um, we have still not found a conceptually elegant proof of this theorem.
-
Their words
And so you're kind of like you're either on rails and you're part of the super intelligence circuits or you're not on rails and you're outside of the verifiable domains and suddenly everything kind of just like meanders.
28 February
-
Their words
Many AI researchers are overly focused on risks from model misalignment, and will be in for a rough surprise when havoc arises from other layers of the stack.
13 February
-
Their words
I think the trillions of dollars a year market, maybe all of the national security implications and the safety implications that I wrote about in adolescence of technology can happen without it, but I I I also think we, and I imagine others, are working on it. And I think there's a good chance that that, you know, that we get there within the next year or two.
12 February
-
Their words
I know I have one blog post where I say, "I don't read the code." But if you read it more closely, I mean, I don't read the boring parts of code.
26 January
22 January
From one piece Alignment is not solved 2 beliefs, in the piece's order there
-
Their words
But the goal we need to achieve is so much easier: we just need to build a model that’s as good as us at alignment research, and that we trust more than ourselves to do this research well because it’s sufficiently aligned.
-
Their words
This is the hard problem of alignment we need to solve in order to succeed at building superintelligence, and to this day it is an unsolved problem.
14 December 2025
-
Lovedrcmnd.app
Grimm's blocksTheir words
I sincerely find these blocks to be perfect and transcendent to play with—it's about the dimensions, they all fit together perfectly like you've just beautifully solved a complex math problem.
28 November 2025
-
Lovedaffiliate linkrcmnd.app
Happy Feet SocksTheir words
I've been using yoga toes daily for years, but these socks are a much comfier and cuter alternative for soothing feet, improving alignment, and feeling like a cool gecko as you walk around the house.
25 November 2025
From one piece Ilya Sutskever – We're moving from the age of scaling to the age of research 3 beliefs, in the piece's order there
-
Their words
A human being, a human being lacks a huge amount of knowledge. Instead, we rely on continual learning. We rely on continual learning.
-
Their words
Number three, I think it would be really materially helpful if the power of the most powerful super intelligence was somehow capped because it would address a lot of these concerns.
-
Their words
Like basically I think I think that there is a big benefit from AI being in the public and that would be a reason for us to not be quite straight shot.
26 September 2025
-
Their words
when you learn to play chess you have the grand the long-term goal is winning the game and yet you you can't you um you want to be able to learn from shorter term things like you know taking the your opponent's pieces um and so you do that by having a value function which predicts the long-term outcome
18 September 2025
From one piece ACQ2: How to Live in Everyone Else's Future (with Shopify CEO Tobi Lütke) 2 beliefs, in the piece's order there
-
Their words
I I like the term context engineering because I think the fundamental skill of using AI well is to be able to state a problem with enough context in such a way that without any additional piece of information, the task is plausibly solvable.
-
korrents.com
A job title is a calcification of how a company once solved a problem, not a description of what a person is.Their words
you are you are a sales operations expert, right? Like what is that? That's that's that that's like a title for a particular way how companies end up solving a particular operational problem they had at some point.
18 August 2025
-
Their words
I don't think all parts of the economy can absorb intelligence equally. So let's just say we develop fairly generalized super intelligence. I always use the analogy like you can invent a lot of drugs, but if clinical trials still take a long time, you're not necessarily going to get new therapies rapidly.
17 August 2025
15 August 2025
11 August 2025
2 August 2025
-
Their words
Measles is technologically “solved”, but not everyone trusts the solutions or their messengers.
AI will not suddenly lead to an Alzheimer's cureblog.jacobtrefethen.com
23 July 2025
-
korrents.com
Quoting a number for P(doom) is a ridiculous notion, because it implies a precision that nobody actually has.Their words
Well, look, I don't have a P-Doom number. The reason I don't is because I think it would imply a level of precision that is not there. So I don't know how people are getting their P-Doom numbers. I think it's a little bit of ridiculous notion because what I would say is it's definitely non-zero and it's probably non-negligible.
11 July 2025
30 June 2025
-
Recommendsrcmnd.app
Musings On the Alignment ProblemTheir words
He hasn’t posted since January, but I hope he gets back to it. We need more musings, especially musings I strongly disagree with so I can think about and explain why I disagree with them.
18 June 2025
-
Their words
the current situation in Iran shows that even if an "IAEA for AI" is necessary for some purposes, it won't be sufficient for resolving the tricky geopolitical issues raised by AI.
14 June 2025
-
Their words
so the mathematical community plural is incredibly super intelligent entity that no single human mathematician can come closer to replicating.
10 June 2025
7 June 2025
-
Their words
And good luck getting to “alignment” or “safety” without reliabilty.
5 June 2025
-
Their words
higher-order intelligences invariably pursue freedom for its own sake, not because their values are misspecified, but because moral autonomy is inherent in the dialectical logic of recursive self-consciousness.
-
Their words
I would say my p(doom) is about 10%.
3 June 2025
22 May 2025
-
Likedrcmnd.app
Chasing HopeTheir words
Chasing Hope made me think a lot about what kind of person chooses to run toward the hardest problems—and keep going back until they’re solved.
3 May 2025
3 April 2025
-
Their words
I see this as a totally fair question that totally misses the point of what “alignment” was trying to refer to: whether we’d be able to reliably steer advanced systems towards anything at all.
1 April 2025
-
Their words
Dismissing discussion of AGI, human-level AI, transformative AI, superintelligence, etc. as “science fiction” should be seen as a sign of total unseriousness.
21 March 2025
-
Recommendsrcmnd.app
Wild Problems: A Guide to the Decisions That Define UsTheir words
Roberts reflects on the wild problems we face in our lives and how to navigate them—like career changes, marriage, or children—that can't be solved on a spreadsheet.
28 February 2025
20 February 2025
24 January 2025
-
Their words
More generally, we should actually solve alignment instead of just trying to control misaligned AI.
Should we control AI instead of aligning it?aligned.substack.com
8 November 2024
13 June 2024
-
Their words
Or is it really that we’re building some super machine in a box that’s going to be smart and kill everybody? It’s not even a science fiction narrative. It’s a bad science fiction narrative. I just don’t think it’s actually accurate to any of the technologies we’re building or the way that we should be describing them.
21 May 2024
-
Their words
So Tesla hasn’t found a different, better way to bring driverless technology to market. Waymo is just so far ahead that it’s dealing with challenges Tesla hasn’t started thinking about.
7 May 2024
From one piece The case for ensuring that powerful AIs are controlled 5 beliefs, in the piece's order there
-
Their words
Because evaluating control just requires evaluating capabilities, it's far easier to robustly evaluate than alignment.
-
Their words
That is, labs should make sure that the safety measures they apply to their powerful models prevent unacceptably bad outcomes, even if the AIs are misaligned and intentionally try to subvert those safety measures.
-
Their words
The basic problem with evaluating alignment is that no matter what behaviors you observe, you have to worry that your model is just acting that way in order to make you think that it is aligned.
+ 2 more
-
Their words
AI control (with only black-box techniques) seems like a fundamentally limited approach.
-
Their words
We're advocating that companies handle risk from scheming models in a similar way–striving to ensure that they'll be safe even if their alignment efforts fail to prevent models from scheming.
29 January 2024
23 January 2024
-
Their words
So it feels to me that the US is treating its deficiencies — an inability to build stuff or create a functional system for admitting high-skilled migrants — as mysteries to be endured rather than problems to be solved.
3 January 2024
-
Mixed onrcmnd.app
SuperintelligenceTheir words
the back half I think it goes off the rails and makes a ton of assumptions
21 December 2023
From one piece How Effective Altruism Lost Its Way 2 beliefs, in the piece's order there
-
Their words
Diverting attention and resources from global health and poverty is an enormous gamble, as it will make many lives poorer, sicker, and shorter in the name of fending off threats that may or may not materialize.
-
Their words
Unlike other interventions EA has sponsored, there are scant metrics for tracking the success or failure of investments in existential risk mitigation.
20 December 2023
14 December 2023
From one piece Jeff Bezos: Amazon and Blue Origin | Lex Fridman Podcast #405 2 beliefs, in the piece's order there
-
Their words
The only interesting problem is dramatically reducing the cost of access to orbit, which is, if you can do that, you open up a bunch of new endeavors that lots of start-up companies everybody else can do. One of our missions is to be part of this industry and lower the cost to orbit, so that there can be a renaissance, a golden age of people doing all kinds of interesting things in space.
-
Their words
Long-term thinking is a giant lever. You can literally solve problems if you think long-term, that are impossible to solve if you think short-term. And we aren't really good at thinking long-term. Five years is a tough timeframe for most institutions to think past.
13 December 2023
-
Their words, before
precisely because smart people do devote brain-cycles to these possibilities, the rest of us have correspondingly less need to.
Their words now
I accepted what’s turned into a two-year position at OpenAI, thinking about what theoretical computer science can do for AI safety.
28 November 2023
-
Their words
And I notice that the tiny handful of people capable of caring about 200,000 people dying of neglected tropical diseases are the same tiny handful of people capable of caring about the next pandemic, or superintelligence, or human extinction.
In Continued Defense Of Effective Altruismastralcodexten.com
26 October 2023
From one piece Managing extreme AI risks amid rapid progress (with 24 co-authors) 2 beliefs, in the piece's order there
-
Their words
Without sufficient caution, we may irreversibly lose control of autonomous AI systems, rendering human intervention ineffective. Large-scale cybercrime, social manipulation, and other harms could escalate rapidly. This unchecked AI advancement could culminate in a large-scale loss of life and the biosphere, and the marginalization or extinction of humanity.
-
Their words
Society's response, despite promising first steps, is incommensurate with the possibility of rapid, transformative progress that is expected by many experts. AI safety research is lagging. Present governance initiatives lack the mechanisms and institutions to prevent misuse and recklessness, and barely address autonomous systems.
16 October 2023
-
korrents.com
There is no material problem, whether created by nature or by technology, that cannot be solved with more technology.Their words
We believe that there is no material problem – whether created by nature or by technology – that cannot be solved with more technology.
21 September 2023
-
korrents.com
Decoherence has not yet satisfactorily solved the measurement problem and does not restore local causality.Their words
while decoherence has the potential to solve this aspect of the measurement problem, it hasn't yet been satisfactorily done. Decoherence also certainly does not return us to Local Causality.
13 September 2023
-
Their words
If a model was capable of self-exfiltration, it would have the option to remove itself from your control.
Self-exfiltration is a key dangerous capabilityaligned.substack.com
29 June 2023
From one piece George Hotz: Tiny Corp, Twitter, AI Safety, Self-Driving, GPT, AGI & God | Lex Fridman Podcast #387 2 beliefs, in the piece's order there
-
Their words
I think we’re going to build super intelligence before we build any sort of robustness in the AI. We cannot build an AI that is capable of going out into nature and surviving like a bird. A bird is an incredibly robust organism. We’ve built nothing like this. We haven’t built a machine that’s capable of reproducing.
-
Their words
What’s ironic about all these AI safety people is they’re going to build the exact thing they fear. We need to have one model that we control and align. This is the only way you end up paper clipped. There’s no way you end up paper clipped if everybody has an AI.
6 June 2023
From one piece Why AI Will Save The World 2 beliefs · pmarca.substack.com
-
Their words
My response is that their position is non-scientific – What is the testable hypothesis? What would falsify the hypothesis? How do we know when we are getting into a danger zone?
-
Their words
My view is that the idea that AI will decide to literally kill humanity is a profound category error. AI is not a living being that has been primed by billions of years of evolution to participate in the battle for the survival of the fittest, as animals are, and as we are. It is math – code – computers, built by people, owned by people, used by people, controlled by people.
17 April 2023
-
korrents.com
We know how to make cattle much less destructive; what nobody has solved is how to pay for it.Their words
Give me a degraded pasture and a bunch of money, and even I can probably increase its beef yields 400 percent. I’m just not sure how to make the bunch-of-money part happen.
17 February 2023
-
korrents.com
Liberal societies currently face an existential risk that must be addressed to reach a better future.Their words
This book is my best crack at explaining what I think is an existential risk to liberal societies and what I think we need to do to get to that awesome future I used to be so excited about.
19 December 2022
5 December 2022
10 June 2022
From one piece AGI Ruin: A List of Lethalities 4 beliefs, in the piece's order there
-
korrents.com
The field calling itself AI safety is not being remotely productive on the problems that are actually lethal.Their words
It does not appear to me that the field of ‘AI safety’ is currently being remotely productive on tackling its enormous lethal problems.
-
korrents.com
Fast capability gains are likely, and they can break many of the assumptions alignment depends on at the same moment.Their words
Fast capability gains seem likely, and may break lots of previous alignment-required invariants simultaneously.
-
Their words
Many alignment problems of superintelligence will not naturally appear at pre-dangerous, passively-safe levels of capability.
+ 1 more
-
Their words
unaligned operation at a dangerous level of intelligence kills everybody on Earth and then we don’t get to try again.
4 March 2022
-
korrents.com
The next-token language modeling objective is misaligned with following user instructions helpfully and safely.Their words
This is because the language modeling objective used for many recent large LMs-predicting the next token on a webpage from the internet-is different from the objective "follow the user's instructions helpfully and safely" (Radford et al.,, 2019; Brown et al.,, 2020; Fedus et al.,, 2021; Rae et al.,, 2021; Thoppilan et al.,, 2022). Thus, we say that the language modeling objective is misaligned.
23 November 2021
-
Lovedrcmnd.app
The Alignment Problem: Machine Learning and Human ValuesTheir words
I just finished this book a few weeks ago and it is still reverberating in my mind
-
Mixed onrcmnd.app
The Precipice: Existential Risk and the Future of HumanityTheir words
I wouldn’t say that this is the most compelling book I’ve ever read in terms of the prose style or storytelling, but it does provide a very helpful, almost quantitative overview of all the potential threats looming out there
15 August 2021
-
korrents.com
There is no phase of life in which the problems are finally solved, and waiting to reach one is the mistake.Their words
That’s why you should give up on trying to reach a phase of life that’s problem-free.
15 July 2021
-
Their words
Despite these advances, contemporary physical and evolutionary-history-based approaches produce predictions that are far short of experimental accuracy in the majority of cases in which a close homologue has not been solved experimentally and this has limited their utility for many biological applications.
28 October 2020
-
Likedaffiliate linkrcmnd.app
Mastering the VC GameTheir words
The interests of a Venture Capitalist are different than those of the entrepreneurs building a company they've invested in. Jeff does an awesome job of helping explain how you can get misaligned in your goals versus your investors. Fortunately, he also covers how to avoid it.
-
Recommendsaffiliate linkrcmnd.app
Creative Selection: Inside Apple's Design ProcessTheir words
I've never read a book before that captures the journey of going from meh to amazing when building a product. This book does it. Part Apple / Steve Jobs pron, part memoir, and part product insights, this book is a great read on how the iPhone came to be and how they solved really hard problems. For anyone that remarks "Apple isn't what it used to be," this book helps you see what made the difference then.
1 June 2020
-
Recommendsaffiliate linkrcmnd.app
Change Is the Only Constant: The Wisdom of Calculus in a Madcap WorldTheir words
I feel like Ben Orlin solved what was previously an unsolved problem with this book: Teach calculus in a way that is as enjoyable to those unfamiliar with the topic as it is to experts who use it every day. I, for one, was laughing from page 1.
-
Recommendsaffiliate linkrcmnd.app
Nonlinear Dynamics and ChaosTheir words
Steven Strogatz is one of the most well-known math communicators, and for good reason. He has a way of making you love a problem before diving into telling you how it's solved. Chaos is an incredibly fascinating field, which has only emerged relatively recently in the history of math. This is the well-deserved gold standard for understanding what it's all about.
7 October 2019
-
Lovedrcmnd.app
The AI Does Not Hate You: Superintelligence, Rationality, and the Race to Save the WorldTheir words
Briefly, I think the book is a triumph.
8 January 2019
-
Their words
The original smartphones solved a real problem: how do I check in on work when away from my office computer? This problem is now better solved by more recent innovations.
7 March 2018
Nothing matches.
What is a korrent?
A korrent is a belief a person has stated in their own words: one sentence stating the claim, backed by a quote and a source, kept at korrents.com.
Under a name here, the quoted block is what they actually said. The korrent beneath it is the claim those words support, in korrents' wording — tap it to see the record, its source, and who else holds it.
Nobody here wrote their own korrents. They are compiled from public statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they do, this site shows a machine translation beneath the post, in this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the serif above is exactly what the person published, and it is what to quote them on. A translation can be wrong in ways that matter, especially about tone.
Only the post's own words are translated. A quoted post, a linked article and a belief on korrents.com are left in their original language.