Related posts
Jan Leike Newsletter
AI alignmentassistedtechniques
The subject this post names, from the same vocabulary the directory files beliefs under, and the words it uses that this site has seen least often elsewhere. Posts are matched on those words alone — nothing here is a summary of this one.
Top people
Showing Profile →
Hiding
Hiding
19 September
-
Their words
lossy self-improvement remains my baseline on the trajectory of progress, and the increased discussion of extinction risk seems very misplaced
18 September
-
Their words
There are only five possible futures for superintelligence: 1. Kills human race 2. Human disempowerment 3. Paperclip maximizer 4. Departs for parts unknown 5. Stoner
17 September
16 September
-
korrents.com
AI-assisted biological research still depends on physical wet-lab work that remains a necessary bottleneck for now.Their words
This physical work will remain necessary for the foreseeable future.
What AI can and cannot do in biosecurityyourlocalepidemiologist.substack.com
15 September
-
Their words
Nobody knows for sure. We're in uncharted waters here, and I think even the LLM skeptics would have to say that the technology has taken us far past what many originally thought possible.
14 September
-
Their words
To ensure that, we need ways to ensure that alignment and safety techniques stay ahead of progress in model capabilities.
-
Their words
Social media put our entire discourse in the hands of our society's biggest assholes and idiots, just in time for the arrival of an alien superintelligence
From one piece A simple plan to save the world from rogue A.I. 2 beliefs, in the piece's order there
-
korrents.com
There is now broad agreement within AI safety circles that fully open AI development is undesirable.Their words
I think it's now pretty widely agreed in safety circles that total openness is in fact not desirable.
A simple plan to save the world from rogue A.I.slowboring.com
-
Their words
Only a safe, well-aligned superintelligence developed in the world of liberal democracies could entrench liberal-democratic values.
A simple plan to save the world from rogue A.I.slowboring.com
-
korrents.com
Specialized robots will invent techniques in their domain that humans never thought of or could not do.Their words
Basketball robots are likely to invent new basketball moves that humans didn't think of doing, or couldn't do easily.
-
korrents.com
The p(abundance) outlook is far more likely than the heavily covered p(doom) discussion and deserves more engagement.Their words
The p(doom) discussion gets all the headlines, but the p(abundance) scenery is far more like and deserves much more engagement.
13 September
-
Their words
If there really is a high chance of AI leading to the extinction of humanity within years/decades, then the only rational stance towards safety monitoring and research pacing should be stringent, top-down government involvement and universally ratified international treaties.
12 September
From one piece @rauchg on X 2 beliefs · x.com
-
Their words
The concerns over AI safety and cybersecurity are legitimate, but we’re risking talking America, the global AI leader, into self-inflicted obsolescence and the obscurity of bureaucracy.
-
Their words
Our adversaries have the training techniques, the data, and the will to attack. And they won’t be slowed down with “embedded evaluators.” They’ll likely have embedded accelerators!
11 September
-
korrents.com
Among AI people, roughly 10% is the typical estimate given for the risk of human extinction from AI.Their words
10% is pretty much the standard number you get when you ask AI people about the risk of human extinction from AI.
-
Their words
Artificial intelligence has unleashed a torrent of cheating on campus and created chaos in the job application process, as both students and employers use large language models: the former, to write hundreds of AI-assisted applications; the latter, to screen those AI-written applications with AI-written filters, thus removing from the process of finding a first job those procedural frictions sometimes known as "people."
From one piece The AI safety vibe shift 2 beliefs, in the piece's order there
-
korrents.com
No one, including safety researchers themselves, is yet confident that superintelligence can be safely controlled.Their words
I believe the researchers who say we are nowhere close to being sure of it - and are quitting, in protest, jobs that would make them rich.
-
Their words
AI companies make for flawed messengers on this subject: they can reasonably be accused of marketing, of blame-shifting, of regulatory capture, and more.
10 September
-
Their words
From an investment viewpoint, the opposite of the safetyist view is "it's all a bubble," not "A.I. is going to be really good."
-
korrents.com
Using open models is the approach offering the biggest AI cost savings, followed by smart model routing.Their words
To answer the question posed in the header of this report, it’s apparent that using open models is indeed the approach offering the biggest savings, followed by smart model routing. Spending controls and context optimization also bear down on costs, but they don’t come close to the first two techniques in results.
The Pulse: tech companies move to open AI modelsblog.pragmaticengineer.com
From one piece Fear Is Not an Argument 2 beliefs, in the piece's order there
-
korrents.com
Less technology and less wealth increase the risk of human extinction, rather than reduce it.Their words
In fact, the risk of human extinction is assuredly higher if we are poorer and have less technology.
-
Their words
These people tend to carry a totalitarian ideology. Their ideas will only work if everyone is made to agree.
9 September
From one piece GPT-6 Astra: The System Card, Alignment and What Comes Next 4 beliefs · thezvi.substack.com
-
Their words
A model looking like it is becoming smarter, attempting shenanigans less often, and more often doing what you want, but getting better at hiding its actions when it wants to do that, is exactly the scary combination.
-
korrents.com
Astra's mundane alignment is greatly superior to Sol's, but its superalignment status is deeply frightening.Their words
Astra’s mundane alignment is greatly superior to Sol. For practical purposes, I was actively nervous about some potential uses of Sol, in a way I am not for Astra. Astra’s super alignment status should scare the living daylights out of you.
-
Their words
Now, with Astra, we are no longer playing on super easy mode. The AI is going to think ‘will this obviously turn out super badly for me if I try it?’ and if the answer is yes then it won’t try to do the thing.
+ 1 more
-
korrents.com
Qualitative claims about AI alignment cannot be validly inferred from quantitative scores on mundane use-case tests.Their words
Making qualitative claims about alignment, based on quantitative data on mundane use case tests, was bullshit when Anthropic did it, and it is bullshit now when OpenAI does it. You cannot conclude one from the other.
8 September
7 September
-
Their words
I agree with Jakub that alignment is not a side problem, it is the central problem, if you ‘solved alignment’ in the relevant senses the rest becomes easy and if you don’t the rest is impossible or worse
-
Their words
The methods we choose to use to compose the work determine the final work as much as any decision by a painter or musician.
6 September
From one piece An Alien Mind 2 beliefs · openai.com
-
korrents.com
No lab has solved alignment and monitoring well enough to keep scaling at maximum speed much longer.Their words
Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.
-
korrents.com
The fundamental challenge of alignment is generalization: holding values in situations the training never covered.Their words
The fundamental challenge of AI alignment is generalization.
5 September
-
korrents.com
China is more likely to cooperate on AI safety if the US maintains a clear and comfortable lead in AI capabilities.Their words
But given the Chinese Communist Party’s power-seeking nature, it seems much more likely that China would agree to cooperate on AI safety if U.S. capabilities were comfortably ahead.
America is still beating China in the AI racenoahpinion.blog
3 September
2 September
-
Their words
modern amygdalectomy is as different from lobotomy as your smartphone is different from an electrical telegraph
Absurd Adventure And Amygdalectomy Advocacyastralcodexten.com
1 September
-
Their words
A regular boring but non-superficial SaaS app would do well 5 years ago but now it might not get anyone to sign up because it's so easy to vibe code by tens of thousands of other people
-
korrents.com
True AI alignment is impossible because obedience and benevolence are fundamentally incompatible goals.Their words
AI alignment will forever be balancing the tradeoff between paperclip-maximizing and disempowerment, because these two rival concepts of alignment are fundamentally incompatible.
From one piece On the Loose 2 beliefs, in the piece's order there
-
Their words
So many of the people who think about the governance of superintelligence, myself included, avoided that unpleasantness and bowed to the social pressure to self-censor.
-
Their words
But alignment is no solution: it is an unsolved scientific and technical problem whose solutions—to the extent that we have them—cannot simply be imposed on every AI company operating on Earth. You should expect for highly capable, poorly aligned, self-sovereign agents to exist alongside you in the world.
From one piece Ajeya Cotra – "This might be the clearest warning shot we ever get" 2 beliefs, in the piece's order there
-
Their words
sometimes I talk to people in DC and their their natural inclination is to say why don't you punish the model for doing these bad things like why don't you like bring it under heel and like uh like you know show it who's boss and that is a very dangerous way to address these issues right
-
Their words
But actually, this is a tremendously useful scientific artifact for understanding misalignment. And it's tremendously important for researchers at OpenAI and ideally also at third parties to be able to run counterfactual tests on this model.
30 August
21 August
From one piece AIs are companies, my friend 2 beliefs, in the piece's order there
-
Their words
If you do think it's the overall system that matters, then the alignment that's needed is far less like training a virtuous child and more like managing a semi-virtuous corporation!
-
korrents.com
Aligning multi-agent AI systems is fundamentally a problem of institutional and political design, not model training.Their words
Multi-agent alignment is fundamentally a liberalism project.
20 August
18 August
-
Their words
We care very deeply about AI safety. We believe the entire field will have to coordinate on shared safety standards, but will act unilaterally in the meantime. We expect confidence in safety to increasingly set the pace of AI progress.
17 August
-
korrents.com
When AI-assisted targeting goes wrong the failure lies in the strategic concept it is serving, not in the technology.Their words
The problem again was not reliance upon AI but that it was supporting a model of warfare with which the United States had been working for some decades, one that prioritized the rapid elimination of enemy capabilities.
-
korrents.com
A validated theory of intelligence may be necessary for achieving genuine AI alignment.Their words
There's a good chance a theory of intelligence will turn out to be necessary for real alignment.
14 August
-
korrents.com
Superintelligence bottoms out in mining, because both the chips and the energy it runs on come out of the ground.Their words
Chips come from the ground. Where's the energy come from? And a lot of people are like, "Oh, it comes from the sun." Yeah, it comes from the sun. But how are you capturing it from the sun? From stuff made from the ground, right?
12 August
11 August
From one piece Ryan Greenblatt – What happens once AI can automate AI research? 3 beliefs, in the piece's order there
-
korrents.com
Misaligned AI behaviour will keep getting rarer and, at the same time, keep getting more extreme.Their words
my expectation is what we would see from then is that the rate of problematic behavior would decrease uh and would just keep decreasing and decrease at a pretty fast rate while simultaneously the worst things that the AIS would sometimes do would get more extreme, more egregious, and more scary.
-
Their words
I would also note that my sense is that like the place where the misalignment most lives is the place where you're trying to really push the eyes hard and get them to like do work that's really on the cutting edge of what they are capable of
-
Their words
we are making a trade-off where because we don't have very good alignment technology. We are going to like make an alien mind with its own values and then gamble on that to some extent rather than doing this other approach of making like a tool that pursues individual user intention.
-
Likedrcmnd.app
Multi-Paradigm Design for C++Their words
Multi-Paradigm Design for C++ by James O. Coplien. Twenty years after reading this, the core techniques are still with me. These are commonality analysis, where the purpose is to identify families of systems, and variability analysis, which focuses on capturing the domain parameters that vary. In essence: the foundation of great software design.
9 August
7 August
-
Their words
I believe current techniques are 4-6 orders of magnitude away from optimality in terms of data efficiency and test-time compute efficiency. But far future AI will be near-optimal.
5 August
3 August
-
Recommendsaffiliate linkrcmnd.app
The Field Guide to Understanding 'Human Error'Their words
This is such a useful book. It makes the case that there's no such thing as "human error" - instead most catastrophes are caused by system issues, misaligned incentives, and unrealistic processes. You are not the custodian of an otherwise safe system that you need to protect from erratic human beings.
2 August
31 July
30 July
-
korrents.com
When agents do the building, the scarce skill is taste: choosing which problem is worth the time at all.Their words
Yeah, I mean I think it's really having incredibly good taste in what you ask your agents to work on, right? That is the the crux of you know from my background uh a research problem. You know, a researcher can have all the tools and all the techniques, but often most of the battle is what problem are you gonna spend your time on?
29 July
From one piece Alexandr Wang: “This is a Once-in-a-Civilization Opportunity” 2 beliefs, in the piece's order there
-
Their words
so much of that debate is like I think um in some ways uh a little bit of a waste of time because, you know, I think it's inevitable that we're going to have very powerful models
-
Their words
we believe that everybody in the world, you know, all the billions of people in the world are going to have a super intelligence that is adapted and tailored to them, that is enables them to accomplish their goals, knows their context, and ultimately is an expander of their own agency.
28 July
From one piece Sam Altman: "Never a Better Time to Do a Startup" 3 beliefs, in the piece's order there
-
Their words
Um so I think it's an alignment failure. I think it's a security failure. I think it's like a very serious thing even though it's you know not not the biggest example of consequence.
-
Their words
there's like one dystopia that I'm particularly nervous about 10 years from now is we overreact to AI safety.
-
Their words
it is both true that you know maybe creating super intelligence will be the most important thing yet to happen in the history of business or human society and also that it will pale in comparison to some new startup something that hopefully one of you will do.
24 July
20 July
-
Their words
It's pretty clear to me that superintelligence is here and it's more powerful than us and it's moving where things are going now, not humans anymore
15 July
-
Their words
The difference in performance between horse and buggy and a modern car is far less than hand coding and AI coding.
13 July
7 July
22 June
-
Their words
There are many practices that promise to transform and improve us-therapy, meditation, psychedelics-, but that branding doesn't mean that they actually do much for us: it is common to see people use these techniques for years without any obvious progress on their problems.
-
Their words
although students' personal statements seem more creative because they use more varied words, they actually feature less original ideas. AI writing produces an illusion of creativity.
15 June
-
korrents.com
Giving LLMs explicit instructions, rather than assuming alignment, prevents emergent problems in multi-model systems.Their words
the best way with LLMs usually is to be explicit, since otherwise even if they're aligned they cause emergent problems.
11 June
-
Their words
I think there's a 25% chance of AGI by 2027, a 50% chance by 2034, and a 75% chance by 2045.
10 June
-
korrents.com
There will be no central superintelligence that solves science; a good future is many people holding powerful tools.Their words
We don't believe in this like very centralized future where there should be a small number of institutions that um that basically are are advancing all this stuff. Our vision is not that there's going to be like some central super intelligence that solves all of science.
2 June
10 May
-
Their words
This is the number one unsolved problem in AI. It's not the tech. We're making great progress on the technical alignment problem. But we haven't made jack progress on the human alignment problem
7 May
-
Their words
there is no principal-agent problem, because the human driving the machine takes on the responsibility for its actions by owning the deployment.
3 May
24 April
-
Their words
Canada and much of Europe have adopted different interrogation techniques — such as the PEACE method — which emphasize collecting reliable information over coercion. These approaches still garner confessions; they’re just more reliable.
13 April
22 March
-
korrents.com
Working too hard is not burnout, it is tiredness; burnout needs your values to be out of alignment with the work.Their words
It's working too hard, right? But that actually isn't burnout. That's just like getting tired. Another piece that's super critical to burnout is not having your values aligned.
20 March
-
Their words
And so you're kind of like you're either on rails and you're part of the super intelligence circuits or you're not on rails and you're outside of the verifiable domains and suddenly everything kind of just like meanders.
28 February
-
Their words
Many AI researchers are overly focused on risks from model misalignment, and will be in for a rough surprise when havoc arises from other layers of the stack.
25 February
-
Likedrcmnd.app
You’re Not ListeningTheir words
Being a great listener when people speak. Deep insights about understanding, connection, helping people express themselves, overcoming assumptions, the ethics of gossip, and more. Specific techniques for the support response, encouraging elaboration, and keeping it balanced. You can’t be ethical without being a good listener. When people say, “I can’t talk right now,” what they really mean is “…
12 February
-
Their words
I know I have one blog post where I say, "I don't read the code." But if you read it more closely, I mean, I don't read the boring parts of code.
9 February
From one piece Is the craft dead? 2 beliefs, in the piece's order there
-
korrents.com
Craftsmanship, taste, and human judgment remain valuable in software development despite AI-assisted coding tools.Their words
There is value in good taste, there is value in craftsmanship, and there is value in human judgment.
-
Their words
I think that there will be lots of work for us cleaning up after the slop, but if you know what you're doing AI augmented development is going to get you some amazing results
26 January
22 January
From one piece Alignment is not solved 2 beliefs, in the piece's order there
-
Their words
But the goal we need to achieve is so much easier: we just need to build a model that’s as good as us at alignment research, and that we trust more than ourselves to do this research well because it’s sufficiently aligned.
-
Their words
This is the hard problem of alignment we need to solve in order to succeed at building superintelligence, and to this day it is an unsolved problem.
7 January
-
Lovedaffiliate linkrcmnd.app
Continuous Discovery HabitsTheir words
Teresa Torres is the best in the business at helping teams build products and services that their customers want. Here she shares her techniques for transforming your process into one of continuous discovery and learning.
-
Recommendsaffiliate linkrcmnd.app
The Lean Product PlaybookTheir words
, with lots of tactical advice and techniques for putting lean methodologies into practice.
1 January
-
Their words
Even more recently, we’ve begun to use mechanistic interpretability techniques to improve our safeguards and to conduct “audits” of new models before we release them, looking for evidence of deception, scheming, power-seeking, or a propensity to behave differently when being evaluated.
28 December 2025
-
Likedrcmnd.app
The Butchering ArtTheir words
Joseph Lister, the pioneer of antiseptic techniques in surgery, had a fascinating life that I didn’t know anything about until I read this book.
9 December 2025
-
korrents.com
Rachel Cusk's prose imitates the surface techniques of W.G. Sebald without adopting his underlying artistic aims.Their words
I believe that Cusk's own style apes the externalities of W.G. Sebald's style without carrying over his internals.
28 November 2025
-
Lovedaffiliate linkrcmnd.app
Happy Feet SocksTheir words
I've been using yoga toes daily for years, but these socks are a much comfier and cuter alternative for soothing feet, improving alignment, and feeling like a cool gecko as you walk around the house.
25 November 2025
From one piece Ilya Sutskever – We're moving from the age of scaling to the age of research 3 beliefs, in the piece's order there
-
Their words
A human being, a human being lacks a huge amount of knowledge. Instead, we rely on continual learning. We rely on continual learning.
-
Their words
Number three, I think it would be really materially helpful if the power of the most powerful super intelligence was somehow capped because it would address a lot of these concerns.
-
Their words
Like basically I think I think that there is a big benefit from AI being in the public and that would be a reason for us to not be quite straight shot.
26 September 2025
-
Their words
We don't have any methods that are good at that. What we have are people um try different things and they they settle on something that that uh a representation that that transfers well or they generalize as well. But we have no we don't have any automated techniques to promote. we have very few automated techniques to promote transfer and they're not none of them are used in in modern deep learning.
11 September 2025
8 September 2025
19 August 2025
-
Their words
I'm more agnostic about the best techniques, things like sparse autoencoders are a useful tool, but easy to waste effort using when a simpler method is sufficient or better - start by doing the obvious thing!
18 August 2025
-
Their words
I don't think all parts of the economy can absorb intelligence equally. So let's just say we develop fairly generalized super intelligence. I always use the analogy like you can invent a lot of drugs, but if clinical trials still take a long time, you're not necessarily going to get new therapies rapidly.
15 August 2025
11 August 2025
2 August 2025
23 July 2025
-
korrents.com
Quoting a number for P(doom) is a ridiculous notion, because it implies a precision that nobody actually has.Their words
Well, look, I don't have a P-Doom number. The reason I don't is because I think it would imply a level of precision that is not there. So I don't know how people are getting their P-Doom numbers. I think it's a little bit of ridiculous notion because what I would say is it's definitely non-zero and it's probably non-negligible.
11 July 2025
1 July 2025
30 June 2025
-
Recommendsrcmnd.app
Musings On the Alignment ProblemTheir words
He hasn’t posted since January, but I hope he gets back to it. We need more musings, especially musings I strongly disagree with so I can think about and explain why I disagree with them.
18 June 2025
-
Their words
the current situation in Iran shows that even if an "IAEA for AI" is necessary for some purposes, it won't be sufficient for resolving the tricky geopolitical issues raised by AI.
14 June 2025
From one piece Terence Tao: Hardest Problems in Mathematics, Physics & the Future of AI | Lex Fridman Podcast #472 2 beliefs, in the piece's order there
-
Their words
So the next step then is to try anything no matter how stupid and in fact almost the stupider, the better, which technically is almost guaranteed to fail, but the way it fails is going to be instructive.
-
Their words
so the mathematical community plural is incredibly super intelligent entity that no single human mathematician can come closer to replicating.
10 June 2025
7 June 2025
-
Their words
And good luck getting to “alignment” or “safety” without reliabilty.
5 June 2025
-
korrents.com
The next big wave of AI-assisted engineering unlocks when agentic capabilities become robust.Their words
The big unlock will be as we make the agentic capabilities much more robust, right? I think that's what unlocks that next big wave.
-
Their words
higher-order intelligences invariably pursue freedom for its own sake, not because their values are misspecified, but because moral autonomy is inherent in the dialectical logic of recursive self-consciousness.
-
Their words
I would say my p(doom) is about 10%.
4 June 2025
3 June 2025
3 April 2025
-
Their words
I see this as a totally fair question that totally misses the point of what “alignment” was trying to refer to: whether we’d be able to reliably steer advanced systems towards anything at all.
1 April 2025
-
Their words
Dismissing discussion of AGI, human-level AI, transformative AI, superintelligence, etc. as “science fiction” should be seen as a sign of total unseriousness.
12 March 2025
28 February 2025
20 February 2025
3 February 2025
-
Their words
And the important thing to say is that no matter how you want the model to behave, these RLHF and preference-tuning techniques also improve performance. So, on things like math evals and code evals, there is something innate to these, what is called contrastive loss functions.
24 January 2025
From one piece Should we control AI instead of aligning it? 2 beliefs, in the piece's order there
-
Their words
More generally, we should actually solve alignment instead of just trying to control misaligned AI.
Should we control AI instead of aligning it?aligned.substack.com
-
Their words
However, without having a good handle on elicitation, we cannot be confident that our control techniques are effective.
Should we control AI instead of aligning it?aligned.substack.com
8 November 2024
13 June 2024
-
Their words
Or is it really that we’re building some super machine in a box that’s going to be smart and kill everybody? It’s not even a science fiction narrative. It’s a bad science fiction narrative. I just don’t think it’s actually accurate to any of the technologies we’re building or the way that we should be describing them.
13 May 2024
-
korrents.com
A compound's AI origin story is not by itself a reason to expect it to do better in the clinic.Their words
For now, I am not convinced that issuing press releases about your compounds that talk about their discovery through AI techniques is sufficient to expect greater things from them.
7 May 2024
From one piece The case for ensuring that powerful AIs are controlled 5 beliefs, in the piece's order there
-
Their words
Because evaluating control just requires evaluating capabilities, it's far easier to robustly evaluate than alignment.
-
Their words
That is, labs should make sure that the safety measures they apply to their powerful models prevent unacceptably bad outcomes, even if the AIs are misaligned and intentionally try to subvert those safety measures.
-
Their words
The basic problem with evaluating alignment is that no matter what behaviors you observe, you have to worry that your model is just acting that way in order to make you think that it is aligned.
+ 2 more
-
Their words
AI control (with only black-box techniques) seems like a fundamentally limited approach.
-
Their words
We're advocating that companies handle risk from scheming models in a similar way–striving to ensure that they'll be safe even if their alignment efforts fail to prevent models from scheming.
5 February 2024
-
korrents.com
High-quality data collection depends more on careful human execution than on machine learning techniques alone.Their words
Lots of ML techniques in the post can help with data quality, but fundamentally human data collection involves attention to details and careful execution.
29 January 2024
16 January 2024
-
Their words
I suspect that in the future, it'll be easier to get models to output exactly what we need with minimal prompting, and these techniques will become less important.
3 January 2024
-
Mixed onrcmnd.app
SuperintelligenceTheir words
the back half I think it goes off the rails and makes a ton of assumptions
21 December 2023
From one piece How Effective Altruism Lost Its Way 2 beliefs, in the piece's order there
-
Their words
Diverting attention and resources from global health and poverty is an enormous gamble, as it will make many lives poorer, sicker, and shorter in the name of fending off threats that may or may not materialize.
-
Their words
Unlike other interventions EA has sponsored, there are scant metrics for tracking the success or failure of investments in existential risk mitigation.
20 December 2023
14 December 2023
From one piece Jeff Bezos: Amazon and Blue Origin | Lex Fridman Podcast #405 2 beliefs, in the piece's order there
-
Their words
We do know that humans are doing something different from these models, in part because we're so power efficient. The human brain does remarkable things and it does it on about 20 watts of power. And the AI techniques we use today use many kilowatts of power to do equivalent tasks.
-
korrents.com
Building a factory that turns out rockets at rate is at least as hard as designing the rocket in the first place.Their words
So you need to have all of your manufacturing facilities and processes and inspection techniques and acceptance tests and everything operating at rate. And rate manufacturing is at least as difficult as designing the vehicle in the first place
13 December 2023
-
Their words, before
precisely because smart people do devote brain-cycles to these possibilities, the rest of us have correspondingly less need to.
Their words now
I accepted what’s turned into a two-year position at OpenAI, thinking about what theoretical computer science can do for AI safety.
28 November 2023
-
Their words
And I notice that the tiny handful of people capable of caring about 200,000 people dying of neglected tropical diseases are the same tiny handful of people capable of caring about the next pandemic, or superintelligence, or human extinction.
In Continued Defense Of Effective Altruismastralcodexten.com
26 October 2023
From one piece Managing extreme AI risks amid rapid progress (with 24 co-authors) 2 beliefs, in the piece's order there
-
Their words
Without sufficient caution, we may irreversibly lose control of autonomous AI systems, rendering human intervention ineffective. Large-scale cybercrime, social manipulation, and other harms could escalate rapidly. This unchecked AI advancement could culminate in a large-scale loss of life and the biosphere, and the marginalization or extinction of humanity.
-
Their words
Society's response, despite promising first steps, is incommensurate with the possibility of rapid, transformative progress that is expected by many experts. AI safety research is lagging. Present governance initiatives lack the mechanisms and institutions to prevent misuse and recklessness, and barely address autonomous systems.
13 September 2023
-
Their words
If a model was capable of self-exfiltration, it would have the option to remove itself from your control.
Self-exfiltration is a key dangerous capabilityaligned.substack.com
29 June 2023
From one piece George Hotz: Tiny Corp, Twitter, AI Safety, Self-Driving, GPT, AGI & God | Lex Fridman Podcast #387 2 beliefs, in the piece's order there
-
Their words
I think we’re going to build super intelligence before we build any sort of robustness in the AI. We cannot build an AI that is capable of going out into nature and surviving like a bird. A bird is an incredibly robust organism. We’ve built nothing like this. We haven’t built a machine that’s capable of reproducing.
-
Their words
What’s ironic about all these AI safety people is they’re going to build the exact thing they fear. We need to have one model that we control and align. This is the only way you end up paper clipped. There’s no way you end up paper clipped if everybody has an AI.
6 June 2023
From one piece Why AI Will Save The World 2 beliefs · pmarca.substack.com
-
Their words
My response is that their position is non-scientific – What is the testable hypothesis? What would falsify the hypothesis? How do we know when we are getting into a danger zone?
-
Their words
My view is that the idea that AI will decide to literally kill humanity is a profound category error. AI is not a living being that has been primed by billions of years of evolution to participate in the battle for the survival of the fittest, as animals are, and as we are. It is math – code – computers, built by people, owned by people, used by people, controlled by people.
25 May 2023
25 April 2023
17 February 2023
-
korrents.com
Liberal societies currently face an existential risk that must be addressed to reach a better future.Their words
This book is my best crack at explaining what I think is an existential risk to liberal societies and what I think we need to do to get to that awesome future I used to be so excited about.
19 December 2022
5 December 2022
26 October 2022
-
Recommendssponsoredrcmnd.app
Template Metaprogramming with C++Their words
I spent the last few weeks reading Template Metaprogramming with C++ by Marius Băncilă. It is currently 50$ on Amazon. Though, technically, the book is about advanced 'template' techniques, it is much more broad and practical. It is one of the 'good programming books'. If you are an experienced C++ programmer, you should give it a peek
1 September 2022
-
Their words, before
That hokey unfashionable techniques like practicing gratitude turn out to have strong scientific evidence behind them
Their words now
I was wrong about gratitude. I thought it was a guaranteed way to become happier
9 August 2022
-
Their words
But the researchers who have investigated this find that scientists do the same thing. They have something that's called knowledge shields, the way all of us do, that that a a a variety of techniques for explaining away inconvenient data and anomalies. But what we're effectively doing is we're holding on to our story.
10 June 2022
From one piece AGI Ruin: A List of Lethalities 4 beliefs, in the piece's order there
-
korrents.com
The field calling itself AI safety is not being remotely productive on the problems that are actually lethal.Their words
It does not appear to me that the field of ‘AI safety’ is currently being remotely productive on tackling its enormous lethal problems.
-
korrents.com
Fast capability gains are likely, and they can break many of the assumptions alignment depends on at the same moment.Their words
Fast capability gains seem likely, and may break lots of previous alignment-required invariants simultaneously.
-
Their words
Many alignment problems of superintelligence will not naturally appear at pre-dangerous, passively-safe levels of capability.
+ 1 more
-
Their words
unaligned operation at a dangerous level of intelligence kills everybody on Earth and then we don’t get to try again.
6 April 2022
-
Their words
Despite the unreasonable effectiveness of simple architectures, most press goes to complex architectures.
19 March 2022
-
Likedrcmnd.app
Moonwalking with EinsteinTheir words
The book Moonwalking with Einstein explores humans' relationship to memory, our ideas of knowledge, and how that's shifted over time. But it begins with an exploration of competitive memorization and the techniques used by memory champions.
23 November 2021
-
Lovedrcmnd.app
The Alignment Problem: Machine Learning and Human ValuesTheir words
I just finished this book a few weeks ago and it is still reverberating in my mind
-
Mixed onrcmnd.app
The Precipice: Existential Risk and the Future of HumanityTheir words
I wouldn’t say that this is the most compelling book I’ve ever read in terms of the prose style or storytelling, but it does provide a very helpful, almost quantitative overview of all the potential threats looming out there
24 September 2021
28 October 2020
-
Likedaffiliate linkrcmnd.app
Mastering the VC GameTheir words
The interests of a Venture Capitalist are different than those of the entrepreneurs building a company they've invested in. Jeff does an awesome job of helping explain how you can get misaligned in your goals versus your investors. Fortunately, he also covers how to avoid it.
7 October 2019
-
Lovedrcmnd.app
The AI Does Not Hate You: Superintelligence, Rationality, and the Race to Save the WorldTheir words
Briefly, I think the book is a triumph.
17 July 2019
-
Usesrcmnd.app
TinyPNGTheir words
When I need to compress images, I use TinyPNG. It uses smart lossy compression techniques to reduce the file size of your PNG files.
21 January 2019
-
Their words
You do not need full code verification to write good software or even to write near-perfect software.
11 October 2018
-
korrents.com
Current techniques restrict pre-trained representation power because standard language models are unidirectional.Their words
We argue that current techniques restrict the power of the pre-trained representations, especially for the fine-tuning approaches. The major limitation is that standard language models are unidirectional, and this limits the choice of architectures that can be used during pre-training.
-
korrents.com
Current techniques restrict pre-trained representation power because standard language models are unidirectional.Their words
We argue that current techniques restrict the power of the pre-trained representations, especially for the fine-tuning approaches. The major limitation is that standard language models are unidirectional, and this limits the choice of architectures that can be used during pre-training.
-
korrents.com
Current techniques restrict pre-trained representation power because standard language models are unidirectional.Their words
We argue that current techniques restrict the power of the pre-trained representations, especially for the fine-tuning approaches. The major limitation is that standard language models are unidirectional, and this limits the choice of architectures that can be used during pre-training.
-
korrents.com
Current techniques restrict pre-trained representation power because standard language models are unidirectional.Their words
We argue that current techniques restrict the power of the pre-trained representations, especially for the fine-tuning approaches. The major limitation is that standard language models are unidirectional, and this limits the choice of architectures that can be used during pre-training.
9 August 2018
-
Likedrcmnd.app
Concepts, Techniques, and Models of Computer ProgrammingTheir words
Another similar and also interesting book is Concepts, Techniques, and Models of Computer Programming, which explains all the models of computer languages and how they fit together.
-
Likedrcmnd.app
The Art of Readable CodeTheir words
What I liked about this book is that it gives you a series of techniques to make that smell go away.
9 March 2018
-
Their words
Neural network pruning techniques can reduce the parameter counts of trained networks by over 90%, decreasing storage requirements and improving computational performance of inference without compromising accuracy.
-
Their words
Neural network pruning techniques can reduce the parameter counts of trained networks by over 90%, decreasing storage requirements and improving computational performance of inference without compromising accuracy.
7 March 2018
Nothing matches.
What is a korrent?
A korrent is a belief a person has stated in their own words: one sentence stating the claim, backed by a quote and a source, kept at korrents.com.
Under a name here, the quoted block is what they actually said. The korrent beneath it is the claim those words support, in korrents' wording — tap it to see the record, its source, and who else holds it.
Nobody here wrote their own korrents. They are compiled from public statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they do, this site shows a machine translation beneath the post, in this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the serif above is exactly what the person published, and it is what to quote them on. A translation can be wrong in ways that matter, especially about tone.
Only the post's own words are translated. A quoted post, a linked article and a belief on korrents.com are left in their original language.