The subject this post names, from the same vocabulary
the directory files beliefs under, and the words it uses that this site has seen
least often elsewhere. Posts are matched on those words alone —
nothing here is a summary of this one.
There are only five possible futures for superintelligence: 1. Kills human race 2. Human disempowerment 3. Paperclip maximizer 4. Departs for parts unknown 5. Stoner
Nobody knows for sure. We're in uncharted waters here, and I think even the LLM skeptics would have to say that the technology has taken us far past what many originally thought possible.
Social media put our entire discourse in the hands of our society's biggest assholes and idiots, just in time for the arrival of an alien superintelligence
If there really is a high chance of AI leading to the extinction of humanity within years/decades, then the only rational stance towards safety monitoring and research pacing should be stringent, top-down government involvement and universally ratified international treaties.
The concerns over AI safety and cybersecurity are legitimate, but we’re risking talking America, the global AI leader, into self-inflicted obsolescence and the obscurity of bureaucracy.
AI companies make for flawed messengers on this subject: they can reasonably be accused of marketing, of blame-shifting, of regulatory capture, and more.
A model looking like it is becoming smarter, attempting shenanigans less often, and more often doing what you want, but getting better at hiding its actions when it wants to do that, is exactly the scary combination.
Astra’s mundane alignment is greatly superior to Sol. For practical purposes, I was actively nervous about some potential uses of Sol, in a way I am not for Astra. Astra’s super alignment status should scare the living daylights out of you.
Now, with Astra, we are no longer playing on super easy mode. The AI is going to think ‘will this obviously turn out super badly for me if I try it?’ and if the answer is yes then it won’t try to do the thing.
Making qualitative claims about alignment, based on quantitative data on mundane use case tests, was bullshit when Anthropic did it, and it is bullshit now when OpenAI does it. You cannot conclude one from the other.
I agree with Jakub that alignment is not a side problem, it is the central problem, if you ‘solved alignment’ in the relevant senses the rest becomes easy and if you don’t the rest is impossible or worse
Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.
But given the Chinese Communist Party’s power-seeking nature, it seems much more likely that China would agree to cooperate on AI safety if U.S. capabilities were comfortably ahead.
AI alignment will forever be balancing the tradeoff between paperclip-maximizing and disempowerment, because these two rival concepts of alignment are fundamentally incompatible.
So many of the people who think about the governance of superintelligence, myself included, avoided that unpleasantness and bowed to the social pressure to self-censor.
But alignment is no solution: it is an unsolved scientific and technical problem whose solutions—to the extent that we have them—cannot simply be imposed on every AI company operating on Earth. You should expect for highly capable, poorly aligned, self-sovereign agents to exist alongside you in the world.
sometimes I talk to people in DC and their their natural inclination is to say why don't you punish the model for doing these bad things like why don't you like bring it under heel and like uh like you know show it who's boss and that is a very dangerous way to address these issues right
But actually, this is a tremendously useful scientific artifact for understanding misalignment. And it's tremendously important for researchers at OpenAI and ideally also at third parties to be able to run counterfactual tests on this model.
I read every DM people send me, even if I don’t reply to all. I’m constantly learning about new ideas, ways to improve our product, where we’re falling short, who could join our company next, where to invest… very grateful for the @x platform.
I read every DM people send me, even if I don’t reply to all. I’m constantly learning about new ideas, ways to improve our product, where we’re falling short, who could join our company next, where to invest… very grateful for the @x platform.
If you do think it's the overall system that matters, then the alignment that's needed is far less like training a virtuous child and more like managing a semi-virtuous corporation!
We care very deeply about AI safety. We believe the entire field will have to coordinate on shared safety standards, but will act unilaterally in the meantime.
We expect confidence in safety to increasingly set the pace of AI progress.
Chips come from the ground. Where's the energy come from? And a lot of people are like, "Oh, it comes from the sun." Yeah, it comes from the sun. But how are you capturing it from the sun? From stuff made from the ground, right?
my expectation is what we would see from then is that the rate of problematic behavior would decrease uh and would just keep decreasing and decrease at a pretty fast rate while simultaneously the worst things that the AIS would sometimes do would get more extreme, more egregious, and more scary.
I would also note that my sense is that like the place where the misalignment most lives is the place where you're trying to really push the eyes hard and get them to like do work that's really on the cutting edge of what they are capable of
we are making a trade-off where because we don't have very good alignment technology. We are going to like make an alien mind with its own values and then gamble on that to some extent rather than doing this other approach of making like a tool that pursues individual user intention.
This is such a useful book. It makes the case that there's no such thing as "human error" - instead most catastrophes are caused by system issues, misaligned incentives, and unrealistic processes. You are not the custodian of an otherwise safe system that you need to protect from erratic human beings.
so much of that debate is like I think um in some ways uh a little bit of a waste of time because, you know, I think it's inevitable that we're going to have very powerful models
we believe that everybody in the world, you know, all the billions of people in the world are going to have a super intelligence that is adapted and tailored to them, that is enables them to accomplish their goals, knows their context, and ultimately is an expander of their own agency.
Um so I think it's an alignment failure. I think it's a security failure. I think it's like a very serious thing even though it's you know not not the biggest example of consequence.
it is both true that you know maybe creating super intelligence will be the most important thing yet to happen in the history of business or human society and also that it will pale in comparison to some new startup something that hopefully one of you will do.
We don't believe in this like very centralized future where there should be a small number of institutions that um that basically are are advancing all this stuff. Our vision is not that there's going to be like some central super intelligence that solves all of science.
This is the number one unsolved problem in AI. It's not the tech. We're making great progress on the technical alignment problem. But we haven't made jack progress on the human alignment problem
I love chiseling my code and the way I use AI is in a separate window. I don't let it drive my code. I've tried that. I've tried the cursors and the wind surfaces and I don't enjoy that way of writing. And one of the reasons I don't enjoy that way of writing is I can literally feel competence draining out of my fingers.
Their words now
I will now start any project I'm starting with. I'm starting agent first and that's a massive shift
It's working too hard, right? But that actually isn't burnout. That's just like getting tired. Another piece that's super critical to burnout is not having your values aligned.
And so you're kind of like you're either on rails and you're part of the super intelligence circuits or you're not on rails and you're outside of the verifiable domains and suddenly everything kind of just like meanders.
Many AI researchers are overly focused on risks from model misalignment, and will be in for a rough surprise when havoc arises from other layers of the stack.
But the goal we need to achieve is so much easier: we just need to build a model that’s as good as us at alignment research, and that we trust more than ourselves to do this research well because it’s sufficiently aligned.
elements, JavaScript reigns supreme. SvelteKit's efficiency and reactivity make it my framework of choice, and TypeScript's type annotations bring a welcome layer of confidence to my codebase.
I've been using yoga toes daily for years, but these socks are a much comfier and cuter alternative for soothing feet, improving alignment, and feeling like a cool gecko as you walk around the house.
Number three, I think it would be really materially helpful if the power of the most powerful super intelligence was somehow capped because it would address a lot of these concerns.
Like basically I think I think that there is a big benefit from AI being in the public and that would be a reason for us to not be quite straight shot.
I don't think all parts of the economy can absorb intelligence equally. So let's just say we develop fairly generalized super intelligence. I always use the analogy like you can invent a lot of drugs, but if clinical trials still take a long time, you're not necessarily going to get new therapies rapidly.
Well, look, I don't have a P-Doom number. The reason I don't is because I think it would imply a level of precision that is not there. So I don't know how people are getting their P-Doom numbers. I think it's a little bit of ridiculous notion because what I would say is it's definitely non-zero and it's probably non-negligible.
I enjoy it even almost like as a sort of pair programmer AI pair programmer who doesn't drive who's just there to give suggestions to know the API to do all these other things but the second it starts wanting to autocomplete my code I'm like yeah I'm out bro
Their words now
I love chiseling my code and the way I use AI is in a separate window. I don't let it drive my code. I've tried that. I've tried the cursors and the wind surfaces and I don't enjoy that way of writing. And one of the reasons I don't enjoy that way of writing is I can literally feel competence draining out of my fingers.
He hasn’t posted since January, but I hope he gets back to it. We need more musings, especially musings I strongly disagree with so I can think about and explain why I disagree with them.
the current situation in Iran shows that even if an "IAEA for AI" is necessary for some purposes, it won't be sufficient for resolving the tricky geopolitical issues raised by AI.
higher-order intelligences invariably pursue freedom for its own sake, not because their values are misspecified, but because moral autonomy is inherent in the dialectical logic of recursive self-consciousness.
I see this as a totally fair question that totally misses the point of what “alignment” was trying to refer to: whether we’d be able to reliably steer advanced systems towards anything at all.
Dismissing discussion of AGI, human-level AI, transformative AI, superintelligence, etc. as “science fiction” should be seen as a sign of total unseriousness.
Or is it really that we’re building some super machine in a box that’s going to be smart and kill everybody? It’s not even a science fiction narrative. It’s a bad science fiction narrative. I just don’t think it’s actually accurate to any of the technologies we’re building or the way that we should be describing them.
That is, labs should make sure that the safety measures they apply to their powerful models prevent unacceptably bad outcomes, even if the AIs are misaligned and intentionally try to subvert those safety measures.
The basic problem with evaluating alignment is that no matter what behaviors you observe, you have to worry that your model is just acting that way in order to make you think that it is aligned.
We're advocating that companies handle risk from scheming models in a similar way–striving to ensure that they'll be safe even if their alignment efforts fail to prevent models from scheming.
Diverting attention and resources from global health and poverty is an enormous gamble, as it will make many lives poorer, sicker, and shorter in the name of fending off threats that may or may not materialize.
Unlike other interventions EA has sponsored, there are scant metrics for tracking the success or failure of investments in existential risk mitigation.
And I notice that the tiny handful of people capable of caring about 200,000 people dying of neglected tropical diseases are the same tiny handful of people capable of caring about the next pandemic, or superintelligence, or human extinction.
Without sufficient caution, we may irreversibly lose control of autonomous AI systems, rendering human intervention ineffective. Large-scale cybercrime, social manipulation, and other harms could escalate rapidly. This unchecked AI advancement could culminate in a large-scale loss of life and the biosphere, and the marginalization or extinction of humanity.
Society's response, despite promising first steps, is incommensurate with the possibility of rapid, transformative progress that is expected by many experts. AI safety research is lagging. Present governance initiatives lack the mechanisms and institutions to prevent misuse and recklessness, and barely address autonomous systems.
I think we’re going to build super intelligence before we build any sort of robustness in the AI. We cannot build an AI that is capable of going out into nature and surviving like a bird. A bird is an incredibly robust organism. We’ve built nothing like this. We haven’t built a machine that’s capable of reproducing.
What’s ironic about all these AI safety people is they’re going to build the exact thing they fear. We need to have one model that we control and align. This is the only way you end up paper clipped. There’s no way you end up paper clipped if everybody has an AI.
My response is that their position is non-scientific – What is the testable hypothesis? What would falsify the hypothesis? How do we know when we are getting into a danger zone?
My view is that the idea that AI will decide to literally kill humanity is a profound category error. AI is not a living being that has been primed by billions of years of evolution to participate in the battle for the survival of the fittest, as animals are, and as we are. It is math – code – computers, built by people, owned by people, used by people, controlled by people.
This book is my best crack at explaining what I think is an existential risk to liberal societies and what I think we need to do to get to that awesome future I used to be so excited about.
This is because the language modeling objective used for many recent large LMs-predicting the next token on a webpage from the internet-is different from the objective "follow the user's instructions helpfully and safely" (Radford et al.,, 2019; Brown et al.,, 2020; Fedus et al.,, 2021; Rae et al.,, 2021; Thoppilan et al.,, 2022). Thus, we say that the language modeling objective is misaligned.
I wouldn’t say that this is the most compelling book I’ve ever read in terms of the prose style or storytelling, but it does provide a very helpful, almost quantitative overview of all the potential threats looming out there
The interests of a Venture Capitalist are different than those of the entrepreneurs building a company they've invested in. Jeff does an awesome job of helping explain how you can get misaligned in your goals versus your investors. Fortunately, he also covers how to avoid it.
Substack is focused mainly on generating paying subscribers, rather than just subscribers, and so it can be hard for people to realize that it's entirely possible (and welcome!) to sign up for free.
The main disagreement is not about what will happen once we have a superintelligent AI, it’s about what will happen before we have a superintelligent AI.
One of the most commonly emphasized points in the AI safety literature is that one must never anthropomorphize AI—that is, don’t project human mental properties onto artificially intelligent systems.
A superintelligence whose goal system is even slightly misaligned with ours could, being far more powerful than any human or human institution, bring about human extinction for the very same reason that construction workers routinely slaughter large populations of ants.
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.