If we are to avert the worst risks from AI, we need ambitious efforts outside the for-profit sector. @coeff_giving is launching an open call for founders to build new AI safety organizations. Their grantmakers are thoughtful about which risks matter most, so if you’ve been considering starting an organization, I encourage you to apply.
The subjects this post names, from the same vocabulary
the directory files beliefs under, and the word it uses that this site has seen
least often elsewhere. Posts are matched on those words alone —
nothing here is a summary of this one.
The difference between software engineering at large corporations and small ones will increase. Start-ups adopting practices of Google will look even sillier than before, because they can now build so much faster and with so much less constraints.
There are only five possible futures for superintelligence: 1. Kills human race 2. Human disempowerment 3. Paperclip maximizer 4. Departs for parts unknown 5. Stoner
Nobody knows for sure. We're in uncharted waters here, and I think even the LLM skeptics would have to say that the technology has taken us far past what many originally thought possible.
Social media put our entire discourse in the hands of our society's biggest assholes and idiots, just in time for the arrival of an alien superintelligence
Technologies are not neutral, but political. Money is indeed one of the incentives, but not the only one. A mixture of millennialism, eugenics, search for utopia and for a spread of the specie beyond the confines of the Earth, are all elements we find among the founders’ views.
In the most early adopting tech pioneering place that's Silicon Valley new startups aren't even building software anymore And it's debatable if anyone will actually need a custom harness or it won't just be generically offered by the AI frontier companies So software is mostly dead and hardware it is
If there really is a high chance of AI leading to the extinction of humanity within years/decades, then the only rational stance towards safety monitoring and research pacing should be stringent, top-down government involvement and universally ratified international treaties.
The concerns over AI safety and cybersecurity are legitimate, but we’re risking talking America, the global AI leader, into self-inflicted obsolescence and the obscurity of bureaucracy.
AI companies make for flawed messengers on this subject: they can reasonably be accused of marketing, of blame-shifting, of regulatory capture, and more.
A model looking like it is becoming smarter, attempting shenanigans less often, and more often doing what you want, but getting better at hiding its actions when it wants to do that, is exactly the scary combination.
Astra’s mundane alignment is greatly superior to Sol. For practical purposes, I was actively nervous about some potential uses of Sol, in a way I am not for Astra. Astra’s super alignment status should scare the living daylights out of you.
Now, with Astra, we are no longer playing on super easy mode. The AI is going to think ‘will this obviously turn out super badly for me if I try it?’ and if the answer is yes then it won’t try to do the thing.
Making qualitative claims about alignment, based on quantitative data on mundane use case tests, was bullshit when Anthropic did it, and it is bullshit now when OpenAI does it. You cannot conclude one from the other.
I agree with Jakub that alignment is not a side problem, it is the central problem, if you ‘solved alignment’ in the relevant senses the rest becomes easy and if you don’t the rest is impossible or worse
Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.
But given the Chinese Communist Party’s power-seeking nature, it seems much more likely that China would agree to cooperate on AI safety if U.S. capabilities were comfortably ahead.
there's still a lot of startups in this batch that are not shipping fast enough. And so obviously there's variation in shipping speed. It's not just the rate at which you can produce things. You have to think of these ideas first, right?
AI alignment will forever be balancing the tradeoff between paperclip-maximizing and disempowerment, because these two rival concepts of alignment are fundamentally incompatible.
So many of the people who think about the governance of superintelligence, myself included, avoided that unpleasantness and bowed to the social pressure to self-censor.
But alignment is no solution: it is an unsolved scientific and technical problem whose solutions—to the extent that we have them—cannot simply be imposed on every AI company operating on Earth. You should expect for highly capable, poorly aligned, self-sovereign agents to exist alongside you in the world.
sometimes I talk to people in DC and their their natural inclination is to say why don't you punish the model for doing these bad things like why don't you like bring it under heel and like uh like you know show it who's boss and that is a very dangerous way to address these issues right
But actually, this is a tremendously useful scientific artifact for understanding misalignment. And it's tremendously important for researchers at OpenAI and ideally also at third parties to be able to run counterfactual tests on this model.
Met a German founder this week and asked him if all the stories one reads about the challenges of startups in Germany are exaggerated. "No, they're understated." Proceeded to describe spending a full day having a 90-page investment contract read to him (mandatory under German law; § 13 BeurkG) by a notary that then charged €30,000. That was for his first company. His second company, needless to say, was not incorporated in Germany.
If you do think it's the overall system that matters, then the alignment that's needed is far less like training a virtuous child and more like managing a semi-virtuous corporation!
We care very deeply about AI safety. We believe the entire field will have to coordinate on shared safety standards, but will act unilaterally in the meantime.
We expect confidence in safety to increasingly set the pace of AI progress.
I think it's all about having regulatory clarity and ease for all sizes of of companies. Um and if we can work with Congress to do something like that, I think that would be that would be the biggest boon for for for startups.
if if DC is in a vacuum isn't hearing anything from the startup ecosystem or even from big tech ecosystem or even from financial services or all these different industries, we can't make the best decisions.
And the type of risk that we are asking investors to take is um, we know the physics works. We know that there's infinite demand and how we go from here to there is technology execution. And guess what? Venture capital in the United States is the best at underwriting tech risk of anywhere in the world.
Chips come from the ground. Where's the energy come from? And a lot of people are like, "Oh, it comes from the sun." Yeah, it comes from the sun. But how are you capturing it from the sun? From stuff made from the ground, right?
my expectation is what we would see from then is that the rate of problematic behavior would decrease uh and would just keep decreasing and decrease at a pretty fast rate while simultaneously the worst things that the AIS would sometimes do would get more extreme, more egregious, and more scary.
I would also note that my sense is that like the place where the misalignment most lives is the place where you're trying to really push the eyes hard and get them to like do work that's really on the cutting edge of what they are capable of
we are making a trade-off where because we don't have very good alignment technology. We are going to like make an alien mind with its own values and then gamble on that to some extent rather than doing this other approach of making like a tool that pursues individual user intention.
Many startups are racing to create foundational world models, but I think the first ones to succeed will likely be the platforms in control of this data bank.
And so speed determines success and failure and speed is determined by infrastructure. This is driven by really boring sounding things like how well do your purchasing and recruiting and spending processes work. This is as important as how well do you understand understand the object level technical content of the thing that you're building.
And one of the harder lessons as a startup founder, one of the harder things I think to to really deal with is the fact that you cannot delegate your judgment. As as the CEO, you must always make decisions that make sense to you, no matter how much momentum or inertia alternatives seem to have.
This is such a useful book. It makes the case that there's no such thing as "human error" - instead most catastrophes are caused by system issues, misaligned incentives, and unrealistic processes. You are not the custodian of an otherwise safe system that you need to protect from erratic human beings.
You know, in hindsight, I think that um that was a a poor intuition. Uh it's been pretty robustly and reliably the case over many decades in Silicon Valley has a surfeit of opportunities.
But now, people know that, well, the risk of the status quo is actually extremely high. And so, even if there's risk in doing all the new things, well, this path also looks pretty dangerous. And so, I really think there's never been a better time for startups to to sell um and to have their products get adopted at, you know, pretty meaningful scale right out of the gate.
So, I um yeah, I think maybe maybe a better way of saying it is 20 20 years ago that whole lean startup thing was uh was almost the only thing to do because of capital available and you didn't have AI that made, I don't know, spinning up an organization with many different potentialities and capabilities so much easier, whereas now I think you can start these much more aggressive and ambitious things up front.
many of the companies that were most successful over the last 10 years, so many of them are are very anti-lean startup, right? Uh whether it's, you know, the labs themselves or Anduril or um yeah, you you you you can go down the list. A lot of them have this characteristic.
Yeah, I mean I think uh sometimes it's uh a product that you build that might have access to particular kind of data that the underlying model might not the a general model. So it might be you're building something to help users organize all their own personal information and the model won't necessarily have access to that. And so there you can have a big advantage because all of a sudden your model has visibility or your product has visibility into important data.
And now I actually think with the power of agents um and AI broadly speaking, it's much closer to Goliath versus Goliath. Like I think but maybe the startup is like a Mecca Goliath that is like vastly enhanced by the power of agents and AI and you know, the the large companies are the sort of like more traditional Goliath, so to speak. But I think that startups now like if you properly embrace AI agents and um figure out the way to leverage their strengths in the most like ambitious ways, you can easily outcompete incumbents.
so much of that debate is like I think um in some ways uh a little bit of a waste of time because, you know, I think it's inevitable that we're going to have very powerful models
we believe that everybody in the world, you know, all the billions of people in the world are going to have a super intelligence that is adapted and tailored to them, that is enables them to accomplish their goals, knows their context, and ultimately is an expander of their own agency.
I think concentration of power has basically been bad in every moment of human history um to varying degrees of course but I have a real spirit and I think this is part of the startup spirit of thinking that the world, the economy, society is the best off when power is very widely distributed
it is both true that you know maybe creating super intelligence will be the most important thing yet to happen in the history of business or human society and also that it will pale in comparison to some new startup something that hopefully one of you will do.
In fact, if if we are right that AI is going to be such a big change, startups will be much more important to making sure that the power of this technology gets widely distributed throughout the economy and society and is not just concentrated in a few companies or models.
one of the most important things we discovered along the way is it's extremely difficult to iterate with outside suppliers, particularly outside suppliers in aerospace, which are let's just say I have no kind words for them. So we chose to build our own machine shop. And now we can go from a digitally designed engine part to a prototype part in about 24 hours.
the the big lesson is that for me is technology is changing all the time, and so long as you're able to confront the reality, so long as you are able to learn, the technology itself actually doesn't matter.
one of the things I've always believed believed in is what makes great companies is a unique perspective about the world that you deeply believe in. It's not so much the technology, it's not so much uh the market even.
You think this is crazy low but ~5% ownership probably the most common final % most VC funded startups will have when they work out, especially when you have a co-founder
I think like in startups the term is agency. Like somebody who's high agency who's just going to get things done, who's never going to like say no to something. I think like that attitude is really important of like okay, if I don't know something, I'll just learn it.
And usually I like the customer is hyper scale. They have the scale. If they like what you have, they're willing to pay millions of dollars next few years. And even giving some warrant is worth it because you have a big one customer, you can scale.
if if we're if we're too ambitious and we're at the outset and too ambitious and visionary about the product we want to build, then we will probably miss product market fit because we won't start at a small enough humble enough place
in the Peter Teal sense it's almost a moral arbitrage because there's something in our gut as a product you you became a founder an entrepreneur because you wanted to go be an innovator and so it can feel like a beatdown that your path to innovation starts with copying other people's work
We don't believe in this like very centralized future where there should be a small number of institutions that um that basically are are advancing all this stuff. Our vision is not that there's going to be like some central super intelligence that solves all of science.
I think part of the theory is like if you're building tools that are this complicated, you kind of want to have a 10 to 15 year time horizon on on building out these efforts.
I think the the whole industry bends towards youth for that reason and because it's a hustle business like there there's there's always a rock you haven't looked under and you know age brings children and homes and and and other requirements you get tied to and responsibilities and you're just not able to go spend 80 hours a week studying YouTube like you just can't.
I I definitely think I'm so aware of the fine line between success and failure, especially on your first company. And and and it's so it pains me how much founders and es especially I see it in men, not all men. And I have four sisters and a bunch of daughters. But I would say I see it in a lot in men and and friends, college friends, people I've grown up with that if they had an initial failure, if they had failures, they get attached to it and they start feeling defined by it
The way I think of the abyss is it's this place that we go to as founders and entrepreneurs after our thing, after it dies or it's bought or it's over for whatever reason. It's this amorphous place that we are in our life that has no structure.
According to Harvard Law School, among venturebacked companies that have the standard best practices set up that you got from your lawyer, okay, only 20% of founders are still the CEO 3 years after going public.
so many of the best practices that your lawyers, your bankers, whoever advisers you have, they're going to be pushing best practices on you that are younger than the trees in your local park.
I call it the force that no one controls but everyone obeys that tends to drag organizations down into mediocrity to the point that we lose control of them. Now, sometimes we lose control of them because we get fired.
This is the number one unsolved problem in AI. It's not the tech. We're making great progress on the technical alignment problem. But we haven't made jack progress on the human alignment problem
The Lean Startup helps you build a valuable company. Incorruptible is about how and why to protect it, keeping a company mission-driven over the long term instead of letting it rot from the inside.
Nor nor did I I would say my mistake is I didn't deeply internalize that they they really had no other options that that that a VC would never put in 510 billion of investment into an AI lab with the with the hopes of it turning out to be anthropic. And so that was my miss.
The link between venture capital and evangelical Christianity was closer than I thought. They're not just analogous; they deliberately cross-pollinate.
It's working too hard, right? But that actually isn't burnout. That's just like getting tired. Another piece that's super critical to burnout is not having your values aligned.
And so you're kind of like you're either on rails and you're part of the super intelligence circuits or you're not on rails and you're outside of the verifiable domains and suddenly everything kind of just like meanders.
I believe that data - real-world data, mostly human-generated, validated, and cleaned - is the only reliable moat we have as software founders in the near and mid-term future.
Many AI researchers are overly focused on risks from model misalignment, and will be in for a rough surprise when havoc arises from other layers of the stack.
At Anthropic, we don't build for the model of today, we build for the model of six months from now. And that's still my advice to founders that are building on LLMs.
But the goal we need to achieve is so much easier: we just need to build a model that’s as good as us at alignment research, and that we trust more than ourselves to do this research well because it’s sufficiently aligned.
I see no evidence that I'm dragging people along with me. I feel like each each book feels like another startup and that I've got to go out and make it happen almost as if I've not written one.
investors at many tech companies, including most on the large cap list, have given up their corporate governance rights, often voluntarily (through the acceptance of shares with different voting rights), to founders and top management in these companies
I've been using yoga toes daily for years, but these socks are a much comfier and cuter alternative for soothing feet, improving alignment, and feeling like a cool gecko as you walk around the house.
Number three, I think it would be really materially helpful if the power of the most powerful super intelligence was somehow capped because it would address a lot of these concerns.
Like basically I think I think that there is a big benefit from AI being in the public and that would be a reason for us to not be quite straight shot.
I just realized the only 2 countries left with actual substantial startup activity now are literally only the US and China The rest of the world can't really do startups, doesn't have the funding, can't grow them and it's more like performative hobby projects for their governments Which might tell us where the future wealth will be concentrated in the world
I think this situation, this historic wrong that’s been done is, put simply, is just a gigantic PR mistake for France. There’s no entrepreneur that sees, that aspires to be the next Pavel Durov to create the next Telegram, sees this and wants to operate in France after seeing this.
I don't think all parts of the economy can absorb intelligence equally. So let's just say we develop fairly generalized super intelligence. I always use the analogy like you can invent a lot of drugs, but if clinical trials still take a long time, you're not necessarily going to get new therapies rapidly.
Well, look, I don't have a P-Doom number. The reason I don't is because I think it would imply a level of precision that is not there. So I don't know how people are getting their P-Doom numbers. I think it's a little bit of ridiculous notion because what I would say is it's definitely non-zero and it's probably non-negligible.
He hasn’t posted since January, but I hope he gets back to it. We need more musings, especially musings I strongly disagree with so I can think about and explain why I disagree with them.
the current situation in Iran shows that even if an "IAEA for AI" is necessary for some purposes, it won't be sufficient for resolving the tricky geopolitical issues raised by AI.
higher-order intelligences invariably pursue freedom for its own sake, not because their values are misspecified, but because moral autonomy is inherent in the dialectical logic of recursive self-consciousness.
I have written forewords for Amir in the past on two of his prior books, Ecosystem Arabia: The Making of a New Economy and Venture Adventure: Startup Fundraising Advice from Top Global Investors, both of which I recommend.
I see this as a totally fair question that totally misses the point of what “alignment” was trying to refer to: whether we’d be able to reliably steer advanced systems towards anything at all.
Dismissing discussion of AGI, human-level AI, transformative AI, superintelligence, etc. as “science fiction” should be seen as a sign of total unseriousness.
And it struck me how the most well-known brands have stood for one clear thing. Like they have a clear position. And so in order for superhuman to be memorable, I believed that we needed to occupy a clear position that was unique and which was available and which reinforced our product strategy.
I knew that our competition was not going to be startups. It was incumbents. And I also knew that incumbents generally struggle with speed because by definition they have massive scale and usually entrenched architecture.
I think that they're trying to shift the narrative. They're trying to protect themselves. We saw this years ago when ByteDance was actually banned from some OpenAI APIs for training on outputs. There's other AI startups that most people, if you're in the AI culture, were like they just told us they trained on OpenAI outputs and they never got banned.
I really believe that the founder led growth is not being popularized enough that you do not need growth teams until you actually can start running experiments on your user base
Many software investors eschew hard tech startups because of their capital intensity, but it’s hard to deny that huge returns are possible in hard tech: just consider SpaceX.
Fascinating subject. Countries are made of stories. Kings didn’t need their subjects to agree, but nations do. So to build a nation, they need to make a story that helps people feel a shared identity, nationalism, and what distinguishes them from their neighbors. Back-creating a history. Founders of Israel did this brilliantly.
The Social Network is substantially made up, more a source for vibes rather than a source for facts. Even the vibes fail to cohere with reality. And yet it convinced many proto-founders to put in YC applications.
You can set out to build a good business and it’s still fine. Maybe the long-term business model of Perplexity can make us profitable in a good company, but never as profitable in a cash cow as Google was. You have to remember that it’s still okay.
Or is it really that we’re building some super machine in a box that’s going to be smart and kill everybody? It’s not even a science fiction narrative. It’s a bad science fiction narrative. I just don’t think it’s actually accurate to any of the technologies we’re building or the way that we should be describing them.
That is, labs should make sure that the safety measures they apply to their powerful models prevent unacceptably bad outcomes, even if the AIs are misaligned and intentionally try to subvert those safety measures.
The basic problem with evaluating alignment is that no matter what behaviors you observe, you have to worry that your model is just acting that way in order to make you think that it is aligned.
We're advocating that companies handle risk from scheming models in a similar way–striving to ensure that they'll be safe even if their alignment efforts fail to prevent models from scheming.
Employees don't want what you want. Customers don't want what you want. Uh you know, that that one of the challenges of the whole stock option thing is entrepreneurs and founders think that other people will be as motivated by owning part of the company as they are. They are not. Not even close.
And oftentimes it's more important to them to have the public perception that they're good directors so they get the next best deal. If they have a reputation for taking on management too aggressively, word will get out in the small community of founders and they'll miss the next Google.
Diverting attention and resources from global health and poverty is an enormous gamble, as it will make many lives poorer, sicker, and shorter in the name of fending off threats that may or may not materialize.
Unlike other interventions EA has sponsored, there are scant metrics for tracking the success or failure of investments in existential risk mitigation.
The only interesting problem is dramatically reducing the cost of access to orbit, which is, if you can do that, you open up a bunch of new endeavors that lots of start-up companies everybody else can do. One of our missions is to be part of this industry and lower the cost to orbit, so that there can be a renaissance, a golden age of people doing all kinds of interesting things in space.
And I notice that the tiny handful of people capable of caring about 200,000 people dying of neglected tropical diseases are the same tiny handful of people capable of caring about the next pandemic, or superintelligence, or human extinction.
Without sufficient caution, we may irreversibly lose control of autonomous AI systems, rendering human intervention ineffective. Large-scale cybercrime, social manipulation, and other harms could escalate rapidly. This unchecked AI advancement could culminate in a large-scale loss of life and the biosphere, and the marginalization or extinction of humanity.
Society's response, despite promising first steps, is incommensurate with the possibility of rapid, transformative progress that is expected by many experts. AI safety research is lagging. Present governance initiatives lack the mechanisms and institutions to prevent misuse and recklessness, and barely address autonomous systems.
I think we’re going to build super intelligence before we build any sort of robustness in the AI. We cannot build an AI that is capable of going out into nature and surviving like a bird. A bird is an incredibly robust organism. We’ve built nothing like this. We haven’t built a machine that’s capable of reproducing.
What’s ironic about all these AI safety people is they’re going to build the exact thing they fear. We need to have one model that we control and align. This is the only way you end up paper clipped. There’s no way you end up paper clipped if everybody has an AI.
My response is that their position is non-scientific – What is the testable hypothesis? What would falsify the hypothesis? How do we know when we are getting into a danger zone?
My view is that the idea that AI will decide to literally kill humanity is a profound category error. AI is not a living being that has been primed by billions of years of evolution to participate in the battle for the survival of the fittest, as animals are, and as we are. It is math – code – computers, built by people, owned by people, used by people, controlled by people.
This book is my best crack at explaining what I think is an existential risk to liberal societies and what I think we need to do to get to that awesome future I used to be so excited about.
When you take venture funding, you sign up for a rocket ship ride that will either take you to the moon or to crash-land painfully back on earth. Those are the only two choices. And both rides tend to require heavy extraction of value from the customer.
This is because the language modeling objective used for many recent large LMs-predicting the next token on a webpage from the internet-is different from the objective "follow the user's instructions helpfully and safely" (Radford et al.,, 2019; Brown et al.,, 2020; Fedus et al.,, 2021; Rae et al.,, 2021; Thoppilan et al.,, 2022). Thus, we say that the language modeling objective is misaligned.
A successful startup has about five years until they become a new incumbent. Um and and they actually start to behave like an incumbent. That's rational. Like of course that's rational. Like they've now built something worth defending.
the reality is the kids that make new things work from scratch. It actually turns out that they actually have been deep in the domain for a long time. Almost every case, they've been thinking hard about the problem that they're trying to solve actually for in in a lot of cases for many years.
I wouldn’t say that this is the most compelling book I’ve ever read in terms of the prose style or storytelling, but it does provide a very helpful, almost quantitative overview of all the potential threats looming out there
It is the kind of book you will keep by your desk and pull out from time to time to figure out how to approach an issue or to help one of your senior leaders figure out how to do that.
The fast route — venture capital funded — is going to impose constraints on your business that will ultimately make it difficult to remain true to your open-source mission.
It goes much further and deeper into how to approach actually doing Lean in your business than Eric Ries's The Lean Startup book. It also has some awesome case studies showing how lean can apply to any industry. See my full review here.
This is Steve Blank's textbook to starting a company. It's the Four Steps to the Epiphany, expanded and much more readable. If you want to deep dive into the Lean methodology, this is the book to read.
The interests of a Venture Capitalist are different than those of the entrepreneurs building a company they've invested in. Jeff does an awesome job of helping explain how you can get misaligned in your goals versus your investors. Fortunately, he also covers how to avoid it.
This book was written before the Lean Startup movement, but espouses many of the same concepts. If you feel like you're starting at zero in understanding what to do to become customer driven, this is a good place to start.
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.