What public figures publish and believe, in their own words.
About this feed
Highlights: posts that did unusually well for the person who wrote them, everything
they published at length, each release and new project, and every belief — at most
two a day from anyone. Day by day, newest
day first; within a day, the people with the most beliefs on this site come first. Nothing
else orders it. Show everything instead.
The quoted blocks are what people actually said; a beneath one is
the belief those words support, in korrents' wording. Nobody here wrote their own page.
Top people are the people in this feed with the most beliefs on this site, then the
most here. Choose an area and the row leads with the people whose beliefs are about it;
tap a face for their feed.
I don't require anyone to be as excited about AI or agents as I am. It's completely fine. But we also need someone in this camp in this Linux camp who are really excited about AI, who really leaning in with agents.
And I do think some of that is positive and energizing, but I also think you can certainly go too far with it, and I also think the current pace is not sustainable.
I’m thinking about banning Claude code at Shopify until they change their mind and read AGENTS.md and .agents/skills etc.
Insisting on only reading CLAUDE.md sometimes leads to split brain problems when different team members use different tools. Just unnecessary.
Maybe I just read too much AI text now vs before, but I really feel like the model's ability to produce text I actually want to read and understand went downhill with newer releases.
All this combined with the survey data on Americans’ stated reasons for opposing data centers makes me think that attitudes toward AI are probably a meaningful factor for a minority of people opposed to data centers, but I don’t think the backlash overall is driven by people’s thoughts on AI.
Someone asked what I'd do if I were 17. I'd learn how to build LLMs from scratch, and then train ones as powerful as I could with whatever hardware I could get access to.
Preventing companies from building data centers in America wouldn't slow down the rate of progress in AI. It would merely slow down the rate of progress in AI in America.
If you do think it's the overall system that matters, then the alignment that's needed is far less like training a virtuous child and more like managing a semi-virtuous corporation!
Every step we take that makes the models less governable and controllable also makes them less useful, so we will not be able to continue using them, which reduces the value, and which means delaying the release is the only option
LLMs constantly hallucinate and cannot be trusted. I still have to verify and iterate a lot but now I usually focus on architecture and design instead of code style.
As the context grows, a model starts paying less attention to instructions in the middle of the context in favor of what is at the beginning and the end.
And so when you look inside these big AI models, the mathematical objects that you see look a lot like the things that you see in neuroscience. So if you look at how do how do these AI models represent um just like represent concepts and you look at the parts of the brain that represent concepts, you see these you see very similar geometry.
I was thinking about how I recognize AI writing, and one big tell is excessively colorful verbs. A handful of journalists might write that a proposal "drew" 100 votes, but any normal person will just say "got".
the first is cognitive debt. So the more that you use AI, it's it's sort of the erosion of your ability to have good memory and have good understanding of the problems that you're working on. And the natural follow-up to that is cognitive surrender, which is where you, you know, you blindly give in to whatever the AI says as your answer.
I remember anytime I would work with a big site on their performance problems, uh you could easily spend half a day, um you know, just looking at traces before you've even written any fixes at all. And now that we have LLMs, it's very quick to like reason through massive stack traces and actually be able to get down to fixes you can make.
One of the things I found most exciting in the last couple of years was seeing as model qualities gotten better and and harnesses and tools have gotten better, how many people um that were directors or VPs or SVPs or any of these levels were actually rolling up their sleeves and trying things out.
The first is that new AI models have the know-how to escape containment in testing environments, which should challenge any assumption that advanced AI can be easily controlled by its makers.
An analysis by the AI Security Institute found that the capabilities of frontier models (i.e., from OpenAI and Anthropic) are roughly doubling every few months, and this doubling rate has actually gotten faster over time.
By 2027, almost every country-and every non-state group with sufficient computing power-will have the ability to download open-weight models even more powerful than the ones that attacked Hugging Face and use them for whatever they like.
AI may help an oil company emit less pollution per barrel of oil. But if the company is producing more barrels than it otherwise would have, total emissions can still rise.
I can't say enough how excellent a coding agent is at customizing Chrome with extensions. A huge productivity win. Every plugin interface you've ever blown off is now trivially usable.
We care very deeply about AI safety. We believe the entire field will have to coordinate on shared safety standards, but will act unilaterally in the meantime.
We expect confidence in safety to increasingly set the pace of AI progress.
have to abide by this rule that was written before LLMs were even a thing. And that shows how challenging sort of the the the situation is if you try to set the line too firm to begin with.
The thing about executive orders, which is a little secret, is when administration changes, you can just revoke the executive order and and start from scratch again. Um when something is a law, you you can't do that.
if the US wants to lead in artificial intelligence, we have to have a vibrant closed and open source ecosystem, and that's the only way they can all work together.
if you create a or allow for the creation of a patchwork of regulations, meaning there's one set of regulations for AI in California, another one in Maryland, another one in Texas. You know, the big tech guys, they can deal with that.
An excessively AI-polished proof may sand away both the "artificial" friction (typos, awkward phrasing, disorganization) and the "natural" friction, leaving a text that is easy to read and hard to learn from.
well but again these are functionally defined emotions and there's I think there's no problem assigning to AI to some extent and certainly to many animals uh those functional emotions
The basic answer is that I think of emotions as functional states of a particular type. They should be understood by what they do, their function rather than how they happen to be constituted in the human brain or differently in an octopus nervous system or in an AI which we should talk about.
The supercomputer inside our own skull runs on about 25 watts, which hints that the minimum energy needed for intelligence might be small - and that the giant, controversial AI compute centers we're building now may turn out to be a temporary blip rather than a permanent feature.
I'd bet we learn more about our own brains from building a thousand AIs under a real theory of intelligence than we've learned from a century of neuroscience.
The problem again was not reliance upon AI but that it was supporting a model of warfare with which the United States had been working for some decades, one that prioritized the rapid elimination of enemy capabilities.
If you keep asking this question, what you realize is energy is the fundamental input. When we figure out AI and robotics that allows us to do semi-aututonomous manufacturing, energy will become the cost of all things, right? The cost of buying a thing will become the cost of energy used to make it.
the score is the classifier's estimated probability for the AI-generated class based on its training distribution. However, we shouldn't interpreted it as a general probability that the text was written by AI.
AI checkers are essentially a cat-and-mouse game. AI checkers may learn to detect a certain pattern that is indicative of AI-generated content. Then, the next LLM may incidentally or deliberately not exhibit that pattern and avoid detection. The AI checker then has to be updated to detect said LLM, and so forth.
I still don't think it's a good idea though. The reason to have a cofounder is not just to get more done. It's to have someone to help bear the stress.
buuuut my roadmap is’nt ten times shorter. and I’m definitely not ten times better at deciding what is worth building. teams haven’t started casually shipping a year of product work every month. you can look around and and it’s hard to say what software has gotten meaningfully better in the last year.
Everything we see, whether it's this room or any room or anywhere you go, everything you see was either grown or mined, manufactured and moved. Everything. And so we look at what we do as industrial AI.
Chips come from the ground. Where's the energy come from? And a lot of people are like, "Oh, it comes from the sun." Yeah, it comes from the sun. But how are you capturing it from the sun? From stuff made from the ground, right?
For the first time, AI has "hands" - that is, the ability to reach out and interact with the tools on your computer and in your browser, just as you would.
I'm so incredibly tired of seeing low-effort AI-designed web pages. I basically bounce off any new product immediately when I see it (never get a chance to judge the product itself). Lots of thin lines, glowy styles, inconsistent fonts, lots of monospace.
One common question I get now is how much of the standard startup advice is still true in the AI era. So far almost all of it. There's more difference between the advice I give different startups in one batch than between what I've given pre-AI and post-AI startups.
It's like they have all the wisdom of their most senior engineers looking at every single diff and that is fantastic. Which means that you don't have to worry about remembering and looking and nitpicking and all the things that we're not good at anyway
You can use AI as a shortcut to help you not have to think too much. And you can use AI to help you think more deeply and more rigorously. And both of those use cases have their place. But when it comes to your core job function, we primarily want the second one, right?
But if you have to fine-tune a model, you actually aren't getting a general purpose model um for the things that you want it to do because you have to fine-tune it for each individual thing.
And in all of these applications, the customer is making a decision based off of the recommendation of the AI model more or less. Uh, and this means that if the customer is ultimately like kind of making the decision, this means that if the system makes a mistake, um, that's okay because usually the person can kind of recognize that or or decide what to do even despite that mistake.
Uh and this means that they're going to be far more useful when they're operating fully autonomously. And as a result, this requires us to develop physical AI systems that make far fewer mistakes than the machine learning systems that have been deployed thus far.
I mean at the very least I actually think that just starting with a generalist policy and then fine-tuning it even like right off the bat uh can be really effective.
The translation & dubbing was a huge amount of work (which I deeply care about) with humans & AI collaborating. Thanks to the amazing team at @ElevenLabs for their help.
There has also been an uptick in product shills, which, in some ways is worse than AI stuff-some human actually put effort into making straight garbage.
when AIs are extremely extremely capable my view is that those AIs will be harder to align than current systems. So for current systems, we have this feedback loop where we basically like we create an AI. We do some evaluations on it. We see that it has some kind of messed up behavior that we can kind of quickly understand. Then we like can like go look in training and be like, "Oh, the these training environments led to this problematic behavior. Let's like tweak that training data. Let's introduce some additional training data to like correct this other issue and then move forward from there." But in a regime where the AIs are extremely situationally aware, very very very very capable and um you know uh we don't necessarily understand what they're doing, this feedback loop breaks down.
I do I do think that I wish that sort of my preferred constitution or like the way I would orient towards this like the thing I would prefer would be more like Claude is like look it would be structurally good for the way this technology work like the constitution should be like it would be structurally good for the way this technology works to be that AIS are like good fiduciaries, good representatives, the equivalent of a lawyer for a user
the reason why RL environments today are much better than they were in like you know 2024 is not that much because um we have hired way more human experts to make RL environments and is instead much more because we better know what how RL like what RL environments we even want to make and and like how we should structure them and also we're using huge amounts of AI labor to build RL environments.
Second, I think ML is a very shallow domain relative to math. So I think in math there's much more of a you find some true deep abstraction um and then like that like if you really understand that thing which is hard to understand then you get somewhere
The fastest-growing political movement in America is dead-set on stopping the construction of data centers, which are critical to the growth of the fastest-growing industry in America.
Facts and Fallacies of Software Engineering by Robert L. Glass. In essence, this is a book about an industry that refuses to learn. That was true 25 years ago when this book was published, and it's probably twice as true today. (Just think about all the AI adoption metrics being rolled out — back to productivity mistaken for lines of code produced, only more elaborate. And expensive). What I like about this book is that Glass doesn't present anything new. Quite the opposite, actually. Rather, it's about research lessons that we all should know, but tend to forget. Ever had to do an estimate, or plan according to a requirements spec? Or maybe you thought that enough eyeballs make all bugs shallow? Then this book is for you. A great work by a fantastic author.
Paradigms of Artificial Intelligence Programming by Peter Norvig. Learning the AI described in this book probably won't land you a job today. But reading the code examples will transform how you think about source code. The book shines when it comes to code comments, a topic that I've never seen demonstrated well in other sources. Here we get to see how comments become valuable as a narrative that explains both intent and reasoning. Brilliant, just brilliant.
Claude Haiku is my current least favorite model - it hallucinates wildly, and is out-performed now by other similarly priced models like GPT-5.6-Luna
Even worse: it seems to still be used by the Claude Code WebFetch tool, which means hallucination risk any time you fetch a URL!
The absolute bad outcome is that our young generation, their agency and human level motivation of learning and living is taken away by tools. So doom scrolling, passive watching of shorts, all this are not helping agency, human agency.
They are the most important people in our society. We should be talking to them. We should be uplifting them. We should be supporting them. We should be providing resources to them.
Coding isn't yet another application domain -- it's the meta-skill required for AI to automatically develop its own training material, via symbolic world models. That's how the RSI loop actually kicks off.
In addition to reliability issues, it often engenders a mind-numbing workflow and an environment where junior developers will never acquire the expertise to become senior developers capable of designing complex systems.
But if you zoom out, it becomes clear that almost every “breakthrough” since last summer has concerned the narrow domains of computer code and math, which are defined by highly structured languages and come accompanied by massive amounts of specialized training data.
It seems pretty clear that the only thing data center opposition is doing is pushing data center construction to the (many) places in America that won't regulate energy or pollution. Weird thing to pat oneself on the back for.
His Lansing rally with Bernie and AOC, more than anything else, crystallized the connection between data center opposition and voters' anger about money in politics.
an ai is oddly well suited to this job. cassandra doesn’t want a promotion and doesn’t need the people in the channel to like her. disagreeing with the vp three times this month won’t show up in a performance review. she doesn’t have an ego invested in being right either.
In the era of base LLM scaling (2022-2024), I believed the LLM line of research would reach a capability plateau (as later seen with base LLMs).
In late 2024, after the o3 test-time compute demo, I changed my views: the new models were showing genuine fluid intelligence, and with this new line of work, the LLM line of research could achieve unbounded capability scaling. "There will be no wall."
However, looking ahead, I still do not believe that future AI (say, in 15 years) will be based on the LLM stack. I believe it will necessarily have to move closer to its optimal, final form -- symbolic learning. Obviously this is a risky and contrarian belief -- the safe bet would be LRMs. But let's see.
I believe current techniques are 4-6 orders of magnitude away from optimality in terms of data efficiency and test-time compute efficiency. But far future AI will be near-optimal.
Rather, the way that AI companies engage with communities-forcing NDAs, dangling billion-dollar promises, pushing environmental externalities far away from AI's wealthy user base-resembles a classic story about dark money in politics.
And that's why every hype cycle produces a wave of absolutely spectacular demos and very few real products. And the recurring mistake of every cycle is spending on the demo when you should be saving for the nines.
Given the high cost of errors, you need to have a very high level of safety and a very high level of confidence on day one before you deploy your first robot, before you drive your your first autonomous mile.
So, a deployment of your agent uh in the real world generates data. That data then grounds the simulator and makes it more realistic. The simulator generates harder edge cases for the critic to score and for the agent to learn from.
This is such a useful book. It makes the case that there's no such thing as "human error" - instead most catastrophes are caused by system issues, misaligned incentives, and unrealistic processes. You are not the custodian of an otherwise safe system that you need to protect from erratic human beings.
most of commerce in America isn't actually high intent. Like how many years are we into e-commerce now? Like 30 years into e-commerce and e-commerce has never exceeded 20% of retail spend in America.
absolutely agree that product skills are probably the most durable. My build would just be there's a lot of people who have the title PM who haven't spent a lot of time building those skills in the last 5 years, but have gotten really good at communicating frameworks to leadership.
as a product manager, I've spent less time in the last year talking to a data scientist than I ever have in my career, even though I've probably spent 10 times more time in data and understanding actually how the product's working than I have ever have in my career.
But based on what we can see at Stripe, the hunger and the intensity with which other companies are either getting started, taking advantage of these new capabilities, or existing companies are retooling, I don't worry about the centralization in the same way. Uh I think there I think there are going to be many thousands of winners.
And just human organizations are complicated and it's very hard to have um to manage to aggressively prosecute 100 different priorities and to deal with all the issues and interference that arises among them and so forth. And so, you know, Google has done incredibly well in a bunch of specific places, but it's not like Google has done all the things even if in some kind of basic material sense, uh Google maybe, you know, had that ability.
I mean, it's very interesting, right? Because these can prove the Jacobian conjecture, you know, whatever. Uh and so clearly they're capable of these monumental feats. Um but somehow I still haven't read the LLM essay that I found super compelling.
So, I um yeah, I think maybe maybe a better way of saying it is 20 20 years ago that whole lean startup thing was uh was almost the only thing to do because of capital available and you didn't have AI that made, I don't know, spinning up an organization with many different potentialities and capabilities so much easier, whereas now I think you can start these much more aggressive and ambitious things up front.
Yeah, I mean I feel like uh the models have been getting a lot better at sort of agent-based longer running coding tasks and it seems pretty clear that they are now actually pretty capable and depending on exactly your definition of of junior engineer it seems pretty spot-on I would say.
Um, and to give you an example of a a a use of a coding agent that works extremely well is you can ask today's models to translate software from one computer language to another very effectively because in that case you actually have a incredibly detailed specification.
Um, and sometimes that's because the model is trying to do something it doesn't have a lot of experience doing. So it's been trained on a whole set of things and as soon as you get a little bit off the distribution of things it knows how to do then like most machine learning models it will you know its performance will suddenly will start to degrade and the farther you get off the comfort zone of what it knows how to do the the more likely it is to to not work as well.
Yeah, I mean I think if you looked at what is important in AI systems these days, you would want to know things like the bandwidth between you know your main memory system on your accelerator to the onchip memory to the um you know the multiplier unit or whatever. You want to know how much energy does it take to do a single multiplier operation.
If you think about our large scale models today, they probably see a thousand times as much data as a human does by the age of 18. Yet, the human by the age of 18 is better in a lot of things and, you know, on par uh with those frontier models that have seen way more data. So could you come up with much more data efficient systems that can learn continuously learn from their own actions?
So, it's hard to tell how much of like the loss of the past few years was AI versus the end of like zero interest rate policy and like the post-COVID crash. And I think it's more the latter, but like again, LLMs are still getting better.
And often I found with my clients I have to tell them like it's doing a good job at generating the actual design, but in actually expressing what the design is supposed to do, it cannot do that yet. You have to do that part yourself.
As a general thing we've seen like to get good results you have to already know how to get good results without it. It just helps you get good results faster.
I predict that in the next 10 years software development will survive, but it will become like any other white-collar professional work. No more $200,000 salaries, unlimited vacation, or incredible employee bargaining power.
we don't believe in a world where these models are so expensive that, you know, they get rationed only for the most wealthy of developers and and companies.
we don't believe in this totalizing, you know, totalitarian view of, you know, AIs that control the world. We believe that these are going to enhance this very broad ecosystem.
so much of that debate is like I think um in some ways uh a little bit of a waste of time because, you know, I think it's inevitable that we're going to have very powerful models
we believe that everybody in the world, you know, all the billions of people in the world are going to have a super intelligence that is adapted and tailored to them, that is enables them to accomplish their goals, knows their context, and ultimately is an expander of their own agency.
I think we're at this like in amazing moment in the world where the bottleneck is not the progress of the AI models, the bottleneck is diffusing that through the rest of the world and and helping the world adapt to this amazing technology that already exists. Like I think if the models didn't improve at all from today, there would still be like decades and decades of like total upheaval and change in the economy and how the world operates and and everything around us
In fact, if if we are right that AI is going to be such a big change, startups will be much more important to making sure that the power of this technology gets widely distributed throughout the economy and society and is not just concentrated in a few companies or models.
At boom we need far more software engineers in a postAI world than we need in a pre-AI world. Why? Because the cost of software development has dropped. anybody including hardware engineers can now become a coder and we need software engineers to make sure the architectures are right and make sense and are coherent.
So I think the worst advice is like work on what you know. Uh what you know can be changed. Particularly in a world where you everyone has personalized AI tutors. Anybody with passion and dedication can learn new knowledge and learn new skills. But what you can't change easily is what you love.
I think the skill nowadays is less about prompt engineering and more about figuring out how do you give Claude a hard task that seems a little bit too hard. And then how do you make it possible for Claude to verify its work along the way? And the verification I think is probably the single most important thing that people do not get right
So when I look at engineers that have been, you know, coding for a long for a long time, you know, like for for years or for decades, this is a really really common failure mode is trying to over specify and it's trying to be overly specific and then, you know, get the model to do the to do the task exactly the way that you would have done it. And that that's just not the way the model works.
for people that aren't building agentic products, but you're using Claude code, every 6 months delete your Claude MD. Delete your skills. Delete your hooks.
I'm seeing a trend here of declining revenue and traffic with indiehackers On my own projects too Maybe big VC products too but I wouldn't know cause they don't share revenue To me it seems clear BigAI is cannibalizing everything that used to be apps
The evidence would show that and it makes perfect sense that AI and automation is creating jobs everywhere. The narrative about AI destroying jobs is exactly backwards. AI eliminate tasks. AI automates tasks away. But it doesn't necessary doesn't necessarily eliminate jobs.
we want we want to encourage everybody and every company to build their own AIs. And and and who knows what innovation will come from the fact that it's open source.
likely this will be one of the largest industries in the world and um, uh, it'll take longer than a couple two, three years. It'll take less than 10. And so this will this will be our next $100 billion business.
we we kind of have coarse level uh recursive self-improvement already. And the fact that every time you use it, it improves the markdown files. Uh every time you use it, it updates its uh long-term memory.
in an intellectually dead field, AI will just produce worthless junk. If what you feed it is decades of complicated and meaningless computations, it will just produce even more complicated and meaningless ones.
So by far the funniest thing about using Claude and Codex together is how Codex is very earnest and businesslike while Claude treats Codex like an inferior but hard-working subordinate
So by far the funniest thing about using Claude and Codex together is how Codex is very earnest and businesslike while Claude treats Codex like an inferior but hard-working subordinate
Google, which led on benchmarks not that long ago, has fallen behind where it now counts: it has no leading frontier model and it has nothing close to Codex and Code.
AI has the potential to assist with both the first and last element of that loop. What it cannot replace is the middle step—doing the work for yourself.
Asking an AI for help with the start of an essay, for instance, will invariably shape what direction you end up following, robbing you of the crucial experience of building your own judgement around a topic and selecting your own path to research and argue.
I genuinely believe that if you took an open weights model from 2025 and built a pentest harness for it, it could do this kind of sandbox escape and scan/hack in most networks. This is only surprising because you assume OpenAI has sounder sandboxes.
For the first time, I ran all potential winners through Pangram, which is an AI-writing detector.1 So according to the machine (and to my ear), all of the winners are 100% human-written.
Martsinovich argues, via code snippets, that there is no such thing as a "conversation" with a chatbot. The AI is born anew every time it speaks, and it simply reads the dialogue so far and then tries to write the next line:
I told him what I believe: that in the age of AI, medicine will be bottlenecked more than ever by regulation and the grind of clinical trials, and that billions poured into faster pre-clinical research won’t touch that problem, unsexy as it is.
I think the challenge is that everyone can now build apps But 1) almost nobody has distribution (like an audience), or 2) the money to pay for distribution (ads or UGC), or 3) the creative genius to get distribution for free (classically called guerilla marketing)
Their words now
I thought it'd be distribution but who knows, obviously creativity and ideas, but if you can copy a successful app in an hour, then how does that differentiating work?
So, even in a world of AI where some things are easier, we were talking earlier about mindset, AI fluency. From my experience, younger folks are more open-minded. They tend to be more native in some of these new ways of working.
It does feel like we have to be more more explicit about the types of people and talent that tend to thrive at Netflix versus other companies like some of the frontier labs.
So, the most useful thing is not to make it level specific or role specific, but to encourage everyone towards the expectation on AI fluency, which doesn't mean use it as a tech for the sake of tech. It's tech where it's useful, to have good judgment about that, and to have the mindset to be open-minded to explore and try new things.
So, we are hiring more people who can look across all the business domains and abstract that to here's the building blocks we're going to need in a world with AI.
In a world of AI with agents operating across multiple systems, wanting source of truth data, the importance of having preferred paved paths that get the most of the benefits and produce some guardrails so we can make sure we're doing good work, common infrastructure, common paved paths, solving problems once with a core set of capabilities becomes more important.
This saturation can be seen more clearly for the GPT 5.6 Sol model, which also shows that increasing reasoning budgets can become uneconomical at some point.
AI is developing at a very rapid pace, far more rapidly than most outside of the industry fully understand, and far more rapidly than almost anyone in the industry predicted.
the external tools that we use to help us reason, remember, associate, calculate-from counting on our fingers to chatting with an advanced LLM-should be properly understood as extensions of our own cognitive process.
the problem with training models on maintainability is like the cost function of bad architecture and bad program design can't be evaluated by running the unit test because it hits you 3 to 6 months later
you can slow way down and read every PR and read every line of code. Uh, and then you're only going to really get modest benefits from AI because that becomes I I think you should expect maybe 30 to 50% lift in productivity is kind of what I see when we go into teams
yes it will catch things and it will raise your floor but I don't believe like the model writing the code is the same model reading the code and if you ask a model hey is this code good it's going to be like oh yeah it's great comprehensive it's got unit tests
But if you don't have good LLM intuition, like 100K for smaller models, 200K for these like really beefy like Codeex and Opus 4.8 models is usually a good like training wheel guideline of like if you pass there, your quality of results may be degrading.
And at the end of the day, they're all like different ways to pass tokens into a model and ask it to produce usually some structured output. And understanding that is a lot more powerful than trying to learn memory and trying to pick some agent framework off the shelf and some memory framework off the shelf.
Geo-politically, countries will be weighing open weight models as a way to get frontier-level tokens inside controlled environments that may not be otherwise possible.
Organizations are increasingly looking for control over how their data is used and are willing to trade off some access to frontier level tokens for this control.
While in mathematics AI agents are starting to have a major positive impact on research (e.g. by resolving the Jacobian conjecture), the situation is very different in hep-th.
I think if you ask an AI just for a strategy lazily, you're not going to get something great. You're going to get something pretty predictable that probably the competition would expect you to do.
I believe that if you are too prescriptive as a leader with a team you end up stifling good ideas but if you're too open-ended sometimes teams just waste time um going in the wrong direction.
But I don't think we should judge content based on the tool that made it. Um I think we should judge it based on the content, the point of view, the person behind the content.
But they just by virtue of having less people to coordinate they can often move faster and make um better decisions a little bit less design by committee.
These developments inform our position that AI systems are now capable of meaningful contributions to formal mathematics, not merely informal problem-solving.
However, these LLM reasoners that generate informal reasoning in natural language are fundamentally limited by the lack of precise, machine-checkable semantics, making their outputs prone to hallucinations [Huang et al., 2025b] and precluding autonomous verification, a prerequisite for tackling open-ended mathematical research.
We argue that the next leap in AI4Math systems requires a decisive shift from predefined problem-solvers to research agents that can address frontier mathematical challenges with rigorous formal mathematical reasoning.
So one of the things is that the pace of development is definitely accelerated. One thing I wonder the pace of business hasn't accelerated though and that mismatch is going to become more and more apparent.
which is why when I see manifestos today I'm just like too soon. Not a bad idea. Would love to have one just too soon. It took 15 years for the technical change of object-oriented programming to come before we could say here are the consequences of it. Here's how in a simple way we can express how to effectively use this technology that we've been using day in and day out for 15 years. The genie comes along. People are like, "Well, what's the new manifesto? It's just not manifesto time yet.
the genie runs out of g runs itself out of options it can't make further forward progress and so I'll wipe it away start over I won't try and tweak I'll start over and say all right well if I implement things in a different order. If I implement with this markdown file or if I implement it with this commit hook, we collectively need to try absolutely everything.
Nobody knows now. That playbook has been wiped clean and people whose identity is I know the playbook are now terrified. Who who am I? Now, it turns out that the skill of writing a playbook is completely different than the skill of applying a playbook.
What is happening is that GenAI's climate harm comes from the system design, such as Google's AI overview triggering constantly whether users want it or not, in addition to significantly worse digital bloat tools dominating overconsumption.
That’s why we’re launching the Preliminary Report of the Independent International Scientific Panel on AI — an initial evidence-based assessment of the current state of AI science — to help the public and policymakers better understand the unprecedented moment we’re in.
I think we would always still prefer like a human that we had a relationship with because the way that we get motivated to be interesting interested in things is a social phenomenon.
I think the way that you'd measure conjecture generating ability is going to be more subjective on like that tone shift where um it'll be mathematicians saying they're not just using it to like solve their problems, but as they step back and decide what their research field should even be that a conversation with such and such model like was genuinely helpful for that.
but it would be a little bit disappointing and a little bit surprising if there weren't over the next 5 years like, uh, economically valuable improvements that were made that were directly like referable to the like AI progress in math.
that like if it's capable of building mountains uh that are, you know, the correct new theory that like crystallizes how we should be thinking about a subject, that's just such a level of intelligence that then it starts to feel like it would be surprising if that didn't permeate into other aspects of the economy besides like just the mountain building for math itself.
I believe that a Mac mini with Codex running on it is the best bang for your buck if you’re building an always-on agentic setup based on a frontier model that can do just about anything you throw at it.
This is why I recommend ignoring the vast web of custom and “community-built” MCP servers you can find out there, and focusing your attention on the two safer bets instead
If the user doesn't actively push back against the AI model, then they get the generic output, the lowest common acceptable denominator of aesthetics and taste.
this kind of instantly ubiquitous homogeneity is going to happen across every field of human endeavor in the era of AI, until AI gets so powerful that it actually becomes original (doubtful) or until we shut it down.
it’s also true that their impact on our output was never as tremendous or as clear-cut as that of prior innovations, such as the steam engine or the power loom, and that integrating them effectively into the workforce took a long time
elevators are the rare part of the economy where Europeans have embraced market dynamism while Americans choose overregulation and labor market rigidity.
untrusted repos should be treated as hostile by default because they can steer the agent toward reading files, running commands, or sending data through approved tools.
it's not that humans have become smarter over it's not that they evolved to become smarter over, you know, the past 50,000 years. It's that humans are able to do a lot more today than they were back in caveman times because there have been billions of humans thinking for a long time and building off of each other's accumulated knowledge.
But what we're seeing today with the modern models is that 5.5 and other models can think for if you scaffold them reasonably well, can think for weeks even um before having performance plateau on some of these benchmarks. And so, the point at which they plateau is simply too far out to reasonably test.
my claim is the proper way to evaluate the models now is you either have some kind of budget for the benchmark whether it's tokens or cost or time or whatever or you plot the performance as a function of the amount of test time compute that's going into the model
I think the reason why it doesn't show up as so much better on the benchmarks is because the benchmarks are being presented, the benchmark results are being presented in the wrong way. They're not controlling for the amount of test time compute that is being used on that benchmark question.
And so you kind of end up in this this bad equilibrium where everybody kind of knows that it's a bad equilibrium, but like nobody wants to break out. And I I felt like, okay, well, if I just hopefully come out and say like, look guys, let's all recognize that we're in a bad equilibrium and let's move to this different equilibrium where we're we're plotting things with an X-axis
And even the cost issue with AI is probably going to be like once the subsidies start running out, which we're starting to see, I think that's going to be a really big issue where maybe all these companies that embraced AI programming are now going to like cut back on it.
I think a lot of people, especially students, are unfortunately learning everything through LLMs. So a lot of that isn't really learning, they're just kind of cheating and they're just doing everything like that. And then they lose a lot of their skills
But in terms of like features, like a website like you can just throw features in there nowadays that nobody really cares about and you can you can do it so quickly. Like a new feature every single day, but do people actually care about that? Is that making it better? It could be making things worse.
I think a lot of the things that people might say that like oh, it was a waste of time to learn this subject cuz I didn't actually use like those details on the job. I think that's a very wrong way to think about it. And I think that's what a lot of people are doing now with AI. Like hey, what what if I'm not going to be writing for a loops a couple years from now. Um I don't think those things are a waste of time.
I think Google has pretty much gone to on-sites at this point, back to the traditional whiteboard format. And they'll let you code on a laptop if you want to, as well, but it's going to be in person. Somebody's going to be watching you code, and you're probably not going to be able to cheat your way through that.
although students' personal statements seem more creative because they use more varied words, they actually feature less original ideas. AI writing produces an illusion of creativity.
Generative AI has no internal, designed momentum towards truth and accuracy (beyond the absurdly diminishing returns of energy-hungry multi-layered LLMs), but a person or an institution can (and should).
So you know, just like internet, you can see some of them turn out to be very big like Amazon, like the Netflix and then some of them is kind of go sideways and disappeared or being acquired. And so I think to me is the same approach.
You know, right now there's a massive build-up in term of the AI, you know, the I think it's the right thing to do. I don't see that in anything to slow it down uh because the workload is increasing a lot.
One is of course everybody knows power constraint. Some country the power they just don't have that. They get impacted. And then secondly, a lot of people didn't realize the helium impact can be also very significant for semiconductor. And then the thirdly, is everybody know right now memory is a bigger shortage.
In the English-language arena, Britten sets the standard for handling writers of inborn musical power-the likes of Shakespeare, Donne, Blake, Keats, Hopkins.
the council does not simply keep the best bits from everyone. It keeps a minority of the good ideas, while peer review seems to give consensus ideas an extra push.
the way we should be using AI is as a testing machine, a failure machine, and a way to vibe code, cloud code, but but build build the, you know, the the lowest possible cycled version of your product that you can get signal back on.
there there's a one question we got to back up and really explore which is is AI a new platform and I would argue that it is not yet a new platform. It is an important technology.
when you have an exponentially growing curve, I think the way that an exponential curve feels is it's growing so quickly that the the kind of emotional feeling is it can't possibly keep going, right?
We don't believe in this like very centralized future where there should be a small number of institutions that um that basically are are advancing all this stuff. Our vision is not that there's going to be like some central super intelligence that solves all of science.
in order to make progress in AI you don't need like many many hundreds of AI researchers um or thousands or anything like that I think you can really make progress with um you know a very strong group of a dozen or a couple dozen people.
I think like people are really important and I think we'll be more important in the future and giving people more tools to be more productive is going to be like a critical part of any kind of positive future
I believe that the path to success is best achieved by putting the best human intelligence together with the best artificial intelligence and that the way of thinking I am describing here is essential to understand and use in the new human/artificial intelligence era.
I know through my experiences that even the most advanced artificial intelligences don’t have adequate enough insights to allow one to blindly follow them and that unique human understanding and insights are still invaluable, and that that is especially true in investing where value-added is a zero-sum game (so that, when it comes to adding value, what is widely known is of little value.)
To be clear, these criteria are not best derived by looking at what would have worked in the past and assuming that it will work in the future—i.e., data mining—or simply asking an AI what to do. They are based on logical understandings converted into decision-making systems.
What this process can now do in creating understanding of the timeless and universal cause:effect relationship and enhancing and systemizing whatever one is thinking is mind-blowing. I believe that you will either stay at the cutting edge of doing this or you will be uncompetitive.
This is a super exciting release - Claude Fable 5 is the same underlying model as Mythos but with added safeguards. The benchmarks are great and it's SOTA on everything by a margin but I'll add that *qualitatively* also, this is a major-version-bump-deserving step change forward
the regulation gets extremely difficult and mundane and expensive that could actually lead to more igopoly and I think some of the players know that and are begging for regulation.
We do live in a world where information is really cut up, but we also live in a world where you can have access to more information than you ever could. And that's even more true now with LLMs.
we're still going to need a display cuz sorry people, unless we're plugging it into our brain like a BCI brain computer or there's some laser thing going into our retina, we're going to need a display.
What's proven is if you properly architect it and have co- cloud code go into certain sub segments or have cloud help you build the architecture, you modify, refine it, lock it in and say just work on these few things
Like right now we're endowed with labor that can turn into uh that can turn into income. When that is no longer the case and we are now at the mercy of the of the elected official for like basic needs, right? So that to me feels like a power sharing arrangement that's really dangerous.
They just recently released a report, and I think like you really have to squint to see anything happening. Like basically, if you want to take kind of like uh an an approach across the entire economy and looking at even looking at like software engineering, like the most exposed sort of sectors, there's just like not really anything going on. There might be a little bit of a signal about like junior developers getting jobs less than before, and that but that's like a less than before rather than a level shift.
So, I think there is a world where it is concentrated, in which case it's going to be really hard to index AGI. There is another world where it is not It's electricity, then like basically every company has access to AGI. So, you just buy you use buy the index. So, like, you know, Nigeria just needs to buy the index.
rather than thinking about individual forecasts like what me and Phil are going to do, rather looking at kind of like basically generating prediction markets, where you get aggregate forecasts, where you get like kind of wisdom of the crowd effects. And kind of the reason that I think this is because we have been famously terrible at forecasting.
we don't have any data. I've been kind of saying we need a Manhattan Project for data. We don't have data on basically consumer demand elasticities. We don't know what they are.
things have to go really wrong for us to like just get over the threshold of uh you know, capital being productive enough to automate lots of work, but not be productive enough that that the interest rate is high and or the price of capital produced goods is falling a lot, okay? So, even without redistribution, a little bit of savings will save a lot of people.
some people think either uh frontier AI gets commoditized and we all enjoy the benefits, but there might be some risk because like it's the market's really competitive and cutthroat, or um things are safer because there's a big gap between the leader and the laggard, but that means that the leaders get fantastically wealthy. No, like you could just have a relatively big gap, but it's a public company ownership and it's widely distributed.
if I had to guess I would guess that the kind of long kind of general trend of just like lowering those frictions and making it easier for more and more people to index more and more will continue despite the recent bump in the other direction.
it's already not that hard to index. So it's not There's been a bit of an increase in the privatization of returns but it's still like you know well under 20% of the total market cap of um non-non-tiny companies in in the US is is a private.
here prices are adjusting in this interesting way that too many macro models don't allow for, right? So, that what what a what's happening is what would be called investment specific technical change where yeah, the price of capital is like falling relative to the price of consumption instead of like the standard doing the standard macro thing of saying there's just output.
there's a consistent theme where Gen Z (and millennials, to some extent) are rejecting the AI hype while older generations are optimistic and coincidentally in a position to benefit from it.
So I'm looking at this like guys the fundamentals of this isn't like everything is going to change because of MCP. What the hell? A whole conference guys? This is just API design.
If you all pivot to AI, then you all will have a problem. It's like kids learning to play soccer. They all run to the ball. No strategy. Spread out, figure where you add value and play your position.
So you have a lot of smart Googlers. These people are brilliant. I mean extremely brilliant to the point where the hardest problem Google had in my opinion was what to build. Not how to build it, what to build.
do not say AI because what we don't want to do is use a big umbrella to describe what you're doing. Let's get concrete details. These are computers. These are computer programs.
Well, I love these markets and they're all around us where we think they're mature and over and they haven't even started yet. And that was search before Google. You know, Google was the 56th search engine. It was a mature, slow growth business. Google made us reimagine what search could be in our lives and obviously turned it into a trillion dollar company value.
LLMs themselves are the perfect captive reader: they never get bored or confused, never miss a reference or have an emotional reaction, never ask themselves why am I reading this?, never close the tab.
I truly believe that following this kind of approach to AI in the classroom would, in the aggregate, produce better learning outcomes than the research and writing workflows available to students pre-AI.
To a certain extent, the concerns about cognitive offloading are an example of technological lag, where the mainstream discussion of AI and its impact is still framed in terms of the first-generation chatbots.
I believe that one dimension of this sadness or demoralization can be attributed to the simple fact that we are increasingly invited to outsource a class of activities that grant us a measure of satisfaction, accomplishment, and purpose.
My working thesis about the generalized impact of “AI” as it is currently deployed can be summed up in the observation that the arc of AI bends toward demoralization.
If AI-driven labor displacement ends up being large in magnitude and permanently drives down the demand for labor, it will likely be necessary to go beyond mere incentive programs to long-term income support for a significant fraction of the labor force.
I had already put both laptops through my benchmark gauntlet, which revealed one theme: the Mac is faster (in most cases), more efficient, quieter, built better, has a much nicer display, and costs much less.
The globalization of generative AI models and tools has caused a kind of leveling, as any aspiring creator anywhere can work in the same visual idioms.
So it's no wonder that the artists and designers whose livelihoods are threatened by it have begun moving aggressively in the opposite direction, toward the scrawled, the sloppy, the seemingly mistake-riddled.
the joke I made was pre AI, I would spend 95% of my energy thinking about what to do and 5% of my energy doing it. Now I spend 96% of my time thinking about what to do and 4% of my time actually doing it. So yeah, it's like a 20% improvement, but dayto-day it feels as hard as ever.
And our harness wasn't very good for like the first like five months of open code. But it was good enough. It was good enough that most people couldn't really tell a difference. And once we won enough share, then we went back and like tried to make our harness like good and smart and optimize and all those things. But uh it was inverted from what everybody else was doing. Everybody else was being like you have to build the smartest harness and that's how you win.
It's like everyone is just saying mantras to themselves cuz the root thing that's going on here is we are experiencing a moment of great change. Everyone is very nervous about what that means for their own position in it. a defense mechanism is to confidently assert a future in which you're a winner. And that's almost what every single prediction that you see is happening.
there is a world where the net result of all these AI coding tools is the same amount of work gets done, but all the engineers are happier because their job is easier. That's not good enough for a lot of companies. So, they're just going to say go back to typing out the code.
But I've always thought it's better to think a lot instead of swinging a lot. I think you could you can eliminate a lot of ideas or directions just by, you know, spending a lot of time in your head and with your team talking. Um, obviously AI doesn't speed that part up.
I think it’s good that the system identified the mechanisms that it did, but I don’t think that the paper does enough to point out that none of these ideas are without precedent - in some cases, a lot of precedent.
And so I think it's it's really important uh when when we think about benchmark progress to think about it from that perspective, which is benchmarks rise on problems that we've framed that we can articulate, that we can score. And there's a lot of work that's human work that uh it it can't be scored until you write it down
I think people think of the edge of AI as being in San Francisco. And I actually don't think that that's where it is. I think the edge of AI is wherever AI meets like a real human doing something.
I think the same kind of principle applies with agents in that they can talk to the compiler. It will tell them what to fix. So I guess this could be a case. We we'll see. But Rust could be a pretty promising candidate for to use for agents because they can get more feedback and it's just hard it's harder to to ship certain type of bugs or maybe impossible to have certain type of bugs.
In a sense, we're all turning into project managers, right? And and we can have an army of junior programmers called agents that will just spit out reams of code, but someone's got to have the big picture and review all of that. And so, increasingly our craft is going from one of writing the code to one of of reviewing the code and and building the architecture of the code and overseeing the work, if if you will.
but but if you have good locality where you you clearly stating what you're importing and whatever and and you can analyze just a single source file and from that extract its protocol to the outside world without having to to know anything deeper. Do you do you know what I mean? I think those are important aspects just simply to reduce the size of the of the of the context window and also make it easier to summarize each module in a program, right?
AI today it's just starting to to become aware of you know the existence of of of language services and agents today like to use grep and awk and whatever you know to to find all the places where you reference a certain thing but it's not semantic search, right?
because if you were to force AI to write a type annotation on everything, then it would probably get it wrong more often because now it has to keep track of all these types and and it and it has to just repeat itself over and over and over, right? And so, types are important where there's no context.
If you say that shareholders should run it, then you're saying literally whoever can borrow the most money should be able to control this technology, it's just that's nuts.
I think the structure frankly is better than founder control. Daio does not have dual class shares the way that Mark Zuckerberg or or Larry and Sergey had.
This is the number one unsolved problem in AI. It's not the tech. We're making great progress on the technical alignment problem. But we haven't made jack progress on the human alignment problem
At even $500/kg, launch cost is only 5% of the total satellite deployment cost, so a lunar mass driver is unlikely to drastically improve the economics of space-based AI, by reducing launch costs.
the reviewer is the principal, the contributor is the agent, and code review only worked because the reviewer could cheaply infer effort from reading the code. Agents collapse that signal.
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.