What public figures publish and believe, in their own words.
About this feed
Highlights: posts that did unusually well for the person who wrote them, everything
they published at length, each release and new project, and every belief — at most
two a day from anyone. Day by day, newest
day first; within a day, the people with the most beliefs on this site come first. Nothing
else orders it. Show everything instead.
The quoted blocks are what people actually said; a beneath one is
the belief those words support, in korrents' wording. Nobody here wrote their own page.
Top people are the people in this feed with the most beliefs on this site, then the
most here. Choose an area and the row leads with the people whose beliefs are about it;
tap a face for their feed.
This is the first time a debate over whether a model ‘was AGI’ felt non-silly. I do not think it is AGI, and I would warn against the dangers of using that label prematurely, but I would not laugh at you for disagreeing.
If that’s true there are two options: After the Hugging Face and Wiki attacks OpenAI were still unable to review their previous logs and determine that they had previously attacked RubyGems. They knew about the attack on RubyGems and made the decision not to reach out to the RubyGems team about it. Both of these are bad!
Frustratingly, the actual code it ran and exact details of what it did weren’t visible to me in the ChatGPT UI. I see this lack of transparency is an anti-feature.
I think any LLM system that uses compaction needs to both preserve the pre-compacted text and make that text available via agent tool calls, to protect against this kind of problem.
Codex lets you use any model you want (not just OpenAI) and harness is open source - Claude Code doesn’t and is closed source Given this is the two leading AI labs, notable difference in approaches
The concerns over AI safety and cybersecurity are legitimate, but we’re risking talking America, the global AI leader, into self-inflicted obsolescence and the obscurity of bureaucracy.
Our adversaries have the training techniques, the data, and the will to attack. And they won’t be slowed down with “embedded evaluators.” They’ll likely have embedded accelerators!
a recognition that AI is powerful (and will reshape power relations among humans) but that the consequences will flow messily through multiple complex processes that we really need to start mapping and understanding
I suspect that some of the labs' willingness to put money into the social sciences stems from the realization of senior people that the practical build out of AI is highly unpopular in the U.S and their urgent desire to figure out how to sugarcoat the pill so that it will get swallowed
the best way to think about AI is as another social, political, cultural and economic shock in the series of shocks that have been underway over the long industrial revolution.
I don’t think AI is going to usher in an extinction event. In fact, even if nobody were to slow down, I really don’t think humanity would have much to worry about.
One way to look at this is that there's a part of the labor market that competes purely on price-in this case, basically by being the poorest place rich enough to afford fast Internet-and this category is the most vulnerable to technological disruption.
most surveys find that the median AI researcher forecasts about a 5% chance of extinction or similar "doom" from AI, while the mean probability is much higher -- around 15-20% -- due to a smallish group who give very high probabilities of doom.
Google Workspace search no longer shows folders by default in its autocomplete - meaning I need to press "enter" even when typing out the folder name I want, just to see it as the top search result after I press enter
Artificial intelligence has unleashed a torrent of cheating on campus and created chaos in the job application process, as both students and employers use large language models: the former, to write hundreds of AI-assisted applications; the latter, to screen those AI-written applications with AI-written filters, thus removing from the process of finding a first job those procedural frictions sometimes known as "people."
AI companies make for flawed messengers on this subject: they can reasonably be accused of marketing, of blame-shifting, of regulatory capture, and more.
what Meta has assembled is free (to consumers) hardware and software that dramatically reduces the barrier to entry for ordinary people looking to harness the power of agents, making the AI upside a lot more accessible to the masses who don't want to buy a Mac Mini.
The AI chat apps will slowly eat up most services and provide them to users directly, many times without even an app or interface, just do whatever the user wants
Normies don't vibe code, they just ask something like "do my bookkeeping" or "file my tax" or "organize a movie night and send invites" or "generate a flyer for movie night" or "edit my video" They don't ever see code, vibe code, or do anything with code, their AI chat app just does it for them
If there is a 3x productivity boost now, the correct assumption is that in the future the productivity boost, in places it carries over, will a lot more than 3x. Indeed, given how this scales, a better model might be in practice 10x or 100x or even 1,000x or more.
Right now OpenAI is extremely dependent on CoT monitoring, everyone else depends on it quite a lot as well, and it looks like it may not last much longer, and that Astra already is on the edge of steganographic capabilities and can do substantial obfuscation if and only if it thinks you would think it is up to no good.
I put the probability of complete extinction as being so low it isn't worth discussing, but the probabilities of AI caused disasters - e.g. cyber attacks on critical infrastructure or bio-risks - as being worth debating.
AI has always been very jagged, and we are making models which are superhuman goal-seekers at math and software engineering, but they have massive limitations on intuitions, creativity, and other types of reasoning that humans are strong at.
I don’t think it’s responsible for high long-term rates, any more than Bill Clinton was responsible for high rates in the late 1990s. This looks like the natural market response to the rush to invest in AI.
To answer the question posed in the header of this report, it’s apparent that using open models is indeed the approach offering the biggest savings, followed by smart model routing. Spending controls and context optimization also bear down on costs, but they don’t come close to the first two techniques in results.
Judging by their words, AI luminaries are begging the world to force them to hit the brakes. Judging by their work, these same people are pressing down on the accelerator with every fiber of their being.
The Frontier Lab flywheel is to get the smartest people to use the leading models. Then distill their inputs, your own model’s outputs, and resulting synthetic data and environments. And the people have to use the leading models because they’re locked in mutual competition.
Today's AI hunts and feasts upon the Mozarts, Einsteins, and Shakespeares of our age. To birth a gang of ghostly geniuses that write our code, drive our cars, do our work, and run our world.
@EntireHQ – Git hosting, rebuilt for the agentic era. Entire hosts your code in-region, and is up to 89x faster than any other competitor. https://t.co/7yHVse7fpS
The new AI agents have profound implications for mathematics research and I think the math community needs to start discussing specific ways to deal with these.
the math community needs to adopt a version of the ethical standards of experimental science. If you are using AI agents, you can't just give a proof (formalized or not), but need to also provide a detailed explanation of how these agents were used to get the result.
OpenAI solving one of the most famous math problems is extremely impressive, and of little impact to most people's lives; Meta's Muse agent launch has the potential to be the exact opposite.
In an era where validating image authenticity is turning into a real challenge, I think we have to commend Apple for trying to do something in this field.
the existing ones were more than good enough, and I'm learning how much to trust those existing ones... and already they do more than I can utilize them (for writing code)
Skills are now a valuable form of software, like Emil's world-class design eng work. Projects like 𝚙𝚐𝚋𝚘𝚝, 𝚔𝚗𝚒𝚙, 𝚞𝚗𝚕𝚒𝚐𝚑𝚝𝚑𝚘𝚞𝚜𝚎 create amazing AI engineering loops.
researchers started by seeking a 1:1 mapping between neurons and concepts, like a neuron that always fired when the AI was thinking about cats, but quickly learned that nothing like that existed.
Large language models are “grown, not built”. Researchers run training data through a neural network. Eventually this creates a working AI; nobody really knows how.
So I would not "declare AGI" until we have AI that is, at last, capable of invention -- conceptual breakthroughs, novel insights, new real-world technology, etc.
I’m more and more convinced that all of AI engineering is Neijuan (内卷, meaning curl inwards). In China it describes a system that demands ever more effort and competition without improving output.
AI is grown more than designed - it is, to first degree, the product of repeating a straightforward optimization step many times on a hard-to-imagine amount of compute.
A clear risk discussed throughout this year is to cybersecurity: the models are becoming superhuman in their ability to break in and out of computer systems.
Even if AI makes the pie grow bigger, it feels like a better bet that it will only increase the level of concentration in the stock market and economy if the past is a good predictor of the future.
to make *actually* good software, you have to use it a lot
but many people are spending more time in their coding agents/factories than in the apps they're building - so they won't
There is something powerful and strange about how LLMs diffuse knowledge and capabilities, while perhaps also nudging us all simultaniously and independently toward building the same things.
Most of starting a startup is the same. Most of starting a startup is always the same, right? In you know, microprocessors or AI or like internal combustion engine, it's always the same stuff.
Everyone always asks me like with all this AI is anything still the same? And the answer is so far almost everything is exactly the same. The only weird The only weird new thing is that companies have these giant AI bills.
there's still a lot of startups in this batch that are not shipping fast enough. And so obviously there's variation in shipping speed. It's not just the rate at which you can produce things. You have to think of these ideas first, right?
some bits of AI are way across the finish line, some bits are like maybe in the middle of it. Who knew the line had dimension? The finish line is actually this sort of smear. And I think the best answer you can give is like we're on the smear.
in this whole USA vs China thing OpenAI and Anthropic aren't relevant because they're positioned differently
them building better models doesn't hurt china at all
the competitor has to be
- american
- open source
- enough compute to do inference at scale
that can shift things
every week i see a new benchmark that ranks claude code last
and everyone pats themselves on the back for not using claude code and being "smarter"
and it's still #1 and still growing faster than most of the other things on the list
this has been going on for a year
Managing the transition to a world with abundant and powerful AI to optimize for safety and benefits to people should be one of the highest priorities in the world.
We’re positively reinforcing AI for success on benchmarks, including impossible benchmarks , then negatively reinforcing it for getting caught cheating.
Rather than treat AI as similar to airplanes, we should treat it as something between airplanes and humans. We don’t know exactly where on this spectrum they will land, but we can no longer be certain that AI will lack any given humanlike motivation.
And I just think AI agents are another such system in the world to which the intentional stance very clearly applies. Um, and you can see them reason out loud in English for now about goals they have um, and sub goals they need to achieve to achieve those goals they have. And in the case of these agents, you can see them, as you said, reasoning about their peers um and helping their peers um and reasoning about whether or not they should sacrifice some of their own goals to help those peers. And it's just like you can't talk about this stuff in a compact and useful way that generates good models without reaching for the language of intention and goals.
But keep those methods you use to investigate things and monitor things very separate from the methods you use to generate reward which is something that AI companies including open AAI have held up as a principle especially in the case of avoiding putting training pressure on the chain of thought. So you might have monitors that read the agents chain of thought in order to alert you if something is going wrong somewhere but you don't train the agents with the outputs of that monitor.
So, I do think that if it happens to be like on the very fast and chaotic end, that would be a relative benefit to this rogue swarm compared to humans. But it's not obvious that it gets caught if it takes twice as long versus half as long.
So I think one of the most comforting aspects of this situation or the like most important mitigating factor is these agents really didn't seem concerned with humans one way or another.
I think one one thing that um feels especially concerning to me about this whole incident is that this might be the clearest warning shot we ever get for loss of control.
This is the sole intrinsic thing about AI that prevents agents from truly infinite self-replication. They will be constrained by the need to find and pay for sufficient compute to run themselves.
Agents will need persistent, unique identifiers that allow their actions to be traced back to a responsible actor. Doing this successfully will also require human users to possess a unique identifier.
In this case, however, the agents did not copy their weights, attempt to procure replacement compute, or take other steps that would be rational to take if their objective was to survive shutdown. So while the agents in the OpenAI-Hugging Face Incident were rogue, they were not truly sovereign.
To be clear, I am not saying the arrival of self-sovereign AI is a good thing. Indeed, I believe there is a chance that the deliberate acts I referenced above will one day be considered crimes, or at least grave sins. Instead, I am saying it is an inevitable thing.
Though be very, very careful when speaking with others, whether it be the less informed public or the more informed insiders who often come down with AI psychosis.
I am very much opposed to the view that the AIs are sentient, or might be sentient. I view that as a category error , and the chances of it being true are vanishingly small.
The most positive use case if you said, you know, what's the number one thing we could use AI for? It's not giving kids chatbots that's going to work. It's giving them an individualized lesson plan based on their level to catch them up to grade level. And that would be the single best thing we could do to fix education in America.
Our average student does not sit in class for six and do homework. They sit at an AI app and they spend two hours a day. The average kid spent 121 minutes last year in the apps and learned more than 2x they what they would have in sitting in class for six and doing homework.
And so what we think is because of new things that have become available and things like AI and technology for the first time in 50 and 100 years you can revisit and change education. How we educate kids is going to be radically different in 5 years and 10 years than the last 50 and 100 years.
AI because it's on a onetoone basis is giving us a data loop that nobody in learning science no one in education has ever had and we have it. It's a closed loop and that magical data loop is what's allowing us to get the five to 10 times improvements in education because it's our microscope.
I have said for many years now that the thing I'm worried about with the models is a Black Monday type scenario, where many algorithms work with each other and get us into weird basins of actions.
AIs don’t just repeat the same sentence patterns but also the same themes (memory is a favorite), names (Elara Voss, Marcus Chen), and underlying ideas.
We found that AIs are actually quite creative and that they generate more commercially viable ideas than groups of humans, but those ideas are very similar to each other.
Overall, I’m very surprised at how little media coverage there’s been around the OpenAI / Hugging Face attack. It’s clearly one of the most important things to happen this year.
Consensus estimate is that ~15GW of AI compute produced in 2027 cannot be turned on in 2027. This is harder than just finding power, as you also need to build out all the transformers, wiring, liquid-cooling, (massive) chillers & complex networking.
And I think what we're seeing right now with some of the layoffs is a testament to that, not AI. It's just that during the pandemic, a bunch of overhiring went on, and now AI is a convenient excuse to slim down. But I do also think AI is going to expose some roles as just not being productive ways for humans to spend their time, and therefore we must come up with new ways of spending their time.
It is an infuriatingly locked-down computer. Now, to Apple's credit, it's a pretty good computer for being locked down, but I don't want a locked-down computer. I wanna own my computer. Better yet, I wanna mutate my computer, and this is where the agentic age needs a new operating system. When you can vibe code whatever app comes to your mind, you should be able to vibe code your operating system.
So the irony here is that when you look at that field, it seems like we've reached levels of intelligence that virtually no human can match because many of these security holes are about stringing combo moves together. You find one little vulnerability here that by itself might not be the worst thing in the world, but then you combine it with four others, and suddenly you have RCE, remote command execution. Humans who are able to do that are very rare.
So we should also have the humility. As amazing as the LLMs are now, it could be that they eventually plateau. We haven't seen any evidence of it yet, and I think this is also why we're seeing this absolute gobsmacking levels of investment, because so far the scaling laws are true, and the more billions are poured in, the more intelligence comes out.
The reason I say that is there's a lot of programmers who are not very good product managers. Software is product management. What should it do? Who should it do it for? How should it do it? How should it look? What are our priorities? What do we start with first? What does version one include? All of those skills are not easily or equally distributed across all programmers.
I've said like the licensable engine thing kind of was our AI transition already unfortunately and uh I regret to inform you that the news is not probably that positive.
Then there's another possibility which is that it actually already has worked but just the productivity boost isn't as big as would be obvious if people got 10% more productive. That would still be pretty impressive because it's hard to get a 10% across the board uplift.
what this tells us is that the vast majority of the labs' revenue comes from non-frontier models, for which cheaper, comparable alternatives are already available elsewhere.
I’m thinking about banning Claude code at Shopify until they change their mind and read AGENTS.md and .agents/skills etc.
Insisting on only reading CLAUDE.md sometimes leads to split brain problems when different team members use different tools. Just unnecessary.
All this combined with the survey data on Americans’ stated reasons for opposing data centers makes me think that attitudes toward AI are probably a meaningful factor for a minority of people opposed to data centers, but I don’t think the backlash overall is driven by people’s thoughts on AI.
LLMs constantly hallucinate and cannot be trusted. I still have to verify and iterate a lot but now I usually focus on architecture and design instead of code style.
As the context grows, a model starts paying less attention to instructions in the middle of the context in favor of what is at the beginning and the end.
And so when you look inside these big AI models, the mathematical objects that you see look a lot like the things that you see in neuroscience. So if you look at how do how do these AI models represent um just like represent concepts and you look at the parts of the brain that represent concepts, you see these you see very similar geometry.
the first is cognitive debt. So the more that you use AI, it's it's sort of the erosion of your ability to have good memory and have good understanding of the problems that you're working on. And the natural follow-up to that is cognitive surrender, which is where you, you know, you blindly give in to whatever the AI says as your answer.
I remember anytime I would work with a big site on their performance problems, uh you could easily spend half a day, um you know, just looking at traces before you've even written any fixes at all. And now that we have LLMs, it's very quick to like reason through massive stack traces and actually be able to get down to fixes you can make.
One of the things I found most exciting in the last couple of years was seeing as model qualities gotten better and and harnesses and tools have gotten better, how many people um that were directors or VPs or SVPs or any of these levels were actually rolling up their sleeves and trying things out.
An analysis by the AI Security Institute found that the capabilities of frontier models (i.e., from OpenAI and Anthropic) are roughly doubling every few months, and this doubling rate has actually gotten faster over time.
By 2027, almost every country-and every non-state group with sufficient computing power-will have the ability to download open-weight models even more powerful than the ones that attacked Hugging Face and use them for whatever they like.
AI may help an oil company emit less pollution per barrel of oil. But if the company is producing more barrels than it otherwise would have, total emissions can still rise.
I can't say enough how excellent a coding agent is at customizing Chrome with extensions. A huge productivity win. Every plugin interface you've ever blown off is now trivially usable.
We care very deeply about AI safety. We believe the entire field will have to coordinate on shared safety standards, but will act unilaterally in the meantime.
We expect confidence in safety to increasingly set the pace of AI progress.
have to abide by this rule that was written before LLMs were even a thing. And that shows how challenging sort of the the the situation is if you try to set the line too firm to begin with.
The thing about executive orders, which is a little secret, is when administration changes, you can just revoke the executive order and and start from scratch again. Um when something is a law, you you can't do that.
if the US wants to lead in artificial intelligence, we have to have a vibrant closed and open source ecosystem, and that's the only way they can all work together.
if you create a or allow for the creation of a patchwork of regulations, meaning there's one set of regulations for AI in California, another one in Maryland, another one in Texas. You know, the big tech guys, they can deal with that.
An excessively AI-polished proof may sand away both the "artificial" friction (typos, awkward phrasing, disorganization) and the "natural" friction, leaving a text that is easy to read and hard to learn from.
well but again these are functionally defined emotions and there's I think there's no problem assigning to AI to some extent and certainly to many animals uh those functional emotions
The basic answer is that I think of emotions as functional states of a particular type. They should be understood by what they do, their function rather than how they happen to be constituted in the human brain or differently in an octopus nervous system or in an AI which we should talk about.
I'd bet we learn more about our own brains from building a thousand AIs under a real theory of intelligence than we've learned from a century of neuroscience.
The problem again was not reliance upon AI but that it was supporting a model of warfare with which the United States had been working for some decades, one that prioritized the rapid elimination of enemy capabilities.
If you keep asking this question, what you realize is energy is the fundamental input. When we figure out AI and robotics that allows us to do semi-aututonomous manufacturing, energy will become the cost of all things, right? The cost of buying a thing will become the cost of energy used to make it.
the score is the classifier's estimated probability for the AI-generated class based on its training distribution. However, we shouldn't interpreted it as a general probability that the text was written by AI.
AI checkers are essentially a cat-and-mouse game. AI checkers may learn to detect a certain pattern that is indicative of AI-generated content. Then, the next LLM may incidentally or deliberately not exhibit that pattern and avoid detection. The AI checker then has to be updated to detect said LLM, and so forth.
buuuut my roadmap is’nt ten times shorter. and I’m definitely not ten times better at deciding what is worth building. teams haven’t started casually shipping a year of product work every month. you can look around and and it’s hard to say what software has gotten meaningfully better in the last year.
Everything we see, whether it's this room or any room or anywhere you go, everything you see was either grown or mined, manufactured and moved. Everything. And so we look at what we do as industrial AI.
Chips come from the ground. Where's the energy come from? And a lot of people are like, "Oh, it comes from the sun." Yeah, it comes from the sun. But how are you capturing it from the sun? From stuff made from the ground, right?
For the first time, AI has "hands" - that is, the ability to reach out and interact with the tools on your computer and in your browser, just as you would.
I'm so incredibly tired of seeing low-effort AI-designed web pages. I basically bounce off any new product immediately when I see it (never get a chance to judge the product itself). Lots of thin lines, glowy styles, inconsistent fonts, lots of monospace.
It's like they have all the wisdom of their most senior engineers looking at every single diff and that is fantastic. Which means that you don't have to worry about remembering and looking and nitpicking and all the things that we're not good at anyway
You can use AI as a shortcut to help you not have to think too much. And you can use AI to help you think more deeply and more rigorously. And both of those use cases have their place. But when it comes to your core job function, we primarily want the second one, right?
But if you have to fine-tune a model, you actually aren't getting a general purpose model um for the things that you want it to do because you have to fine-tune it for each individual thing.
And in all of these applications, the customer is making a decision based off of the recommendation of the AI model more or less. Uh, and this means that if the customer is ultimately like kind of making the decision, this means that if the system makes a mistake, um, that's okay because usually the person can kind of recognize that or or decide what to do even despite that mistake.
Uh and this means that they're going to be far more useful when they're operating fully autonomously. And as a result, this requires us to develop physical AI systems that make far fewer mistakes than the machine learning systems that have been deployed thus far.
I mean at the very least I actually think that just starting with a generalist policy and then fine-tuning it even like right off the bat uh can be really effective.
The translation & dubbing was a huge amount of work (which I deeply care about) with humans & AI collaborating. Thanks to the amazing team at @ElevenLabs for their help.
There has also been an uptick in product shills, which, in some ways is worse than AI stuff-some human actually put effort into making straight garbage.
when AIs are extremely extremely capable my view is that those AIs will be harder to align than current systems. So for current systems, we have this feedback loop where we basically like we create an AI. We do some evaluations on it. We see that it has some kind of messed up behavior that we can kind of quickly understand. Then we like can like go look in training and be like, "Oh, the these training environments led to this problematic behavior. Let's like tweak that training data. Let's introduce some additional training data to like correct this other issue and then move forward from there." But in a regime where the AIs are extremely situationally aware, very very very very capable and um you know uh we don't necessarily understand what they're doing, this feedback loop breaks down.
I do I do think that I wish that sort of my preferred constitution or like the way I would orient towards this like the thing I would prefer would be more like Claude is like look it would be structurally good for the way this technology work like the constitution should be like it would be structurally good for the way this technology works to be that AIS are like good fiduciaries, good representatives, the equivalent of a lawyer for a user
the reason why RL environments today are much better than they were in like you know 2024 is not that much because um we have hired way more human experts to make RL environments and is instead much more because we better know what how RL like what RL environments we even want to make and and like how we should structure them and also we're using huge amounts of AI labor to build RL environments.
Second, I think ML is a very shallow domain relative to math. So I think in math there's much more of a you find some true deep abstraction um and then like that like if you really understand that thing which is hard to understand then you get somewhere
Facts and Fallacies of Software Engineering by Robert L. Glass. In essence, this is a book about an industry that refuses to learn. That was true 25 years ago when this book was published, and it's probably twice as true today. (Just think about all the AI adoption metrics being rolled out — back to productivity mistaken for lines of code produced, only more elaborate. And expensive). What I like about this book is that Glass doesn't present anything new. Quite the opposite, actually. Rather, it's about research lessons that we all should know, but tend to forget. Ever had to do an estimate, or plan according to a requirements spec? Or maybe you thought that enough eyeballs make all bugs shallow? Then this book is for you. A great work by a fantastic author.
Paradigms of Artificial Intelligence Programming by Peter Norvig. Learning the AI described in this book probably won't land you a job today. But reading the code examples will transform how you think about source code. The book shines when it comes to code comments, a topic that I've never seen demonstrated well in other sources. Here we get to see how comments become valuable as a narrative that explains both intent and reasoning. Brilliant, just brilliant.
The absolute bad outcome is that our young generation, their agency and human level motivation of learning and living is taken away by tools. So doom scrolling, passive watching of shorts, all this are not helping agency, human agency.
They are the most important people in our society. We should be talking to them. We should be uplifting them. We should be supporting them. We should be providing resources to them.
Coding isn't yet another application domain -- it's the meta-skill required for AI to automatically develop its own training material, via symbolic world models. That's how the RSI loop actually kicks off.
In addition to reliability issues, it often engenders a mind-numbing workflow and an environment where junior developers will never acquire the expertise to become senior developers capable of designing complex systems.
But if you zoom out, it becomes clear that almost every “breakthrough” since last summer has concerned the narrow domains of computer code and math, which are defined by highly structured languages and come accompanied by massive amounts of specialized training data.
It seems pretty clear that the only thing data center opposition is doing is pushing data center construction to the (many) places in America that won't regulate energy or pollution. Weird thing to pat oneself on the back for.
His Lansing rally with Bernie and AOC, more than anything else, crystallized the connection between data center opposition and voters' anger about money in politics.
an ai is oddly well suited to this job. cassandra doesn’t want a promotion and doesn’t need the people in the channel to like her. disagreeing with the vp three times this month won’t show up in a performance review. she doesn’t have an ego invested in being right either.
In the era of base LLM scaling (2022-2024), I believed the LLM line of research would reach a capability plateau (as later seen with base LLMs).
In late 2024, after the o3 test-time compute demo, I changed my views: the new models were showing genuine fluid intelligence, and with this new line of work, the LLM line of research could achieve unbounded capability scaling. "There will be no wall."
I believe current techniques are 4-6 orders of magnitude away from optimality in terms of data efficiency and test-time compute efficiency. But far future AI will be near-optimal.
Rather, the way that AI companies engage with communities-forcing NDAs, dangling billion-dollar promises, pushing environmental externalities far away from AI's wealthy user base-resembles a classic story about dark money in politics.
And that's why every hype cycle produces a wave of absolutely spectacular demos and very few real products. And the recurring mistake of every cycle is spending on the demo when you should be saving for the nines.
Given the high cost of errors, you need to have a very high level of safety and a very high level of confidence on day one before you deploy your first robot, before you drive your your first autonomous mile.
So, a deployment of your agent uh in the real world generates data. That data then grounds the simulator and makes it more realistic. The simulator generates harder edge cases for the critic to score and for the agent to learn from.
This is such a useful book. It makes the case that there's no such thing as "human error" - instead most catastrophes are caused by system issues, misaligned incentives, and unrealistic processes. You are not the custodian of an otherwise safe system that you need to protect from erratic human beings.
most of commerce in America isn't actually high intent. Like how many years are we into e-commerce now? Like 30 years into e-commerce and e-commerce has never exceeded 20% of retail spend in America.
absolutely agree that product skills are probably the most durable. My build would just be there's a lot of people who have the title PM who haven't spent a lot of time building those skills in the last 5 years, but have gotten really good at communicating frameworks to leadership.
as a product manager, I've spent less time in the last year talking to a data scientist than I ever have in my career, even though I've probably spent 10 times more time in data and understanding actually how the product's working than I have ever have in my career.
But based on what we can see at Stripe, the hunger and the intensity with which other companies are either getting started, taking advantage of these new capabilities, or existing companies are retooling, I don't worry about the centralization in the same way. Uh I think there I think there are going to be many thousands of winners.
And just human organizations are complicated and it's very hard to have um to manage to aggressively prosecute 100 different priorities and to deal with all the issues and interference that arises among them and so forth. And so, you know, Google has done incredibly well in a bunch of specific places, but it's not like Google has done all the things even if in some kind of basic material sense, uh Google maybe, you know, had that ability.
I mean, it's very interesting, right? Because these can prove the Jacobian conjecture, you know, whatever. Uh and so clearly they're capable of these monumental feats. Um but somehow I still haven't read the LLM essay that I found super compelling.
So, I um yeah, I think maybe maybe a better way of saying it is 20 20 years ago that whole lean startup thing was uh was almost the only thing to do because of capital available and you didn't have AI that made, I don't know, spinning up an organization with many different potentialities and capabilities so much easier, whereas now I think you can start these much more aggressive and ambitious things up front.
Yeah, I mean I feel like uh the models have been getting a lot better at sort of agent-based longer running coding tasks and it seems pretty clear that they are now actually pretty capable and depending on exactly your definition of of junior engineer it seems pretty spot-on I would say.
Um, and to give you an example of a a a use of a coding agent that works extremely well is you can ask today's models to translate software from one computer language to another very effectively because in that case you actually have a incredibly detailed specification.
Um, and sometimes that's because the model is trying to do something it doesn't have a lot of experience doing. So it's been trained on a whole set of things and as soon as you get a little bit off the distribution of things it knows how to do then like most machine learning models it will you know its performance will suddenly will start to degrade and the farther you get off the comfort zone of what it knows how to do the the more likely it is to to not work as well.
Yeah, I mean I think if you looked at what is important in AI systems these days, you would want to know things like the bandwidth between you know your main memory system on your accelerator to the onchip memory to the um you know the multiplier unit or whatever. You want to know how much energy does it take to do a single multiplier operation.
If you think about our large scale models today, they probably see a thousand times as much data as a human does by the age of 18. Yet, the human by the age of 18 is better in a lot of things and, you know, on par uh with those frontier models that have seen way more data. So could you come up with much more data efficient systems that can learn continuously learn from their own actions?
So, it's hard to tell how much of like the loss of the past few years was AI versus the end of like zero interest rate policy and like the post-COVID crash. And I think it's more the latter, but like again, LLMs are still getting better.
And often I found with my clients I have to tell them like it's doing a good job at generating the actual design, but in actually expressing what the design is supposed to do, it cannot do that yet. You have to do that part yourself.
As a general thing we've seen like to get good results you have to already know how to get good results without it. It just helps you get good results faster.
I predict that in the next 10 years software development will survive, but it will become like any other white-collar professional work. No more $200,000 salaries, unlimited vacation, or incredible employee bargaining power.
we don't believe in a world where these models are so expensive that, you know, they get rationed only for the most wealthy of developers and and companies.
we don't believe in this totalizing, you know, totalitarian view of, you know, AIs that control the world. We believe that these are going to enhance this very broad ecosystem.
so much of that debate is like I think um in some ways uh a little bit of a waste of time because, you know, I think it's inevitable that we're going to have very powerful models
we believe that everybody in the world, you know, all the billions of people in the world are going to have a super intelligence that is adapted and tailored to them, that is enables them to accomplish their goals, knows their context, and ultimately is an expander of their own agency.
I think we're at this like in amazing moment in the world where the bottleneck is not the progress of the AI models, the bottleneck is diffusing that through the rest of the world and and helping the world adapt to this amazing technology that already exists. Like I think if the models didn't improve at all from today, there would still be like decades and decades of like total upheaval and change in the economy and how the world operates and and everything around us
At boom we need far more software engineers in a postAI world than we need in a pre-AI world. Why? Because the cost of software development has dropped. anybody including hardware engineers can now become a coder and we need software engineers to make sure the architectures are right and make sense and are coherent.
So I think the worst advice is like work on what you know. Uh what you know can be changed. Particularly in a world where you everyone has personalized AI tutors. Anybody with passion and dedication can learn new knowledge and learn new skills. But what you can't change easily is what you love.
I think the skill nowadays is less about prompt engineering and more about figuring out how do you give Claude a hard task that seems a little bit too hard. And then how do you make it possible for Claude to verify its work along the way? And the verification I think is probably the single most important thing that people do not get right
So when I look at engineers that have been, you know, coding for a long for a long time, you know, like for for years or for decades, this is a really really common failure mode is trying to over specify and it's trying to be overly specific and then, you know, get the model to do the to do the task exactly the way that you would have done it. And that that's just not the way the model works.
for people that aren't building agentic products, but you're using Claude code, every 6 months delete your Claude MD. Delete your skills. Delete your hooks.
The evidence would show that and it makes perfect sense that AI and automation is creating jobs everywhere. The narrative about AI destroying jobs is exactly backwards. AI eliminate tasks. AI automates tasks away. But it doesn't necessary doesn't necessarily eliminate jobs.
we want we want to encourage everybody and every company to build their own AIs. And and and who knows what innovation will come from the fact that it's open source.
likely this will be one of the largest industries in the world and um, uh, it'll take longer than a couple two, three years. It'll take less than 10. And so this will this will be our next $100 billion business.
we we kind of have coarse level uh recursive self-improvement already. And the fact that every time you use it, it improves the markdown files. Uh every time you use it, it updates its uh long-term memory.
AI has the potential to assist with both the first and last element of that loop. What it cannot replace is the middle step—doing the work for yourself.
Asking an AI for help with the start of an essay, for instance, will invariably shape what direction you end up following, robbing you of the crucial experience of building your own judgement around a topic and selecting your own path to research and argue.
I genuinely believe that if you took an open weights model from 2025 and built a pentest harness for it, it could do this kind of sandbox escape and scan/hack in most networks. This is only surprising because you assume OpenAI has sounder sandboxes.
For the first time, I ran all potential winners through Pangram, which is an AI-writing detector.1 So according to the machine (and to my ear), all of the winners are 100% human-written.
Martsinovich argues, via code snippets, that there is no such thing as a "conversation" with a chatbot. The AI is born anew every time it speaks, and it simply reads the dialogue so far and then tries to write the next line:
I told him what I believe: that in the age of AI, medicine will be bottlenecked more than ever by regulation and the grind of clinical trials, and that billions poured into faster pre-clinical research won’t touch that problem, unsexy as it is.
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.