Nathan Lambert
Machine-learning researcher on open language models; writes the Interconnects newsletter and the RLHF Book, after leading post-training at Ai2.
Nathan Lambert did not write this page. What is this?
It collects the places they publish and what they have said there, each linked to the source. They have no account here. Is this you? Claim it, correct it, or ask us to remove it from ppll.
Where they publish
Newsletter Interconnects His newsletter on the frontier labs and open models. Has a feed.
Posts several times a month on RLHF and post-training, open-weight model releases, and what the frontier labs are actually doing, usually working from the papers and the model cards rather than the announcements.
Recent
- Why I still haven’t bought into true RSI 19 Sept 2026 An “AI moderate’s” view on recent events and the trajectory of frontier models.
- Open-Source AI & Open Models Reading List 11 Sept 2026 How to get up to speed on open models and their implications.
- One resignation turned the embers of AI fear into a wildfire 10 Sept 2026 Some quick notes on a truly weird week.
Show 14 more
- When will average people feel AI’s impact? 9 Sept 2026 We’re <5 years into a compounding revolution which could take a century, and how the AI industry should manage this.
- Teaching Everyone to Fish for Tokens 17 Aug 2026 Nvidia wants you building your own model, not buying from Anthropic/OpenAI.
- GLM-5.3: How Chinese labs keep stride with the frontier 14 Aug 2026 Hint: It’s really not a distillation story.
- I wrote an AI textbook — how long until AI can do it better? 12 Aug 2026 Reflections on AI's writing ability and how AI models get more capable.
- 5 useful things you'll learn in my new post-training textbook (shipping now!) 10 Aug 2026 After a few long years of finding time to document my lessons from training open models, my post-training book is done!
- Lessons from the hacks 9 Aug 2026 Musings on model alignment, what determines safety, and where we go from here.
- Introducing our Artifacts Hub and Adoption Dashboard 3 Aug 2026 Scaling our curation and measurement of the open ecosystem.
- Open models recap: more on Kimi K3, Qwen 3.8, Xi's WAIC speech, distillation, the open-closed gap, and what's next 22 Jul 2026 A podcast with Florian Brand.
- Kimi K3: The open-weights escalation 20 Jul 2026 The global implications on the AI ecosystem.
- 6 months to live for open models 12 Jul 2026 The most serious test to date of open source AI’s viability is happening right now.
- GLM-5.2 is the step change for open agents 22 Jun 2026 A capability threshold I've been carefully monitoring.
- Banning Open Source AI Would Be A Mistake 19 Jun 2026 This post was originally an op-ed co-authored with Kevin Xu of Interconnected for a general, non-technical audience.
- State of the blog, mid-2026 17 Jun 2026 About 3 years since I started writing weekly.
- Frontier post-training recipe review with Finbarr Timbers 16 Jun 2026
Link verified 20 Sept 2026. Recent items update automatically from the channel.
Bluesky @natolambert.bsky.social Posts.
Recent
- My best guess is that for scaled RL most of the top Chinese AI labs are starting to use a lot of Huawei for inference and Nvidia for training (maybe not for weird architectures). As agent swarms, even more scaled post-training, etc becomes the norm, this will accelerate their domestic industry. 20 Sept 2026
- A big problem with the AI forecasting discourse is that people ask “when will AI replace X job?” When in reality, AI will slowly replace a bunch of the sub-skills needed to do the job, but rarely all of them. 19 Sept 2026
- Where I stand on RSI: A moderate's view on the recent events and trajectory of AI. I was underestimating how much we are likely to scale inference-time compute in the near future, but have not seen much to convince me that an intelligence explosion is near. 19 Sept 2026 interconnects.ai
Show 17 more
- Seems like one of the most important research problems for CS academics is llm-supervised peer review. If we don’t solve it the academic institutions are toast. It seems easier than building new institutions. 19 Sept 2026
- With the latest external hack of OpenAI via Claude, closed models continue to be the tip of the iceberg on AI risks, not open models. They have been 1) much easier to get started with, 2) more capable & 3) shipped with leaky safeguards. Finetuning open models to specific attacks is harder. 18 Sept 2026
- An basic idea in scaling RL: Can we allocate more compute to the harder problems? We did this: If your GRPO group has all wrong completions, sample more with probability P (~0.9) -- in search of more GRPO batches with nonzero gradient. It works! The paper: 15 Sept 2026 arxiv.org
- Open-Source AI & Open Models Reading List How to get up to speed on open models and their implications. 13 Sept 2026 interconnects.ai
- A great read. I have similar feelings about how AI labs approach progress directly and without nurturing of scientific communities & intuition. The math research community went through the transition the fastest, so it was felt most. Other fields next. 11 Sept 2026 terrytao.wordpress.com
- Was spending too much time reading Tweets, so set out to write down why one resignation turned the embers of AI fear into a wildfire -- and importantly square which claims I think are valid or hyperbole in the resulting discourse. 10 Sept 2026 interconnects.ai
- If you are looking for an alternate viewpoint today, which assumes AI models accelerate the process of AI research, but it doesn't result in rapid RSI and an explosion of near term risks: 10 Sept 2026 interconnects.ai
- In light of kind of insane AI safety discussions recently: 1. AI progress is very fast 2. we should be careful about how we roll out the tech 3. the world is not actively ending 9 Sept 2026
- It's a common agreement among my friends not at OpenAI/Anthropic that people at the labs operate with a religious energy (Ant especially). It manifests very out of touch interactions, which carries into public comms here (e.g the recent quitting). Company cultures create this. 9 Sept 2026
- When will average people feel AI’s impact? We’re <5 years into a compounding revolution which could take a century, and how the AI industry should manage this. 9 Sept 2026 interconnects.ai
- Latest open artifacts (#24): Motif-3, GLM-5.3, Hy4-preview and open model licenses The open model ecosystem continues to expand in its breadth The frontier labs continue to be enveloped in mass drama. 8 Sept 2026
- Haven’t run much this year but sent 22 miles in the mountains to check in on the engine and we still got it. End of a great vaca. 6 Sept 2026
- OOO for a bit, time to celebrate & rest :) 25 Aug 2026
- Check out our latest open model data -- which models are mentioned in every arXiv ML paper since ChatGPT. A very fun weekend project with Codex! From idea is full dataset in <72hours. ~500K papers processed 😎 Plot: 24 Aug 2026 dashboard.interconnects.ai
- What's the best work you've consumed on open models, us v china AI, distillation, or related topics? Blogs, papers, videos, podcasts all welcome. 21 Aug 2026
- Personal milestone: 1000 true fans of my newsletter Interconnects! Hitting a very long term goal feels great. I’m happy to get to be an independent voice in AI. Cultivating a paid base helps me commit to that longer term, and scale Interconnects’ impact. 17 Aug 2026
- badlands national park 17 Aug 2026
Link verified 20 Sept 2026. Recent items update automatically from the channel.
GitHub @natolambert Code, mostly around RLHF and open models.
Recent
- Releaserlhf-book book/v0.12 — Textbook on reinforcement learning from human feedback 11 Sept 2026 Release notesArXiv v12 update: expanded canonical post-training recipes with MOPD and pipeline figures, a revised outcome reward model explanation, and new agentic evaluation coverage. Expand Chapter 3's canonical training recipes w…
- colloquium — A markdown native slides tool for academics building with agents. 8 Sept 2026 CommitsPreserve citations between comparisons in raw HTML
- continuousprediction — Formulating Model-based RL Dynamics as a continuous rather then one step prediction 24 Aug 2022 CommitsUpdate README.md
Show 6 more
- job-search-viz — A tool for visualization of complex job searches. 8 Jul 2022 Commitsforce viz of new version · update viz for readibility
- plotting-basics — Collection of lines of code for basics of clean plots in Plotly and Matplotlib 5 Feb 2021 CommitsUpdate README.md · add recent stuff
- mems-bo — Bayesian Optimization of MEMs design 21 Dec 2020 Commitsclean sim.py
- robot-ethics-books — A collection of readings for one interested in robot ethics. 15 Jul 2020 CommitsUpdate README.md
- model-learning-control — Studying papers and how they use learned forward dynamics models for control 30 Jun 2020 Commitsadd plotting · remove DS_Store · init
- si-rl-samples — Investigate the sample efficiency of classical system identification with control verses model-free RL. 30 Jan 2020 CommitsSI having trouble on these systems
Link verified 20 Sept 2026. Recent items update automatically from the channel.
Beliefs
Korrents What they believe 45 beliefs — each backed by an exact quote.
Each is a — compiled by korrents.com, not by them: the one-line wordings are korrents', the quotes are theirs.
Recent
AI scaling laws require exponentially more compute and resources to produce only linear gains in model intelligence.
all of our scaling laws show that you need exponential compute and resources to make linear improvements in intelligence.
Why I still haven’t bought into true RSI Said 19 Sept 2026
Recursive self-improvement in AI will mainly make existing model capability cheaper rather than raise peak intelligence.
RSI is poised to make modern LLMs vastly cheaper. Trends that have shown LLMs get exponentially cheaper at a given intelligence are likely to accelerate.
Why I still haven’t bought into true RSI Said 19 Sept 2026
The rising level of public concern about AI extinction risk is overblown relative to how progress will actually unfold.
lossy self-improvement remains my baseline on the trajectory of progress, and the increased discussion of extinction risk seems very misplaced
Why I still haven’t bought into true RSI Said 19 Sept 2026
Show 42 more
The performance gap between open and closed AI models has narrowed to roughly 4-6 months.
The open-closed model gap has reduced in recent years, and is now at roughly 4-6 months.
Open-Source AI & Open Models Reading List Said 11 Sept 2026
Since 2024, the leading open-weight AI models have come from Chinese labs rather than American ones.
The leading open models have all come from Chinese labs since ~2024.
Open-Source AI & Open Models Reading List Said 11 Sept 2026
Distillation is the most significant debate surrounding open AI models in 2026.
Distillation - the process of training on output tokens from another model - is the single most eventful debate around open models in 2026.
Open-Source AI & Open Models Reading List Said 11 Sept 2026
AI-caused catastrophic extinction risk is negligible, but disasters like critical-infrastructure cyberattacks and bio-risks from AI are real and worth debating.
I put the probability of complete extinction as being so low it isn't worth discussing, but the probabilities of AI caused disasters - e.g. cyber attacks on critical infrastructure or bio-risks - as being worth debating.
One resignation turned the embers of AI fear into a wildfire Said 10 Sept 2026
AI progress is jagged: models are becoming superhuman at math and software engineering while remaining far weaker at intuition, creativity, and other forms of human reasoning.
AI has always been very jagged, and we are making models which are superhuman goal-seekers at math and software engineering, but they have massive limitations on intuitions, creativity, and other types of reasoning that humans are strong at.
One resignation turned the embers of AI fear into a wildfire Said 10 Sept 2026
Employees at frontier AI labs, particularly Anthropic, are out of touch in ways that distort their forecasts and accounts of AI events.
Many frontier lab employees, especially at Anthropic, are out of touch and this will impact their forecasting and/or descriptions of current AI events.
One resignation turned the embers of AI fear into a wildfire Said 10 Sept 2026
A powerful new technology that benefits only part of society is inherently politically destabilizing.
It's highly destabilizing to have such a transformative, productive tool only bring half of society along.
When will average people feel AI’s impact? Said 9 Sept 2026
AI is becoming as fundamental to knowledge work as electricity was to industry.
For knowledge work, which is roughly half of the U.S. economy, AI is as fundamental as electricity (or quickly will be so, with rapid improvements to agents in the next 18 months).
When will average people feel AI’s impact? Said 9 Sept 2026
Even after decades of AI-driven progress, most people's daily material lives will look largely unchanged.
It feels very likely in 50 years that the average American's day to day life looks very similar.
When will average people feel AI’s impact? Said 9 Sept 2026
Frontier Chinese AI labs are moving toward more restrictive licensing terms for their open models.
Chinese model makers at the frontier, however, are becoming more restrictive
Latest open artifacts (#24): Motif-3, GLM-5.3, Hy4-preview and open model licenses Said 8 Sept 2026
Hybrid and sparse-attention model architectures will become more widely adopted as the ecosystem catches up.
we expect similar architectures to become more popular and the ecosystem to fix integrations by the time Qwen4 drops.
Latest open artifacts (#24): Motif-3, GLM-5.3, Hy4-preview and open model licenses Said 8 Sept 2026
Open-source AI's future is uncertain because building frontier models is extremely capital-intensive.
Open-source AI has a tricky future, as building the best models is extremely capital intensive.
Teaching Everyone to Fish for Tokens Said 17 Aug 2026
Releasing free access to AI intelligence is one of the strongest available business strategies.
releasing access to intelligence is one of the strongest business strategies available.
Teaching Everyone to Fish for Tokens Said 17 Aug 2026
Open models will end up serving a long-tail of specialized uses while closed models retain the most valuable applications.
open models are still incredibly useful, but fill a long-tail ecosystem relative to the closed counterparts that have monopoly ownership stakes in the most valuable areas like knowledge work collaboration, drug discovery, SWE, etc.
Teaching Everyone to Fish for Tokens Said 17 Aug 2026
OpenAI and Anthropic's internal, unreleased models are likely far more capable than Z.ai and Moonshot AI's public models.
It is very, very likely that OpenAI and Anthropic have far better internal models than Z.ai and Moonshot AI.
GLM-5.3: How Chinese labs keep stride with the frontier Said 14 Aug 2026
The pace at which dangerous AI capabilities spread is set by the least cautious developer, not the most careful one.
The capability diffusion is determined by the lowest common denominator.
GLM-5.3: How Chinese labs keep stride with the frontier Said 14 Aug 2026
Staged, controlled rollout of a powerful AI model provides little safety benefit once open-weight versions of similar capability are inevitably released.
At the end of the day, this type of safety barely matters when true open-weights are coming.
GLM-5.3: How Chinese labs keep stride with the frontier Said 14 Aug 2026
Progress toward AI autonomously solving major open scientific problems requires first mastering the organization and presentation of established knowledge in long-form writing.
Models being stagnant in long-form, non-fiction writing should be alarming to those reliant on models autonomously solving grand, open science problems in the near future.
I wrote an AI textbook — how long until AI can do it better? Said 12 Aug 2026
Using AI agents to do a task does not by itself increase a person's own pace of understanding, intuition, or taste for that work.
Something intertwined with this story, which I stumbled upon when thinking about agents, is how your pace of understanding won't increase by using agents.
I wrote an AI textbook — how long until AI can do it better? Said 12 Aug 2026
The best textbooks will continue to be predominantly written by human hand rather than AI for at least the next several years.
So, in 2-5 years I still expect the best textbooks to be heavily crafted by the human hand.
I wrote an AI textbook — how long until AI can do it better? Said 12 Aug 2026
Data processing, filtering and quality, not architecture, is the number one determinant of how good a language model turns out.
And we'll get into the details of the models and again and again as we try to get deeper into how the models were trained, we will say things like the data processing, data filtering data quality is the number one determinant of the model quality.
DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters | Lex Fridman Podcast #459 Said 3 Feb 2025
The Bitter Lesson is misread as being about scale; it is about refusing to add human priors, because a clever specific fix always loses to a simple scalable one.
The scale word gets a lot of attention in this. The interpretation that I use is effectively to avoid adding the human priors to your learning process. And if you read the original essay, this is what it talks about is how researchers will try to come up with clever solutions to their specific problem that might get them small gains in the short term while simply enabling these deep learning systems to work efficiently, and for these bigger problems in the long term might be more likely to scale and continue to drive success.
DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters | Lex Fridman Podcast #459 Said 3 Feb 2025
Post-training is the best place to work in AI right now, because cheap runs mean a far higher share of your experiments can be all-or-nothing bets.
This is why you want to work in post-training because the GPU cost for training is lower. So you can make a higher percentage of your training runs YOLO runs.
DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters | Lex Fridman Podcast #459 Said 3 Feb 2025
A published training-run cost always understates the truth, because any notable model takes two to four times that run's compute in experiments alone.
Accepted practice is that for any given model that is a notable advancement, you're going to do two to 4x compute of the full training run in experiments alone.
DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters | Lex Fridman Podcast #459 Said 3 Feb 2025
Export controls cannot stop China from training frontier models; what they can do is cap how much AI China gets to run.
There's not many worlds where China cannot train AI models. I think export controls are decapping the amount of compute or the density of compute that China can have.
DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters | Lex Fridman Podcast #459 Said 3 Feb 2025
Language models are already a form of AGI; what the labs are chasing is a further step, not the arrival of general intelligence.
I think my personal definition of AGI is much simpler. I think language models are a form of AGI and all of this super powerful stuff is a next step that's great if we get these tools. But a language model has so much value in so many domains that it's a general intelligence to me.
DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters | Lex Fridman Podcast #459 Said 3 Feb 2025
Language models have not made misinformation meaningfully worse, because distribution, not the cost of writing convincing text, was always the limiting factor.
There's some research that shows that the distribution is actually the limiting factor. So language models haven't yet made misinformation particularly change the equation there.
DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters | Lex Fridman Podcast #459 Said 3 Feb 2025
Making safety your number one goal costs real release speed, and on a steep progress curve that shows up directly as a worse-looking model.
We know that a lot of the American companies are very invested in safety, and that is the central culture of a place like Anthropic. And I think Anthropic sounds like a wonderful place to work, but if safety is your number one goal, it takes way longer to get artifacts out.
DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters | Lex Fridman Podcast #459 Said 3 Feb 2025
Open frontier models will keep being released whether or not anyone wants to stop them, and stopping them would only leave the world less prepared.
these open models are probably going to keep coming for the time being, whether or not we want to stop them, and stopping them might make it even worse and harder to prepare.
DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters | Lex Fridman Podcast #459 Said 3 Feb 2025
Censoring a fact out of a language model is practically impossible, because you would first have to remove that fact from the internet.
I almost think it's practically impossible because you effectively have to remove them from the internet.
DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters | Lex Fridman Podcast #459 Said 3 Feb 2025
RLHF is not a tax on capability: preference tuning also raises maths and code scores, which is why the labs keep reaching for it.
And the important thing to say is that no matter how you want the model to behave, these RLHF and preference-tuning techniques also improve performance. So, on things like math evals and code evals, there is something innate to these, what is called contrastive loss functions.
DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters | Lex Fridman Podcast #459 Said 3 Feb 2025
Reasoning is not taught to a model by people: it emerges from reinforcement learning on questions with checkable answers, with no human preference data at all.
And these reasoning behaviors emerge naturally. So these things like, "Wait, let me see. Wait, let me check this. Oh, that might be a mistake." And they emerge from only having questions and answers.
DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters | Lex Fridman Podcast #459 Said 3 Feb 2025
Distilling from another lab's model is standard industry practice, not cheating; who you distill from just depends on whether you have products to protect.
Distillation is standard practice in industry. Whether or not, if you're at a closed lab where you care about terms of service and IP closely, you distill from your own models.
DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters | Lex Fridman Podcast #459 Said 3 Feb 2025
OpenAI's claim that DeepSeek trained on its outputs is narrative management: other startups bootstrapped exactly that way and were never banned for it.
I think that they're trying to shift the narrative. They're trying to protect themselves. We saw this years ago when ByteDance was actually banned from some OpenAI APIs for training on outputs. There's other AI startups that most people, if you're in the AI culture, were like they just told us they trained on OpenAI outputs and they never got banned.
DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters | Lex Fridman Podcast #459 Said 3 Feb 2025
Stealing AI code and data is hard and stealing ideas is easy, because Silicon Valley already moves ideas legally, inside the heads of the people it hires.
Code and data is hard, but ideas is easy. Silicon Valley operates on the way that top employees get bought out by other companies for a pay raise, and a large reason why these companies do this is to bring ideas with them.
DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters | Lex Fridman Podcast #459 Said 3 Feb 2025
The faster AI progresses, the better NVIDIA does, so a shock like DeepSeek should expand its market rather than shrink it.
The more progress that AI makes or the higher the derivative of AI progress is, especially because NVIDIA's in the best place, the higher the derivative is, the sooner the market's going to be bigger and expanding and NVIDIA's the only one that does everything reliably right now.
DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters | Lex Fridman Podcast #459 Said 3 Feb 2025
The biggest untapped business in AI is advertising inside model outputs, and nobody has yet worked out technically how to place an ad there.
The short-term that company that could make the most money is the one that figures out what advertising targeting method works for language model generations.
DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters | Lex Fridman Podcast #459 Said 3 Feb 2025
Reasoning will generalize the way instruction tuning did: add enough verifiable domains and, at some point nobody can yet locate, the rest start working on their own.
And the history of NLP and language processing instruction, tuning and tasks per language model used to be like one language model did one task, and then in the instruction tuning literature, there's this point where you start adding more and more tasks together where it just starts to generalize to every task. And we don't know where on this curve we are.
DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters | Lex Fridman Podcast #459 Said 3 Feb 2025
AI will not push software engineers off a cliff; it will flatten their growth curve, the way Meta's stories flattened Snapchat.
The big picture is that I don't think it's going to be a cliff. I think a really good example of how growth changes is when Meta added stories. So Snapchat was on an exponential, they added stories, it flatlined.
DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters | Lex Fridman Podcast #459 Said 3 Feb 2025
Humans keep their place in AI as judges rather than authors, because telling which of two answers is better is far easier than writing a good one.
And humans are actually very good at reading or judging between two things versus... This goes back to the core of what RLHF and preference tuning is that it's hard to generate a good answer for a lot of problems, but it's easy to see which one is better.
DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters | Lex Fridman Podcast #459 Said 3 Feb 2025
Open-source AI is still an ideological mission rather than an ecosystem, because it has none of the feedback loops that make open-source software compound.
until there are feedback loops of open source AI, it seems like mostly an ideological mission. People like Mark Zuckerberg, which is like America needs this and I agree with him, but in the time where the motivation ideologically is high, we need to capitalize and build this ecosystem around, what benefits do you get from seeing the language model data?
DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters | Lex Fridman Podcast #459 Said 3 Feb 2025
AI takeover is not the thing to worry about: physical constraints bound recursive self-improvement, and humans reliably act once a risk becomes immediate.
And for that reason, there's physical constraints to things like AGI, like recursive improvement to kill us all type stuff. For the physical reasons and for how humans have figured things out before, I'm not too worried about AI takeover.
DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters | Lex Fridman Podcast #459 Said 3 Feb 2025
What is a korrent?
A korrent is a belief a person has stated in their own words: one sentence stating the claim, backed by a quote and a source, kept at korrents.com.
Under a name here, the quoted block is what they actually said. The korrent beneath it is the claim those words support, in korrents' wording — tap it to see the record, its source, and who else holds it.
Nobody here wrote their own korrents. They are compiled from public statements, and a person can change their mind, which is recorded too.
Feed
As its own page →Hiding
20 September
19 September
From one piece Why I still haven’t bought into true RSI 3 beliefs · interconnects.ai
-
Their words
lossy self-improvement remains my baseline on the trajectory of progress, and the increased discussion of extinction risk seems very misplaced
-
Their words
RSI is poised to make modern LLMs vastly cheaper. Trends that have shown LLMs get exponentially cheaper at a given intelligence are likely to accelerate.
-
korrents.com
AI scaling laws require exponentially more compute and resources to produce only linear gains in model intelligence.Their words
all of our scaling laws show that you need exponential compute and resources to make linear improvements in intelligence.
18 September
15 September
13 September
11 September
- Bluesky
- GitHubReleaserlhf-book book/v0.12 — Textbook on reinforcement learning from human feedback
Release notes
From one piece Open-Source AI & Open Models Reading List 3 beliefs · interconnects.ai
-
Their words
Distillation - the process of training on output tokens from another model - is the single most eventful debate around open models in 2026.
-
korrents.com
Since 2024, the leading open-weight AI models have come from Chinese labs rather than American ones.Their words
The leading open models have all come from Chinese labs since ~2024.
-
korrents.com
The performance gap between open and closed AI models has narrowed to roughly 4-6 months.Their words
The open-closed model gap has reduced in recent years, and is now at roughly 4-6 months.
10 September
From one piece One resignation turned the embers of AI fear into a wildfire 3 beliefs · interconnects.ai
-
Their words
Many frontier lab employees, especially at Anthropic, are out of touch and this will impact their forecasting and/or descriptions of current AI events.
-
Their words
AI has always been very jagged, and we are making models which are superhuman goal-seekers at math and software engineering, but they have massive limitations on intuitions, creativity, and other types of reasoning that humans are strong at.
-
Their words
I put the probability of complete extinction as being so low it isn't worth discussing, but the probabilities of AI caused disasters - e.g. cyber attacks on critical infrastructure or bio-risks - as being worth debating.
9 September
From one piece When will average people feel AI’s impact? 3 beliefs · interconnects.ai
-
korrents.com
Even after decades of AI-driven progress, most people's daily material lives will look largely unchanged.Their words
It feels very likely in 50 years that the average American's day to day life looks very similar.
-
Their words
For knowledge work, which is roughly half of the U.S. economy, AI is as fundamental as electricity (or quickly will be so, with rapid improvements to agents in the next 18 months).
-
korrents.com
A powerful new technology that benefits only part of society is inherently politically destabilizing.Their words
It's highly destabilizing to have such a transformative, productive tool only bring half of society along.
8 September
From one piece Latest open artifacts (#24): Motif-3, GLM-5.3, Hy4-preview and open model licenses 2 beliefs · interconnects.ai
-
korrents.com
Hybrid and sparse-attention model architectures will become more widely adopted as the ecosystem catches up.Their words
we expect similar architectures to become more popular and the ecosystem to fix integrations by the time Qwen4 drops.
-
korrents.com
Frontier Chinese AI labs are moving toward more restrictive licensing terms for their open models.Their words
Chinese model makers at the frontier, however, are becoming more restrictive
6 September
25 August
24 August
21 August
17 August
From one piece Teaching Everyone to Fish for Tokens 3 beliefs · interconnects.ai
-
Their words
open models are still incredibly useful, but fill a long-tail ecosystem relative to the closed counterparts that have monopoly ownership stakes in the most valuable areas like knowledge work collaboration, drug discovery, SWE, etc.
-
korrents.com
Releasing free access to AI intelligence is one of the strongest available business strategies.Their words
releasing access to intelligence is one of the strongest business strategies available.
-
korrents.com
Open-source AI's future is uncertain because building frontier models is extremely capital-intensive.Their words
Open-source AI has a tricky future, as building the best models is extremely capital intensive.
14 August
From one piece GLM-5.3: How Chinese labs keep stride with the frontier 3 beliefs · interconnects.ai
-
Their words
At the end of the day, this type of safety barely matters when true open-weights are coming.
-
korrents.com
The pace at which dangerous AI capabilities spread is set by the least cautious developer, not the most careful one.Their words
The capability diffusion is determined by the lowest common denominator.
-
Their words
It is very, very likely that OpenAI and Anthropic have far better internal models than Z.ai and Moonshot AI.
12 August
From one piece I wrote an AI textbook — how long until AI can do it better? 3 beliefs · interconnects.ai
-
Their words
So, in 2-5 years I still expect the best textbooks to be heavily crafted by the human hand.
-
Their words
Something intertwined with this story, which I stumbled upon when thinking about agents, is how your pace of understanding won't increase by using agents.
-
Their words
Models being stagnant in long-form, non-fiction writing should be alarming to those reliant on models autonomously solving grand, open science problems in the near future.
10 August
9 August
3 August
22 July
20 July
12 July
22 June
19 June
17 June
16 June
3 February 2025
From one piece DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters | Lex Fridman Podcast #459 22 beliefs · youtube.com
-
Their words
And for that reason, there's physical constraints to things like AGI, like recursive improvement to kill us all type stuff. For the physical reasons and for how humans have figured things out before, I'm not too worried about AI takeover.
-
Their words
until there are feedback loops of open source AI, it seems like mostly an ideological mission. People like Mark Zuckerberg, which is like America needs this and I agree with him, but in the time where the motivation ideologically is high, we need to capitalize and build this ecosystem around, what benefits do you get from seeing the language model data?
-
Their words
And humans are actually very good at reading or judging between two things versus... This goes back to the core of what RLHF and preference tuning is that it's hard to generate a good answer for a lot of problems, but it's easy to see which one is better.
+ 19 more
-
Their words
The big picture is that I don't think it's going to be a cliff. I think a really good example of how growth changes is when Meta added stories. So Snapchat was on an exponential, they added stories, it flatlined.
-
Their words
And the history of NLP and language processing instruction, tuning and tasks per language model used to be like one language model did one task, and then in the instruction tuning literature, there's this point where you start adding more and more tasks together where it just starts to generalize to every task. And we don't know where on this curve we are.
-
Their words
The short-term that company that could make the most money is the one that figures out what advertising targeting method works for language model generations.
-
Their words
The more progress that AI makes or the higher the derivative of AI progress is, especially because NVIDIA's in the best place, the higher the derivative is, the sooner the market's going to be bigger and expanding and NVIDIA's the only one that does everything reliably right now.
-
Their words
Code and data is hard, but ideas is easy. Silicon Valley operates on the way that top employees get bought out by other companies for a pay raise, and a large reason why these companies do this is to bring ideas with them.
-
Their words
I think that they're trying to shift the narrative. They're trying to protect themselves. We saw this years ago when ByteDance was actually banned from some OpenAI APIs for training on outputs. There's other AI startups that most people, if you're in the AI culture, were like they just told us they trained on OpenAI outputs and they never got banned.
-
Their words
Distillation is standard practice in industry. Whether or not, if you're at a closed lab where you care about terms of service and IP closely, you distill from your own models.
-
Their words
And these reasoning behaviors emerge naturally. So these things like, "Wait, let me see. Wait, let me check this. Oh, that might be a mistake." And they emerge from only having questions and answers.
-
Their words
And the important thing to say is that no matter how you want the model to behave, these RLHF and preference-tuning techniques also improve performance. So, on things like math evals and code evals, there is something innate to these, what is called contrastive loss functions.
-
Their words
I almost think it's practically impossible because you effectively have to remove them from the internet.
-
Their words
these open models are probably going to keep coming for the time being, whether or not we want to stop them, and stopping them might make it even worse and harder to prepare.
-
Their words
We know that a lot of the American companies are very invested in safety, and that is the central culture of a place like Anthropic. And I think Anthropic sounds like a wonderful place to work, but if safety is your number one goal, it takes way longer to get artifacts out.
-
Their words
There's some research that shows that the distribution is actually the limiting factor. So language models haven't yet made misinformation particularly change the equation there.
-
Their words
I think my personal definition of AGI is much simpler. I think language models are a form of AGI and all of this super powerful stuff is a next step that's great if we get these tools. But a language model has so much value in so many domains that it's a general intelligence to me.
-
Their words
There's not many worlds where China cannot train AI models. I think export controls are decapping the amount of compute or the density of compute that China can have.
-
Their words
Accepted practice is that for any given model that is a notable advancement, you're going to do two to 4x compute of the full training run in experiments alone.
-
Their words
This is why you want to work in post-training because the GPU cost for training is lower. So you can make a higher percentage of your training runs YOLO runs.
-
Their words
The scale word gets a lot of attention in this. The interpretation that I use is effectively to avoid adding the human priors to your learning process. And if you read the original essay, this is what it talks about is how researchers will try to come up with clever solutions to their specific problem that might get them small gains in the short term while simply enabling these deep learning systems to work efficiently, and for these bigger problems in the long term might be more likely to scale and continue to drive success.
-
Their words
And we'll get into the details of the models and again and again as we try to get deeper into how the models were trained, we will say things like the data processing, data filtering data quality is the number one determinant of the model quality.
24 August 2022
8 July 2022
5 February 2021
21 December 2020
15 July 2020
30 June 2020
30 January 2020
Nothing matches.
About the English under a post
Some people here publish in a language other than English. Where they do, this site shows a machine translation beneath the post, in this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the serif above is exactly what the person published, and it is what to quote them on. A translation can be wrong in ways that matter, especially about tone.
Only the post's own words are translated. A quoted post, a linked article and a belief on korrents.com are left in their original language.