Noam Brown
AI researcher at OpenAI; built the poker systems Libratus and Pluribus and the Diplomacy agent Cicero, and works on reasoning models.
Noam Brown did not write this page. What is this?
It collects the places they publish and what they have said there, each linked to the source. They have no account here. Is this you? Claim it, correct it, or ask us to remove it from ppll.
Where they publish
Beliefs
Korrents What they believe 18 beliefs — each backed by an exact quote.
Each is a — compiled by korrents.com, not by them: the one-line wordings are korrents', the quotes are theirs.
Recent
Model comparisons understate real progress, because benchmark tables do not control for how much test-time compute each answer used.
I think the reason why it doesn't show up as so much better on the benchmarks is because the benchmarks are being presented, the benchmark results are being presented in the wrong way. They're not controlling for the amount of test time compute that is being used on that benchmark question.
Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown Said 26 Jun 2026
Running a model until its performance plateaus is no longer a usable evaluation rule, because a well-scaffolded model keeps improving for weeks.
But what we're seeing today with the modern models is that 5.5 and other models can think for if you scaffold them reasonably well, can think for weeks even um before having performance plateau on some of these benchmarks. And so, the point at which they plateau is simply too far out to reasonably test.
Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown Said 26 Jun 2026
A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number.
my claim is the proper way to evaluate the models now is you either have some kind of budget for the benchmark whether it's tokens or cost or time or whatever or you plot the performance as a function of the amount of test time compute that's going into the model
Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown Said 26 Jun 2026
Show 15 more
Running a model five times and keeping the best answer buys a higher benchmark score without buying a better model.
Um so if you say okay well we're going to instead of just running this model once we're going to run it five times and take the best of the five responses or like ask a judge which one it thinks is best then you can get much higher scores than that model. And so it's really easy to make something that looks a lot better on paper but is actually not better once you control for the amount of test time compute.
Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown Said 26 Jun 2026
Within a year a model will write an entire state-of-the-art poker solver in one shot — the work of a whole PhD thesis.
And I wouldn't be surprised if you know 6 months or a year from now, the model is able to do zero shot an entire poker solver, basically my entire PhD thesis in one go.
Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown Said 26 Jun 2026
The safety frameworks the labs publish do not account for how much test-time compute a dangerous-capability evaluation was given.
the preparedness frameworks and responsible scaling policies, they don't really account for the amount of test time compute.
Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown Said 26 Jun 2026
A model's capability is now a function of how much money you spend on it, so asking what a model can do means nothing until you name a budget.
The problem is we're in a world now where the capability of the model is a function of how much money you put into it. Basically, if you give it a budget of $10,000, it can do a lot more than what it can do with a budget of $10. Give it a budget of $10 million, you can do even more. At what budget should you evaluate these models? The policies that exist today don't really address that question.
Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown Said 26 Jun 2026
The only way to know what a model can do after running for a month is to actually run it for a month.
And the problem is if you want to evaluate the capabilities of a model, what it can do after running for a month, the only way to be fully sure is to actually run it for a month. And if you want to know after 6 months, the only way to know fully is to run it for six months.
Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown Said 26 Jun 2026
Evaluating a model properly would mean delaying its release, and competitive pressure means no lab will.
It's actually very difficult because, yeah, you would have to the only way to to really do the evaluations is then delay the model release cycle. Um and you know there's a lot of competitive pressure right now to not do that.
Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown Said 26 Jun 2026
An outsider could have disproved the Erdős unit-distance conjecture with a public model, because nobody had tried spending $100,000 of compute on one question.
Um but it would be possible and it would have been possible for somebody to disprove the erdos unit distance conjecture before we did using a general purpose model. And nobody had explored sufficiently what happens if I put $100,000 worth of compute into 5.5 what could it do?
Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown Said 26 Jun 2026
A frontier lab should not spend its researchers harvesting results out of today's models; the payoff is in building the next one.
we are trying to encourage people to not spend all their time just like going through all the mathematical open problems physics problems and um just seeing pushing the models to their limits to see what they can prove or disprove.
Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown Said 26 Jun 2026
Extra thinking time cannot help with factual recall: a model that does not know a date will not know it after a week either.
if you ask a person when was Abraham Lincoln born and they don't know the date. They could sit there, they could think about it for a week, if they if they don't have access to Wikipedia or something, they're not going to be able to do better answering that question if they thought about it for a week compared to 5 seconds. Same with the model.
Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown Said 26 Jun 2026
Models have no research taste yet, which is why they complement researchers rather than replace them.
one thing I see for research in particular is they don't have very good research taste right now and so I think they're actually a very good complement to researchers
Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown Said 26 Jun 2026
A model can make an existing algorithm a hundred times faster and still cannot invent a better one, however long you give it.
And then I was like, okay, can you come up with an algorithm that is better than the algorithms that I came up with or that anybody else came up with and go ahead and like look at all the published work and synthesize that and then try to come up with something novel and it it's not able to do it. And I can give it a lot of time and it's it's still not able to do it.
Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown Said 26 Jun 2026
There will be no overnight intelligence explosion, because a model's best work takes so much test-time compute that time itself is the bottleneck.
and I don't think we're headed to that world largely because of the fact that the models rely so much on large scale test time compute in order to achieve um their greatest intelligence. If you if it requires so much test time on compute to unlock the full capabilities of the model, then that means you're bottlenecked by time
Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown Said 26 Jun 2026
Humanity's advantage is accumulated shared knowledge rather than higher intelligence — and AI models, which vanish with their context window, have none of it.
it's not that humans have become smarter over it's not that they evolved to become smarter over, you know, the past 50,000 years. It's that humans are able to do a lot more today than they were back in caveman times because there have been billions of humans thinking for a long time and building off of each other's accumulated knowledge.
Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown Said 26 Jun 2026
On everyday high-stakes questions, a model's answer can now be trusted more than an expert human's.
And I think they're at a point now where they've actually been at a point for a while now where I feel like I can just trust the outputs arguably more than I could trust the output from from a human,
Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown Said 26 Jun 2026
Every lab knows the benchmark grid is the wrong way to present a model, and publishes it anyway because everybody else does.
And so you kind of end up in this this bad equilibrium where everybody kind of knows that it's a bad equilibrium, but like nobody wants to break out. And I I felt like, okay, well, if I just hopefully come out and say like, look guys, let's all recognize that we're in a bad equilibrium and let's move to this different equilibrium where we're we're plotting things with an X-axis
Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown Said 26 Jun 2026
What is a korrent?
A korrent is a belief a person has stated in their own words: one sentence stating the claim, backed by a quote and a source, kept at korrents.com.
Under a name here, the quoted block is what they actually said. The korrent beneath it is the claim those words support, in korrents' wording — tap it to see the record, its source, and who else holds it.
Nobody here wrote their own korrents. They are compiled from public statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they do, this site shows a machine translation beneath the post, in this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the serif above is exactly what the person published, and it is what to quote them on. A translation can be wrong in ways that matter, especially about tone.
Only the post's own words are translated. A quoted post, a linked article and a belief on korrents.com are left in their original language.