What public figures publish and believe, in their own words.
About this feed
Highlights: posts that did unusually well for the person who wrote them, everything they published at length, each release and new project, and every belief — at most two a day from anyone. Day by day, newest day first; within a day, the people with the most beliefs on this site come first. Nothing else orders it. Show everything instead.
The quoted blocks are what people actually said; a beneath one is the belief those words support, in korrents' wording. Nobody here wrote their own page.
Top people are the people in this feed with the most beliefs on this site, then the most here. Choose an area and the row leads with the people whose beliefs are about it; tap a face for their feed.
Top people in benchmarks
Showing Profile →
Hiding
Hiding
31 July
23 July
-
korrents.com
Google currently has no leading frontier AI model and no agentic coding tool comparable to Codex or Claude CodeTheir words
Google, which led on benchmarks not that long ago, has fallen behind where it now counts: it has no leading frontier model and it has nothing close to Codex and Code.
22 July
15 July
-
Their words
the problem with training models on maintainability is like the cost function of bad architecture and bad program design can't be evaluated by running the unit test because it hits you 3 to 6 months later
30 June
-
Their words
I think the way that you'd measure conjecture generating ability is going to be more subjective on like that tone shift where um it'll be mathematicians saying they're not just using it to like solve their problems, but as they step back and decide what their research field should even be that a conversation with such and such model like was genuinely helpful for that.
26 June
From one piece Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown 5 beliefs, in the piece's order there
-
korrents.com
Running a model five times and keeping the best answer buys a higher benchmark score without buying a better model.Their words
Um so if you say okay well we're going to instead of just running this model once we're going to run it five times and take the best of the five responses or like ask a judge which one it thinks is best then you can get much higher scores than that model. And so it's really easy to make something that looks a lot better on paper but is actually not better once you control for the amount of test time compute.
-
Their words
But what we're seeing today with the modern models is that 5.5 and other models can think for if you scaffold them reasonably well, can think for weeks even um before having performance plateau on some of these benchmarks. And so, the point at which they plateau is simply too far out to reasonably test.
-
Their words
my claim is the proper way to evaluate the models now is you either have some kind of budget for the benchmark whether it's tokens or cost or time or whatever or you plot the performance as a function of the amount of test time compute that's going into the model
+ 2 more
-
Their words
I think the reason why it doesn't show up as so much better on the benchmarks is because the benchmarks are being presented, the benchmark results are being presented in the wrong way. They're not controlling for the amount of test time compute that is being used on that benchmark question.
-
Their words
And so you kind of end up in this this bad equilibrium where everybody kind of knows that it's a bad equilibrium, but like nobody wants to break out. And I I felt like, okay, well, if I just hopefully come out and say like, look guys, let's all recognize that we're in a bad equilibrium and let's move to this different equilibrium where we're we're plotting things with an X-axis
25 June
17 June
-
Their words
In the English-language arena, Britten sets the standard for handling writers of inborn musical power-the likes of Shakespeare, Donne, Blake, Keats, Hopkins.
9 June
-
Lovedrcmnd.app
Claude Fable 5Their words
This is a super exciting release - Claude Fable 5 is the same underlying model as Mythos but with added safeguards. The benchmarks are great and it's SOTA on everything by a margin but I'll add that *qualitatively* also, this is a major-version-bump-deserving step change forward
29 May
-
Likedrcmnd.app
MacBook NeoTheir words
I had already put both laptops through my benchmark gauntlet, which revealed one theme: the Mac is faster (in most cases), more efficient, quieter, built better, has a much nicer display, and costs much less.
25 May
-
korrents.com
Designing a good AI benchmark has become a task only the most capable people can do well.Their words
Creating benchmarks is now a job relegated to the smartest and best of us.
24 May
-
Their words
And so I think it's it's really important uh when when we think about benchmark progress to think about it from that perspective, which is benchmarks rise on problems that we've framed that we can articulate, that we can score. And there's a lot of work that's human work that uh it it can't be scored until you write it down
15 April
-
Their words
Nvidia's computing stack is the best performance per TCO in the world, bar none. Nobody can demonstrate to me that any single platform in the world today has better performance TCO ratio. Not one company.
1 April
-
Their words
we have a very clear definition and expectation of what it is at the staff engineer level because we benchmark ourselves to all the great company out there Google, Facebook and all that
12 February
-
Their words
We show that equilibrium generically occurs at neither the Harberger nor Glaeser-Luttmer benchmark. Cost-minimizing suppliers drive allocations to vertices, not interiors. Corners are not an assumption but an outcome about what cost-minimizing suppliers choose. The correct benchmark is corners, not random, and corners generate qualitatively different welfare properties: losses far larger than either efficient or random distributions, and discontinuous jumps from small parameter perturbations.
30 December 2025
23 November 2025
-
Their words
The benchmark is human performance, not perfection. We sometimes get requirements for 90%+ accuracy.
16 November 2025
-
korrents.com
Fair hardware benchmarks must use the best software and tuning for each hardware option, not hardware specs alone.Their words
The most accurate and realistic evaluation for HW involves selecting the best software and then tuning it, and doing this for all HW options.
5 October 2025
22 June 2025
-
Their words
Since these datasets are likely already part of model training data, we shouldn't rely solely on them to evaluate our Q&A system.
31 May 2025
31 March 2025
22 March 2025
-
korrents.com
A speed difference between languages that share the LLVM backend measures the benchmark author, not the languages.Their words
It's just that Zig, C, C++, all those languages that were being tested, they're all LLVM backends, right? That's the one that actually turns the thing into the executable part. And if there's a variation in speed, it just means in one language you didn't quite express what you are supposed to correctly.
9 March 2025
-
Their words
It's easy to get impressive-looking results if you're comparing against a poorly-tuned baseline, and that observation turns out to explain a surprising fraction of supposed improvements.
19 January 2025
-
korrents.com
Industry benchmarks are a dangerous target because every company defines the underlying metric differently.Their words
So, taking also any benchmark that is derived from competition and saying this is where we should be is also so dangerous because it might not even be applicable depending on how you define the metric.
13 May 2024
From one piece The Evolving Landscape of LLM Evaluation 2 beliefs, in the piece's order there
-
Their words
GPT models perform much better on coding problems released before their pre-training data cut-off.
-
Their words
The time when benchmarks lasted multiple decades has passed. Going forward, we will rely less on public benchmark results.
28 May 2019
-
Recommendsaffiliate linkrcmnd.app
Trillion Dollar Coach: The Leadership Playbook of Silicon Valley's Bill CampbellTheir words
Eric Schmidt, Jonathan Rosenberg and Alan Eagle, Trillion Dollar Coach: The Leadership Playbook of Silicon Valley’s Bill Campbell: I understood Bill Campbell was a behind-the-scenes guy in Silicon Valley, but I had no idea just how influential he was. Bill Gurley of Benchmark noted “I would argue that Bill has had a bigger impact on Silicon Valley than any other single person simply because his reach was so amazingly wide.” That’s a story I want to read.
27 February 2019
9 July 2018
15 November 2017
Nothing matches.
What is a korrent?
A korrent is a belief a person has stated in their own words: one sentence stating the claim, backed by a quote and a source, kept at korrents.com.
Under a name here, the quoted block is what they actually said. The korrent beneath it is the claim those words support, in korrents' wording — tap it to see the record, its source, and who else holds it.
Nobody here wrote their own korrents. They are compiled from public statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they do, this site shows a machine translation beneath the post, in this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the serif above is exactly what the person published, and it is what to quote them on. A translation can be wrong in ways that matter, especially about tone.
Only the post's own words are translated. A quoted post, a linked article and a belief on korrents.com are left in their original language.