Related posts
Brendan Gregg GitHub
The subject this post names, from the same vocabulary the directory files beliefs under, and the words it uses that this site has seen least often elsewhere. Posts are matched on those words alone — nothing here is a summary of this one.
Top people
Showing Profile →
Hiding
Hiding
20 September
17 September
16 September
14 September
12 September
7 September
5 September
4 September
-
Their words
You have to treat Mythos 5.1 as being a Tier 2 manipulator, until and unless you can show that it is not one.
Claude Fable 5.1 and Mythos 5.1: The System Cardthezvi.substack.com
3 September
-
Their words
every week i see a new benchmark that ranks claude code last and everyone pats themselves on the back for not using claude code and being "smarter" and it's still #1 and still growing faster than most of the other things on the list this has been going on for a year
-
Their words
what the benchmark is actually telling you is the shit you think matters does not matter
1 September
-
korrents.com
The pelican-drawing benchmark's correlation with genuine model capability has weakened since 2025.Their words
its connection to how good the models were at other tasks didn’t seem to hold as strongly as it did back in 2025
-
Their words
We’re positively reinforcing AI for success on benchmarks, including impossible benchmarks , then negatively reinforcing it for getting caught cheating.
24 August
23 August
22 August
7 August
6 August
31 July
22 July
30 June
-
Their words
I think the way that you'd measure conjecture generating ability is going to be more subjective on like that tone shift where um it'll be mathematicians saying they're not just using it to like solve their problems, but as they step back and decide what their research field should even be that a conversation with such and such model like was genuinely helpful for that.
26 June
From one piece Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown 4 beliefs, in the piece's order there
-
korrents.com
Running a model five times and keeping the best answer buys a higher benchmark score without buying a better model.Their words
Um so if you say okay well we're going to instead of just running this model once we're going to run it five times and take the best of the five responses or like ask a judge which one it thinks is best then you can get much higher scores than that model. And so it's really easy to make something that looks a lot better on paper but is actually not better once you control for the amount of test time compute.
-
Their words
my claim is the proper way to evaluate the models now is you either have some kind of budget for the benchmark whether it's tokens or cost or time or whatever or you plot the performance as a function of the amount of test time compute that's going into the model
-
Their words
I think the reason why it doesn't show up as so much better on the benchmarks is because the benchmarks are being presented, the benchmark results are being presented in the wrong way. They're not controlling for the amount of test time compute that is being used on that benchmark question.
+ 1 more
-
Their words
And so you kind of end up in this this bad equilibrium where everybody kind of knows that it's a bad equilibrium, but like nobody wants to break out. And I I felt like, okay, well, if I just hopefully come out and say like, look guys, let's all recognize that we're in a bad equilibrium and let's move to this different equilibrium where we're we're plotting things with an X-axis
17 June
-
Their words
In the English-language arena, Britten sets the standard for handling writers of inborn musical power-the likes of Shakespeare, Donne, Blake, Keats, Hopkins.
29 May
-
Likedrcmnd.app
MacBook NeoTheir words
I had already put both laptops through my benchmark gauntlet, which revealed one theme: the Mac is faster (in most cases), more efficient, quieter, built better, has a much nicer display, and costs much less.
25 May
-
korrents.com
Designing a good AI benchmark has become a task only the most capable people can do well.Their words
Creating benchmarks is now a job relegated to the smartest and best of us.
24 May
-
Their words
And so I think it's it's really important uh when when we think about benchmark progress to think about it from that perspective, which is benchmarks rise on problems that we've framed that we can articulate, that we can score. And there's a lot of work that's human work that uh it it can't be scored until you write it down
15 April
-
Their words
Nvidia's computing stack is the best performance per TCO in the world, bar none. Nobody can demonstrate to me that any single platform in the world today has better performance TCO ratio. Not one company.
1 April
-
Their words
we have a very clear definition and expectation of what it is at the staff engineer level because we benchmark ourselves to all the great company out there Google, Facebook and all that
12 February
-
Their words
We show that equilibrium generically occurs at neither the Harberger nor Glaeser-Luttmer benchmark. Cost-minimizing suppliers drive allocations to vertices, not interiors. Corners are not an assumption but an outcome about what cost-minimizing suppliers choose. The correct benchmark is corners, not random, and corners generate qualitatively different welfare properties: losses far larger than either efficient or random distributions, and discontinuous jumps from small parameter perturbations.
23 November 2025
-
Their words
The benchmark is human performance, not perfection. We sometimes get requirements for 90%+ accuracy.
31 May 2025
22 March 2025
-
korrents.com
A speed difference between languages that share the LLVM backend measures the benchmark author, not the languages.Their words
It's just that Zig, C, C++, all those languages that were being tested, they're all LLVM backends, right? That's the one that actually turns the thing into the executable part. And if there's a variation in speed, it just means in one language you didn't quite express what you are supposed to correctly.
9 March 2025
-
Their words
It's easy to get impressive-looking results if you're comparing against a poorly-tuned baseline, and that observation turns out to explain a surprising fraction of supposed improvements.
19 January 2025
-
korrents.com
Industry benchmarks are a dangerous target because every company defines the underlying metric differently.Their words
So, taking also any benchmark that is derived from competition and saying this is where we should be is also so dangerous because it might not even be applicable depending on how you define the metric.
13 May 2024
From one piece The Evolving Landscape of LLM Evaluation 2 beliefs, in the piece's order there
-
Their words
GPT models perform much better on coding problems released before their pre-training data cut-off.
-
Their words
The time when benchmarks lasted multiple decades has passed. Going forward, we will rely less on public benchmark results.
28 May 2019
-
Recommendsaffiliate linkrcmnd.app
Trillion Dollar Coach: The Leadership Playbook of Silicon Valley's Bill CampbellTheir words
Eric Schmidt, Jonathan Rosenberg and Alan Eagle, Trillion Dollar Coach: The Leadership Playbook of Silicon Valley’s Bill Campbell: I understood Bill Campbell was a behind-the-scenes guy in Silicon Valley, but I had no idea just how influential he was. Bill Gurley of Benchmark noted “I would argue that Bill has had a bigger impact on Silicon Valley than any other single person simply because his reach was so amazingly wide.” That’s a story I want to read.
27 February 2019
9 July 2018
15 November 2017
Nothing matches.
What is a korrent?
A korrent is a belief a person has stated in their own words: one sentence stating the claim, backed by a quote and a source, kept at korrents.com.
Under a name here, the quoted block is what they actually said. The korrent beneath it is the claim those words support, in korrents' wording — tap it to see the record, its source, and who else holds it.
Nobody here wrote their own korrents. They are compiled from public statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they do, this site shows a machine translation beneath the post, in this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the serif above is exactly what the person published, and it is what to quote them on. A translation can be wrong in ways that matter, especially about tone.
Only the post's own words are translated. A quoted post, a linked article and a belief on korrents.com are left in their original language.