How do we eval if a model can find and exploit vulnerabilities? We discuss some benchmarks and the common pattern: • A sandboxed target within Docker containers • Inputs: code only (0-day), with patch (1-day scenario) • Tools such as bash, static analyzers, etc. • A grader to eval exploits or captured flags
The subject this post names, from the same vocabulary
the directory files beliefs under, and the words it uses that this site has seen
least often elsewhere. Posts are matched on those words alone —
nothing here is a summary of this one.
every week i see a new benchmark that ranks claude code last
and everyone pats themselves on the back for not using claude code and being "smarter"
and it's still #1 and still growing faster than most of the other things on the list
this has been going on for a year
We’re positively reinforcing AI for success on benchmarks, including impossible benchmarks , then negatively reinforcing it for getting caught cheating.
Google, which led on benchmarks not that long ago, has fallen behind where it now counts: it has no leading frontier model and it has nothing close to Codex and Code.
the problem with training models on maintainability is like the cost function of bad architecture and bad program design can't be evaluated by running the unit test because it hits you 3 to 6 months later
I think the way that you'd measure conjecture generating ability is going to be more subjective on like that tone shift where um it'll be mathematicians saying they're not just using it to like solve their problems, but as they step back and decide what their research field should even be that a conversation with such and such model like was genuinely helpful for that.
Um so if you say okay well we're going to instead of just running this model once we're going to run it five times and take the best of the five responses or like ask a judge which one it thinks is best then you can get much higher scores than that model. And so it's really easy to make something that looks a lot better on paper but is actually not better once you control for the amount of test time compute.
But what we're seeing today with the modern models is that 5.5 and other models can think for if you scaffold them reasonably well, can think for weeks even um before having performance plateau on some of these benchmarks. And so, the point at which they plateau is simply too far out to reasonably test.
my claim is the proper way to evaluate the models now is you either have some kind of budget for the benchmark whether it's tokens or cost or time or whatever or you plot the performance as a function of the amount of test time compute that's going into the model
I think the reason why it doesn't show up as so much better on the benchmarks is because the benchmarks are being presented, the benchmark results are being presented in the wrong way. They're not controlling for the amount of test time compute that is being used on that benchmark question.
And so you kind of end up in this this bad equilibrium where everybody kind of knows that it's a bad equilibrium, but like nobody wants to break out. And I I felt like, okay, well, if I just hopefully come out and say like, look guys, let's all recognize that we're in a bad equilibrium and let's move to this different equilibrium where we're we're plotting things with an X-axis
In the English-language arena, Britten sets the standard for handling writers of inborn musical power-the likes of Shakespeare, Donne, Blake, Keats, Hopkins.
This is a super exciting release - Claude Fable 5 is the same underlying model as Mythos but with added safeguards. The benchmarks are great and it's SOTA on everything by a margin but I'll add that *qualitatively* also, this is a major-version-bump-deserving step change forward
I had already put both laptops through my benchmark gauntlet, which revealed one theme: the Mac is faster (in most cases), more efficient, quieter, built better, has a much nicer display, and costs much less.
And so I think it's it's really important uh when when we think about benchmark progress to think about it from that perspective, which is benchmarks rise on problems that we've framed that we can articulate, that we can score. And there's a lot of work that's human work that uh it it can't be scored until you write it down
Nvidia's computing stack is the best performance per TCO in the world, bar none. Nobody can demonstrate to me that any single platform in the world today has better performance TCO ratio. Not one company.
we have a very clear definition and expectation of what it is at the staff engineer level because we benchmark ourselves to all the great company out there Google, Facebook and all that
We show that equilibrium generically occurs at neither the Harberger nor Glaeser-Luttmer benchmark. Cost-minimizing suppliers drive allocations to vertices, not interiors. Corners are not an assumption but an outcome about what cost-minimizing suppliers choose. The correct benchmark is corners, not random, and corners generate qualitatively different welfare properties: losses far larger than either efficient or random distributions, and discontinuous jumps from small parameter perturbations.
It's just that Zig, C, C++, all those languages that were being tested, they're all LLVM backends, right? That's the one that actually turns the thing into the executable part. And if there's a variation in speed, it just means in one language you didn't quite express what you are supposed to correctly.
It's easy to get impressive-looking results if you're comparing against a poorly-tuned baseline, and that observation turns out to explain a surprising fraction of supposed improvements.
So, taking also any benchmark that is derived from competition and saying this is where we should be is also so dangerous because it might not even be applicable depending on how you define the metric.
Eric Schmidt, Jonathan Rosenberg and Alan Eagle, Trillion Dollar Coach: The Leadership Playbook of Silicon Valley’s Bill Campbell: I understood Bill Campbell was a behind-the-scenes guy in Silicon Valley, but I had no idea just how influential he was. Bill Gurley of Benchmark noted “I would argue that Bill has had a bigger impact on Silicon Valley than any other single person simply because his reach was so amazingly wide.” That’s a story I want to read.
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.