The subject this post names, from the same vocabulary
the directory files beliefs under, and the words it uses that this site has seen
least often elsewhere. Posts are matched on those words alone —
nothing here is a summary of this one.
No matter who is and is not at fault, it is rather alarming that the labs cannot cooperate on something like assigning credit for a mathematical proof. This is a very bad sign and also a wake-up call.
every week i see a new benchmark that ranks claude code last
and everyone pats themselves on the back for not using claude code and being "smarter"
and it's still #1 and still growing faster than most of the other things on the list
this has been going on for a year
It is highly fortunate that the OpenAI agents hacked HuggingFace. This is the only reason we know about all the severe internal failures at OpenAI, and gives us an opportunity to wake up before it is too late.
We’re positively reinforcing AI for success on benchmarks, including impossible benchmarks , then negatively reinforcing it for getting caught cheating.
This will likely serve as a wake-up call to those not paying attention to the US-China space race—which, to be clear, is the vast majority of Americans.
At boom we need far more software engineers in a postAI world than we need in a pre-AI world. Why? Because the cost of software development has dropped. anybody including hardware engineers can now become a coder and we need software engineers to make sure the architectures are right and make sense and are coherent.
Google, which led on benchmarks not that long ago, has fallen behind where it now counts: it has no leading frontier model and it has nothing close to Codex and Code.
the problem with training models on maintainability is like the cost function of bad architecture and bad program design can't be evaluated by running the unit test because it hits you 3 to 6 months later
So I think the the thing I'm most excited is actually like what we call like iterated loops or like slow loops where we basically have a cron job. We have the loop the the the structure of the loop is really easy. It's like run this llinter fix one thing commit and push and then we run that every night in our GitHub actions and we wake up every morning to one PR that makes the codebase a little bit better.
At this point, I had pretty much given up and filed Thunderbolt's "wake from sleep" promise in the "unreliable" folder, right next to Bluetooth Pairing.
I got the CalDigit TS4 as soon as it was released, in 2022. It was perhaps a little bit better but overall I was disappointed to still sometimes encounter a machine that would not wake up.
I think the way that you'd measure conjecture generating ability is going to be more subjective on like that tone shift where um it'll be mathematicians saying they're not just using it to like solve their problems, but as they step back and decide what their research field should even be that a conversation with such and such model like was genuinely helpful for that.
Um so if you say okay well we're going to instead of just running this model once we're going to run it five times and take the best of the five responses or like ask a judge which one it thinks is best then you can get much higher scores than that model. And so it's really easy to make something that looks a lot better on paper but is actually not better once you control for the amount of test time compute.
But what we're seeing today with the modern models is that 5.5 and other models can think for if you scaffold them reasonably well, can think for weeks even um before having performance plateau on some of these benchmarks. And so, the point at which they plateau is simply too far out to reasonably test.
my claim is the proper way to evaluate the models now is you either have some kind of budget for the benchmark whether it's tokens or cost or time or whatever or you plot the performance as a function of the amount of test time compute that's going into the model
I think the reason why it doesn't show up as so much better on the benchmarks is because the benchmarks are being presented, the benchmark results are being presented in the wrong way. They're not controlling for the amount of test time compute that is being used on that benchmark question.
And so you kind of end up in this this bad equilibrium where everybody kind of knows that it's a bad equilibrium, but like nobody wants to break out. And I I felt like, okay, well, if I just hopefully come out and say like, look guys, let's all recognize that we're in a bad equilibrium and let's move to this different equilibrium where we're we're plotting things with an X-axis
In the English-language arena, Britten sets the standard for handling writers of inborn musical power-the likes of Shakespeare, Donne, Blake, Keats, Hopkins.
But if you took something that she loves and you make it one inch better, she might love that more than if you showed her something she's never seen before and didn't wake up knowing that she wanted.
This is a super exciting release - Claude Fable 5 is the same underlying model as Mythos but with added safeguards. The benchmarks are great and it's SOTA on everything by a margin but I'll add that *qualitatively* also, this is a major-version-bump-deserving step change forward
know your goal or suffer a death by a thousand compromises. Because what I had done my whole career and most of us do is compromise to get that next engineer, CTO, investor. You put a jerk on your board because you impressed with their firm name and the valuation and all your friends are going to be impressed and it's going to be so much easier. We make all these compromises and contort ourselves and eventually we wake up and it's company we don't want to work at.
I had already put both laptops through my benchmark gauntlet, which revealed one theme: the Mac is faster (in most cases), more efficient, quieter, built better, has a much nicer display, and costs much less.
And so I think it's it's really important uh when when we think about benchmark progress to think about it from that perspective, which is benchmarks rise on problems that we've framed that we can articulate, that we can score. And there's a lot of work that's human work that uh it it can't be scored until you write it down
Nvidia's computing stack is the best performance per TCO in the world, bar none. Nobody can demonstrate to me that any single platform in the world today has better performance TCO ratio. Not one company.
we have a very clear definition and expectation of what it is at the staff engineer level because we benchmark ourselves to all the great company out there Google, Facebook and all that
We show that equilibrium generically occurs at neither the Harberger nor Glaeser-Luttmer benchmark. Cost-minimizing suppliers drive allocations to vertices, not interiors. Corners are not an assumption but an outcome about what cost-minimizing suppliers choose. The correct benchmark is corners, not random, and corners generate qualitatively different welfare properties: losses far larger than either efficient or random distributions, and discontinuous jumps from small parameter perturbations.
And I spoke with Anthony Beevor once about the attempt of British intelligence to assassinate Hitler, and he had seen some evidence that at the point in time, they dropped those plans because they knew that drugged Hitler or malfunctioning Hitler, which he was after, you know, the summer of 1943, is better for Britain than, you know, killing Hitler and then having to deal with, like, some kind of, you know, maybe the army would have taken over the country
We dropped atomic bombs on Hiroshima and Nagasaki. Those were not military targets. We were not doing anything strategic against the country other than terrorizing the country by killing women and children.
What the Apple paper shows, most fundamentally, regardless of how you define AGI, is that LLMs are no substitute for good well-specified conventional algorithms.
I didn't find 8Sleep to be very reliable in its sleep tracking. The scores don't make as much sense to me when I wake up, and as we saw above they don't correlate very strongly with Whoop or Oura.
It's just that Zig, C, C++, all those languages that were being tested, they're all LLVM backends, right? That's the one that actually turns the thing into the executable part. And if there's a variation in speed, it just means in one language you didn't quite express what you are supposed to correctly.
It's easy to get impressive-looking results if you're comparing against a poorly-tuned baseline, and that observation turns out to explain a surprising fraction of supposed improvements.
So, taking also any benchmark that is derived from competition and saying this is where we should be is also so dangerous because it might not even be applicable depending on how you define the metric.
And then we end up at the bottom of that with this idea of everyday I wake up and I check my phone and I'm like, oh, it's going to be 60 degrees out. Great. And we start thinking that 60 degrees is more real than hot and cold. That thermodynamics, the whole formal structure of thermodynamics is more real than the basic experience of hot and cold that it came from.
In Italy, average life expectancies in the solidly Medieval 1200s were 35-40, while by the year 1500 (definitely Renaissance) life expectancy in Italian city states had dropped to 18.
Eric Schmidt, Jonathan Rosenberg and Alan Eagle, Trillion Dollar Coach: The Leadership Playbook of Silicon Valley’s Bill Campbell: I understood Bill Campbell was a behind-the-scenes guy in Silicon Valley, but I had no idea just how influential he was. Bill Gurley of Benchmark noted “I would argue that Bill has had a bigger impact on Silicon Valley than any other single person simply because his reach was so amazingly wide.” That’s a story I want to read.
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.