There’s no proof that “good code” will matter in the future. The notion of “good code” itself is mostly based on the idea that it’s easy/cheap/efficient for humans to work with. But humans won’t modify the majority of code.
No matter who is and is not at fault, it is rather alarming that the labs cannot cooperate on something like assigning credit for a mathematical proof. This is a very bad sign and also a wake-up call.
the math community needs to adopt a version of the ethical standards of experimental science. If you are using AI agents, you can't just give a proof (formalized or not), but need to also provide a detailed explanation of how these agents were used to get the result.
It's just like we are massively underperforming and people don't believe it when you say 10x 100x but it's actually true and we've seen a lot of proof of it as you point out.
They've got blog posts of we had to rewrite this whole thing because the performance was bad. If it was always hotspots that made your performance bad, you'd never have to rewrite the whole thing. So, we know that that doesn't work anymore.
An excessively AI-polished proof may sand away both the "artificial" friction (typos, awkward phrasing, disorganization) and the "natural" friction, leaving a text that is easy to read and hard to learn from.
Okay, if I'm not going to read this code, how do I know it's going to perform within boundaries of the last code that I generated? That is conformance testing.
Your your models can be leaked, algorithms can be replicated, but hundreds of millions of miles of fully autonomous operations in the real world, backed by evidence-grade evaluation and publicly audited proof, that is much, much more difficult to replicate.
Despite these strides, we argue that current AI4Math systems still largely operate as solvers, excelling at isolated, well-defined proof generation rather than as researchers capable of expanding the boundaries of mathematical knowledge.
And so having anything that's able to give you that green check mark that says even if this is going to be complicated to understand, even if it's going to be a pain, you at the very least know it is correct. Like every other field would kill for that, right?
But a proof can reason about potentially infinite state spaces. So, it can tell you things about like every possible thing that could possibly happen in the entire universe.
One is that the LLMs are getting increasingly good at writing these proofs. And if we don't have to write the proof by hand as humans, it just becomes feasible to do them in situations where previously it would have not been economical.
But also LLMs increase the need for these formal proofs because, you know, we're live coding a bunch of stuff. If we have to manually review all of that code, then that will become the bottleneck.
so so some people are concerned, you know, what if the Riemann hypothesis is proved with a completely incomprehensible proof? I I think once you have the artifact of a proof, we can do a lot of of of of of analysis on it.
I mean, some problems have been basically solved by pure brute force. The four color theorem is is a famous example. Um, we have still not found a conceptually elegant proof of this theorem.
we are moving into a world very quickly this year where proof of work is so important and I mean proof of work not in the Bitcoin sense but your proof of what you have done, your resume. And I don't mean your resume because nobody's going to believe that. I mean the actual work that you did which has to be visible
Extrapolators have a remarkable track record in the AI field, being repeatedly early to trends and capabilities that empiricists believed were still decades away.
so we know deep learning is really bad at this for example we know that if you train on some new thing it will often catastrophically interfere with all the old things that you that you knew
And that's a phase shift, because suddenly it makes sense when you write a paper to write it in Lean first, or through a conversation with AI, which is generally on the fly with you, and it becomes natural for journals to accept.
Eventually, I think it’s kind of hard to imagine, but yes, all of these Nobel Prizes, all of these mathematical proofs, all of these conversations, all of these ideas, all the influence we have on each other, even the AI, eventually will expire.
There's ever more pressure to rebuild society more and more around credentials. Do you have this certificate? Do you have that proof? But companies that are focused on just building great products and doing great things gravitate towards people who do the great work.
If you want to argue that human-level AI is extremely unlikely in the next 20 years, you certainly can, but you should treat that as a minority position where the burden of proof is on you.
But I felt like what this paper showed was that the burden of proof is now on the pessimists. So, that's why we called it the pessimism line. Throughout history, there's been alien pessimists and alien optimists, and they've been yelling at each other, that's all they had to go with.
Despite evaluations, we cannot consider coming powerful frontier AI systems "safe unless proven unsafe". With current testing methodologies, issues can easily be missed. Additionally, it is unclear if governments can quickly build the immense expertise needed for reliable technical evaluations of AI capabilities and societal-scale risks. Given this, developers of frontier AI should carry the burden of proof to demonstrate that their plans keep risks within acceptable limits.
Be skeptical of absolute certainty not backed up with proof. Be skeptical of pronouncements that there’s only one way a court will ever possibly interpret something.
But nobody will ever get to the billions of representative miles necessary to say anything compelling about expected safety until after they actually deploy their fleet.
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.