There’s no proof that “good code” will matter in the future. The notion of “good code” itself is mostly based on the idea that it’s easy/cheap/efficient for humans to work with. But humans won’t modify the majority of code.
Targeting old 16-bit computers, the authors couldn't afford to build a sloppy, wasteful UI, and so by modern standards the game UI is remarkably fast and responsive.
Increased reliance on CFAA prosecutions for public policy goals couldn't possibly be a cursed monkey's paw situation, said no practitioner in this space ever.
No matter who is and is not at fault, it is rather alarming that the labs cannot cooperate on something like assigning credit for a mathematical proof. This is a very bad sign and also a wake-up call.
the math community needs to adopt a version of the ethical standards of experimental science. If you are using AI agents, you can't just give a proof (formalized or not), but need to also provide a detailed explanation of how these agents were used to get the result.
It's just like we are massively underperforming and people don't believe it when you say 10x 100x but it's actually true and we've seen a lot of proof of it as you point out.
They've got blog posts of we had to rewrite this whole thing because the performance was bad. If it was always hotspots that made your performance bad, you'd never have to rewrite the whole thing. So, we know that that doesn't work anymore.
An excessively AI-polished proof may sand away both the "artificial" friction (typos, awkward phrasing, disorganization) and the "natural" friction, leaving a text that is easy to read and hard to learn from.
there's not a technical reason that they couldn't have plaid lunchboxes and shoelaces or be singing a song, but it would probably be more distracting and just take away from that instant recognition that you're aiming for.
Okay, if I'm not going to read this code, how do I know it's going to perform within boundaries of the last code that I generated? That is conformance testing.
This helped the leader shift from wanting to change Jacob-something they couldn't control-to taking responsibility for changing the relationship-something they could.
The team and I use Auto mode exclusively, and have been for many months. I couldn't imagine going back to permission prompts! Really excited to get this out to everyone.
Your your models can be leaked, algorithms can be replicated, but hundreds of millions of miles of fully autonomous operations in the real world, backed by evidence-grade evaluation and publicly audited proof, that is much, much more difficult to replicate.
Despite these strides, we argue that current AI4Math systems still largely operate as solvers, excelling at isolated, well-defined proof generation rather than as researchers capable of expanding the boundaries of mathematical knowledge.
And so having anything that's able to give you that green check mark that says even if this is going to be complicated to understand, even if it's going to be a pain, you at the very least know it is correct. Like every other field would kill for that, right?
I mean it's not actually clear that we couldn't run it as a business if we wanted to. I just think that we'll have a bigger impact by getting this in more scientist hands quicker
And our harness wasn't very good for like the first like five months of open code. But it was good enough. It was good enough that most people couldn't really tell a difference. And once we won enough share, then we went back and like tried to make our harness like good and smart and optimize and all those things. But uh it was inverted from what everybody else was doing. Everybody else was being like you have to build the smartest harness and that's how you win.
But a proof can reason about potentially infinite state spaces. So, it can tell you things about like every possible thing that could possibly happen in the entire universe.
One is that the LLMs are getting increasingly good at writing these proofs. And if we don't have to write the proof by hand as humans, it just becomes feasible to do them in situations where previously it would have not been economical.
But also LLMs increase the need for these formal proofs because, you know, we're live coding a bunch of stuff. If we have to manually review all of that code, then that will become the bottleneck.
so so some people are concerned, you know, what if the Riemann hypothesis is proved with a completely incomprehensible proof? I I think once you have the artifact of a proof, we can do a lot of of of of of analysis on it.
I mean, some problems have been basically solved by pure brute force. The four color theorem is is a famous example. Um, we have still not found a conceptually elegant proof of this theorem.
we are moving into a world very quickly this year where proof of work is so important and I mean proof of work not in the Bitcoin sense but your proof of what you have done, your resume. And I don't mean your resume because nobody's going to believe that. I mean the actual work that you did which has to be visible
Extrapolators have a remarkable track record in the AI field, being repeatedly early to trends and capabilities that empiricists believed were still decades away.
I Who Have Never Known Men by Jacqueline Hartman: A quick and haunting bit of speculative fiction that I still think about all the time. Couldn't put it down.
So this was not only in the Epic of Gilgamesh, but it was also in the Book of Genesis. So what it meant was that it wasn’t… You couldn’t have two stories… It wasn’t two stories about the same thing. It was literary dependence.
so we know deep learning is really bad at this for example we know that if you train on some new thing it will often catastrophically interfere with all the old things that you that you knew
This was just because Hitler was afraid of the power of his army high command, and convinced by Goering's morphine-high vision that he would stop it with the air force, which he couldn't.
So when you make something like chat GPT, you can go from to zero to 100 million users faster than any technology in history because of the other technologies though. like you couldn't have gotten there if not for the build out of the internet, if not for the smartphone.
Ellen Ullman's Close to the Machine is a memoir I couldn't believe I hadn't read earlier. Many of her observations about engineering culture feel as relevant today as when she was a programmer in the 1980s
When this is good and on point, it is very, very good, and provides essential perspective I couldn’t have found elsewhere, especially when doing key interviews.
And that's a phase shift, because suddenly it makes sense when you write a paper to write it in Lean first, or through a conversation with AI, which is generally on the fly with you, and it becomes natural for journals to accept.
Eventually, I think it’s kind of hard to imagine, but yes, all of these Nobel Prizes, all of these mathematical proofs, all of these conversations, all of these ideas, all the influence we have on each other, even the AI, eventually will expire.
There's ever more pressure to rebuild society more and more around credentials. Do you have this certificate? Do you have that proof? But companies that are focused on just building great products and doing great things gravitate towards people who do the great work.
If you want to argue that human-level AI is extremely unlikely in the next 20 years, you certainly can, but you should treat that as a minority position where the burden of proof is on you.
I couldn't let myself be angry or consumed by that kind of stuff because hate is so sticky, it sticks for a lifetime. And there really is only one cure for hate, which is forgiveness. I just don't think you can get rid of it without that.
For me personally, I kept introducing bugs, and I couldn't figure out why. And what I realized is that I developed... I wasn't copiloting well, I was autopiloting much better.
Re-read because I couldn't interest myself in a couple of other pieces of mind candy and thought "you know who did this sort of thing well...". Which got me thinking, again, about the relationship between this series and feminism.
But I felt like what this paper showed was that the burden of proof is now on the pessimists. So, that's why we called it the pessimism line. Throughout history, there's been alien pessimists and alien optimists, and they've been yelling at each other, that's all they had to go with.
o3's improvement over the GPT series proves that architecture is everything. You couldn't throw more compute at GPT-4 and get these results. Simply scaling up the things we were doing from 2019 to 2023 -- take the same architecture, train a bigger version on more data -- is not enough.
Luis and Walter Alvarez, who made that incredible discovery, initially their discovery was based entirely on impact proxies, just as the Younger Dryas is. There was no crater. And for a long time they were disbelieved because they couldn’t produce a crater.
Despite evaluations, we cannot consider coming powerful frontier AI systems "safe unless proven unsafe". With current testing methodologies, issues can easily be missed. Additionally, it is unclear if governments can quickly build the immense expertise needed for reliable technical evaluations of AI capabilities and societal-scale risks. Given this, developers of frontier AI should carry the burden of proof to demonstrate that their plans keep risks within acceptable limits.
Be skeptical of absolute certainty not backed up with proof. Be skeptical of pronouncements that there’s only one way a court will ever possibly interpret something.
If you had given the Romans the designs for a Newcomen steam engine, they couldn’t have built it without developing whole new technologies for the purpose (or casting every part in bronze, which introduces its own problems) and then wouldn’t have had any profitable use to put it to.
I can't remember when exactly I started looking into David Ogilvy – or where I got this book, probably from a secondhand bookstore somewhere – but once I started reading it I couldn't put it down. Ogilvy had such a great personality. Very spirited, opinionated guy with a great sense of occasion.
But nobody will ever get to the billions of representative miles necessary to say anything compelling about expected safety until after they actually deploy their fleet.
It includes all sorts of thought experiments that all but destroys the concept of self. I couldn’t finish the book because it was too hard and jargony.
But it's just a beautifully told story and also they talk about their writing of the book, so it's very self-referential, which is the whole key to Gödel's proof.
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.