The jump from Sol to Astra is larger than the jump from Fable 5 to Fable 5.1. This is a big deal. Astra is the best model for what one would broadly call ‘ambitious projects,’ and likely has the highest raw intelligence factor of any model.
Normies don't vibe code, they just ask something like "do my bookkeeping" or "file my tax" or "organize a movie night and send invites" or "generate a flyer for movie night" or "edit my video" They don't ever see code, vibe code, or do anything with code, their AI chat app just does it for them
Well, spatial intelligence eventually must enable us to both generate what the space is, reason within it, and being able to edit and interact within it.
But the problem is these agents are just naturally pretty sloppy and they're naturally pretty spiky in their capability profiles. And so you wouldn't necessarily even if you noticed a weird error that it made, you wouldn't necessarily jump to the conclusion that it was like because of some sort of malign crazy conspiracy.
Then there's another possibility which is that it actually already has worked but just the productivity boost isn't as big as would be obvious if people got 10% more productive. That would still be pretty impressive because it's hard to get a 10% across the board uplift.
Energy is the fundamental input to uh the quality of human life. U if you look backward into time and say how did humanity go from one standard of living to the next, it's always unlocked by cheaper energy, right?
So, you might be surprised to hear that most state-of-the-art foundation models for robotics have no memory or no context. They're just operating on the current sensor observations, the current camera readings, uh, and predicting actions based off of that.
Um, as I speak, the number of new businesses starting on Stripe is up around a bit under, but around 2x year-over-year, um, which again is the largest relative jump, uh, we've seen.
And so if I say, "Hey, make this change." and the agent makes the change and then it runs the test and then they're broken and then it fixes the test. I have very high confidence the next change I asked it to make, it's going to follow that path again
But now with Claude Code on my VPS in the last year, it just live edits on my production server, which sounds like it should go wrong but it just doesn't, it's very careful and only twice in 12 months messed up which meant my site didn't load for 10 seconds which is OK
That was always the idea, and that that goes back even to the predecessor of Turbo Pascal, this idea that it's not just a compiler. It's an experience, right? I mean, you don't just compile your programs. You also edit them. You also run them. You also debug them. You also have a runtime library. It all has to like fit together.
So we just could not see a way this could happen by chance and once we saw that we really felt quite convinced that this was a real signal and that really somehow there has been natural selection to increase the genetic changes that today manifest themselves as more years of school at predicting more years of schooling.
Like I I write books and as much I wish that books as a format will be like will survive. I do think that people don't read books the same way anymore.
But right now I think it's sitting somewhere between half million and five million lines of code, somewhere in there. Probably more on the half million side right now and with the next drop of an Anthropic model, we're probably going to see it jump up to a few million lines.
Extrapolators have a remarkable track record in the AI field, being repeatedly early to trends and capabilities that empiricists believed were still decades away.
But but that's actually a very weak criterion, right? People thought I was saying like we won't need 90% of the software engineers. Those things are worlds apart, right?
if you think about it, it sort of makes sense because the LLM is an averaging machine, right? It's predicting the most likely that's an averaging kind of a thing. And so when what you're looking for is a kind of average, it's usually pretty good. So, summarization, topics, themes. But when you're asking for like what is interesting and not average, it's actually pretty bad at it.
Scrutiny typically either jumps several points on any 10-point scale after an incident that kills a bunch of people, or increases very slowly over decades.
After many years, Sublime Text remains one of my favourite code editors. I still often use it to edit huge files or to make heavy edits through its powerful find and replace feature.
And I kind of feel like the industry it's it's um it's over it's it's making too big of a jump and it's trying to pretend like this is amazing and it's not. It's slop and I think they're not coming to terms with it and maybe they're trying to fund raise or something like that.
Collapse is such a strong word, and the West has been using this word repeatedly. If I remember correctly, maybe four to five times, maybe even six times since 1980s during the period of China’s fastest growth. I tend to not think that Chinese economy will collapse.
As much as possible, try to not predict what the future may hold, but just wait as long as possible for that future to become the present and show you what it actually needs.
When we resist not knowing, we often leap to conclusions, make decisions without all of the information, or simply guess. Of course, that might make it even more likely that we're wrong, the very thing our subconscious is trying to avoid. Needing to be right comes at the expense of curiosity.
it will take time for deep learning systems and training data to mature enough to be useful for the truly value-adding task of predicting safety and efficacy in humans
I'm not a defensive realist, I'm an offensive realist. And my argument is that states look for opportunities to gain more power, and every time they see, or almost every time they see an opportunity to gain more power, and they think the likelihood of success is high and the cost will not be great, they'll jump at that opportunity.
Friedman’s focus on the money supply has not held up, as Samuelson suggested, but the alternative Keynesian macro models recommended by Samuelson in the same interview have not done better and they were not outperforming simple random walk models of predicting the macroeconomic future.
It is actually all their nuances and quirks and slight annoyances that make this relationship worthwhile. I don’t think we’re going to realize that until it’s too late.
This is because the language modeling objective used for many recent large LMs-predicting the next token on a webpage from the internet-is different from the objective "follow the user's instructions helpfully and safely" (Radford et al.,, 2019; Brown et al.,, 2020; Fedus et al.,, 2021; Rae et al.,, 2021; Thoppilan et al.,, 2022). Thus, we say that the language modeling objective is misaligned.
On the iPad, I write using 1Writer or Drafts, and edit podcasts using the Apple Pencil and Ferrite Recording Studio, which is the best podcast editing app I've found on any platform.
I edit podcasts on the Mac in Logic Pro, edit video in Final Cut Pro, and remove noise from audio tracks using iZotope RX 8. I record all my podcasts using Audio Hijack.
I have reviewed before the effectiveness of peer review at figuring out "what's good" in the context of grant awards, finding that peer review as currently practiced does substantially better than chance at predicting future impact, especially for the most impactful of the papers; but at the same time is is far from perfect, leaving plenty of variance unexplained.
I prefer to edit my writing on paper than on a screen, so I do print out longer essays and chapters and make changes by hand. I have an HP OfficeJet printer.
The work of Philip Tetlock, PhD and others has shown that people who have one big idea to explain everything (“hedgehogs” or ideologues) are very bad at accurately modeling the world, predicting outcomes, and recommending effective actions.
ScreenFlow - I believe the future of my business is in video + video courses. I'm in ScreenFlow almost every day. Sometimes I use it to "hijack audio" during interviews. Sometimes I use it to edit video. It's getting a lot of use.
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.