RSI is poised to make modern LLMs vastly cheaper. Trends that have shown LLMs get exponentially cheaper at a given intelligence are likely to accelerate.
Simple enough for everyone to access Claude's full capabilities. I've been using this experience every day for the last few weeks, and it feels awesome. Simpler, faster, and more powerful.
Inside Anthropic and OpenAI, internal models are improving even faster. Over the summer there was a step change, as Mythos and Astra started kicking off the early stages of recursive self-improvement (RSI).
OpenAI and Anthropic have entered the RSI era. OpenAI Research Progress Is Accelerating Due To OpenAI Research Progress OpenAI has been opening up lately about many things. The most important issue of all is that of the acceleration and automation of AI R&D.
If this is only due to gains in capabilities, that is extremely bad news, and it means CoT monitoring is unlikely to survive for another year unless we find a way to actively improve it, and it might not last six months.
so so yeah, it could be possible now. I think if it's not possible now um it I think it's quite likely to be possible within six months unless there's a dramatic improvement in the security posture
This is very important because in the abstract, I think everyone says, "Yay, productivity. That's good." What did you think, this was all just vibes? No. Productivity means fewer people to do the same number or the same job. Now, the amount of job you want done may increase, and therefore you get more people, but it also may not.
it's really much easier if you can understand how to center a div as they say if you can vertically center a div in HTML then you can probably learn assembly language I would say
uh the we found that that leads to improvement and we saw in the shirt folding example we saw like a quantitative bump from using that sort of imagination compared to not using it. At the same time I think that the model actually performed surprisingly well without that as well.
Coding isn't yet another application domain -- it's the meta-skill required for AI to automatically develop its own training material, via symbolic world models. That's how the RSI loop actually kicks off.
we we kind of have coarse level uh recursive self-improvement already. And the fact that every time you use it, it improves the markdown files. Uh every time you use it, it updates its uh long-term memory.
Take Eroom’s Law: the number of new drugs approved per dollar of R&D has fallen for decades, even as our scientific tools have grown vastly more powerful, the exact opposite of what the existence of more raw capability would predict.
Riders diverted from other lines still benefit from the project, especially in a case like QueensLink, where the diverted riders would enjoy an improvement in trip time to Midtown of about 10-15 minutes each way.
I don't understand why anybody would learn the CAP theorem when that theorem exists because it's just more complete. It's not that much more complicated. I think it's more simple to understand.
we feel this sense of ownership of products that we rely on every day. And we're angry when they change. Even if they change for the better, we're like angry because we go through life faster because of pattern recognition.
I think in a lot of ways the data is actually even more of a constraint and um because if you look at like the scale of these models compared to language models they're smaller but they're smaller because the amount of data is less
the joke I made was pre AI, I would spend 95% of my energy thinking about what to do and 5% of my energy doing it. Now I spend 96% of my time thinking about what to do and 4% of my time actually doing it. So yeah, it's like a 20% improvement, but dayto-day it feels as hard as ever.
I think that, you know, the model that we've put together collectively about the relationships between archaic and modern humans is sort of accreted over time. There was this, you know, idea that modern humans are distinct and that Neanderthals and Denisovans are like sisters of each other. And then over time we developed and detected additional mixture events like this modern human into Neanderthal and then this other ones I didn't even talk about like super divergent lineage going into Denisovans and like all this other stuff.
I don't love the other methods, which is continuous improvement. The problem with continuous improvement, it… First of all, you should engineer something from first principles at the speed, you know, with speed of light thinking. Limit it only by physical limits, and physics limits. And after that, of course you would improve it over time.
if improvement stopped you know here the value of an H100 is now predicated on the value that GPD 5.4 four can get out of it instead of the value that GP4 can get out of it and the margins and all that stuff that these labs are doing and they're in a competitive environment so their margins can't go to infinity. Um so you sort of have this like dynamic that is quite interesting in that an H100 is worth more today than it was 3 years ago.
that work is not being wasted. It's just being stored. You know, it's kind of like complaining about heating an ice cube up a little bit and it not melting yet. It's like, well, you just haven't hit the phase transition. So, but that's where people give up.
the improvement software improvements of really throughput in terms of tokens per dollar per watt that we're able to get uh you know quarter over quarter year over year is massive uh right so it's 5x 10x maybe 40x in some of these cases
we literally at some point had a team that would that was called blockers and they just went and one by one struck them down and each time we saw uh improvement in retention, improvement in activation the metrics for as we addressed each one you could literally see the change in the graph.
I'm more agnostic about the best techniques, things like sparse autoencoders are a useful tool, but easy to waste effort using when a simpler method is sufficient or better - start by doing the obvious thing!
Very rarely do you transform into a simpler problem. So if they can pick up a sense of smell, then they could maybe start competing with a human level of mathematicians.
And for that reason, there's physical constraints to things like AGI, like recursive improvement to kill us all type stuff. For the physical reasons and for how humans have figured things out before, I'm not too worried about AI takeover.
I think my personal definition of AGI is much simpler. I think language models are a form of AGI and all of this super powerful stuff is a next step that's great if we get these tools. But a language model has so much value in so many domains that it's a general intelligence to me.
This "memorize, fetch, apply" paradigm can achieve arbitrary levels of skills at arbitrary tasks given appropriate training data, but it cannot adapt to novelty or pick up new skills on the fly (which is to say that there is no fluid intelligence at play here.)
Their words now
OpenAI's new o3 model represents a significant leap forward in AI's ability to adapt to novel tasks. This is not merely incremental improvement, but a genuine breakthrough, marking a qualitative shift in AI capabilities compared to the prior limitations of LLMs.
o3's improvement over the GPT series proves that architecture is everything. You couldn't throw more compute at GPT-4 and get these results. Simply scaling up the things we were doing from 2019 to 2023 -- take the same architecture, train a bigger version on more data -- is not enough.
My suspicion is that massive improvement in quality of online entertainment crowd out socialising and development of social skills, possibly hurting relationship formation as well as desire for children within marriage.
The cost to discover and develop a drug today is orders of magnitude higher than in the 1950s. Despite this, the probability that a drug entering clinical trials will eventually reach the market has hardly improved in the intervening years.
Any true safety improvement for Tesla is good to have, but is much more likely due to comparison against an "average" vehicle (12 years old in the US) which is much less safe than any recent high-end vehicle regardless of manufacturer, and probably not driven on roads as safe on average as where Teslas are more popular.
Efficiency and invention are sort of at odds, because real invention, Not incremental improvement... Incremental improvement is so important in every endeavor, in everything you do, you have to work hard on also just making things a little bit better. But I'm talking about real invention, real lateral thinking that requires wandering, and you have to give yourself permission to wander.
I suggest using a text-to-image generator like DALL-E 3. However you feel about synthetic media, the pace of improvement is readily apparent from a single prompt
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.