Margins will not climb to 100%. As inference gets cheaper, buyers will ask more of their agents, shifting the equilibrium over time in response to competition.
in this whole USA vs China thing OpenAI and Anthropic aren't relevant because they're positioned differently
them building better models doesn't hurt china at all
the competitor has to be
- american
- open source
- enough compute to do inference at scale
that can shift things
And um if you build a specialized chip for low precision dense linear algebra and can't do anything else that turns out to be really useful for machine learning inference uh even though it can't run Chrome or Word or whatever.
Um, and that's a very very useful general technique is you know inference time compute to perform search over plausible ways of solving the problem that can get much much higher performance or much more reliability in longunning agent flows.
Yeah, I mean it's a little different, but I think uh you're going to see more and more uh uh high performance and um low energy uh inference hardware systems because I think everyone is now realizing that inference is the key to making you know these agent-based systems be available to more and more people and that latency is really important and that specialization of the hardware is a really key way you can make uh things that are more energy efficient and lower latency than more general purpose uh computational devices like say GPUs or TPUs
cuz because we rent GPUs at scale to run the models and we still use middleman by the way. So we're not like going all the way down to the down to the floor. Even for us there are some models the sticker price and the cost to us there's like an 80% margin in there.
There's always negative sentiment that exists for any business that's getting hyped. They have no incentive to correct it. Um so again it's complicated because I know the training costs are a big part of it. Uh the R&D department is is hugely expensive but long-term inference makes sense as a business and I think it it always will.
The demand for inference is growing. So, like I don't think it's linearly growing. I think it might even be exponentially growing. But we haven't made our production of GPUs grow exponentially. That's like kind of a linear process. So as those lines intersect, there's going to be uh tightening.
because if you were to force AI to write a type annotation on everything, then it would probably get it wrong more often because now it has to keep track of all these types and and it and it has to just repeat itself over and over and over, right? And so, types are important where there's no context.
They could release claw slow mode and have an increase in tokens per dollar by a significant amount. Um they could probably like reduce the price of Opus 46 by you know 4x 5x and reduce the speed by another by maybe just like 2x like the curve on inference throughput versus speed is there already just on hm um and yet they don't um because no one actually wants to use a slow model
So when you look at inference at let's say 100 tokens a second for deepseek and kimk 2.5 hopper versus blackwell the performance difference is on the order of 20x
There's There's nothing preventing longer context from working. You just have to train at longer context and then learn to to serve them at inference. And both of those are engineering problems that we are working on and that I would assume others are working on as well.
The economics of orbital “datacenters” or essentially glorified Starlink satellites with a bunch of GPUs attached are likely to be even better than Starlink.
actually it turns out that you can significantly decrease power consumption with a very small reduction in overall compute. So if you if you've got like three really bad days in a row or something, you can actually just like you can dial back your power usage quite a lot without compromising your inference or or um or training.
OpenAI has a fantastic margin. When they're doing inference, their gross margins are north of 75%. So that's a four to five X factor right there of the cost difference, is that OpenAI is just making crazy amounts of money because they're the only one with the capability.
This allows me to run models locally on my MacBook Pro M1 Max. With the 64GB of RAM it has, it’s a pretty potent machine for basic inference despite it being three years old.
I think if we can achieve that amount of inference compute, where it leads to a dramatically better answer as you apply more inference compute, I think that will be the beginning of real reasoning breakthroughs.
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.