Related posts
Peter Steinberger Blog
The the words it uses that this site has seen least often elsewhere. Posts are matched on those words alone — nothing here is a summary of this one.
Top people
Showing Profile →
Hiding
Hiding
20 September
19 September
18 September
17 September
-
korrents.com
As AI inference costs fall, rising customer expectations will keep service margins from converging toward 100%.Their words
Margins will not climb to 100%. As inference gets cheaper, buyers will ask more of their agents, shifting the equilibrium over time in response to competition.
15 September
From one piece @paulg on X 2 beliefs · x.com
-
Their words
Someone needs to define the unit, perhaps using a chain of increasingly hard problems, each pair of which can be solved by a single model.
-
korrents.com
Tokens are not the true unit of AI inference, because better models yield more problem-solving per token.Their words
Although you pay for AI by the token, that's not the unit of inference, because you get more problem solving per token as models improve.
13 September
12 September
11 September
8 September
3 September
-
open sourceAnthropicOpenAIChina
Their words
in this whole USA vs China thing OpenAI and Anthropic aren't relevant because they're positioned differently them building better models doesn't hurt china at all the competitor has to be - american - open source - enough compute to do inference at scale that can shift things
23 August
-
korrents.com
Demand for intelligence is highly elastic — every fall in inference cost is met with rapidly growing usage.Their words
the demand for intelligence is highly elastic: as inference costs fall, usage grows rapidly.
20 August
30 July
From one piece Jeff Dean: The 1% Rule for Building in AI 3 beliefs, in the piece's order there
-
Their words
And um if you build a specialized chip for low precision dense linear algebra and can't do anything else that turns out to be really useful for machine learning inference uh even though it can't run Chrome or Word or whatever.
-
Their words
Um, and that's a very very useful general technique is you know inference time compute to perform search over plausible ways of solving the problem that can get much much higher performance or much more reliability in longunning agent flows.
-
Their words
Yeah, I mean it's a little different, but I think uh you're going to see more and more uh uh high performance and um low energy uh inference hardware systems because I think everyone is now realizing that inference is the key to making you know these agent-based systems be available to more and more people and that latency is really important and that specialization of the hardware is a really key way you can make uh things that are more energy efficient and lower latency than more general purpose uh computational devices like say GPUs or TPUs
18 July
-
Their words
a smaller model at a higher reasoning effort can sometimes reach a similar score as a larger model at a lower reasoning effort
Controlling Reasoning Effort in LLMsmagazine.sebastianraschka.com
30 June
15 June
9 June
4 June
27 May
From one piece Building OpenCode with Dax Raad 3 beliefs, in the piece's order there
-
Their words
cuz because we rent GPUs at scale to run the models and we still use middleman by the way. So we're not like going all the way down to the down to the floor. Even for us there are some models the sticker price and the cost to us there's like an 80% margin in there.
-
Their words
There's always negative sentiment that exists for any business that's getting hyped. They have no incentive to correct it. Um so again it's complicated because I know the training costs are a big part of it. Uh the R&D department is is hugely expensive but long-term inference makes sense as a business and I think it it always will.
-
Their words
The demand for inference is growing. So, like I don't think it's linearly growing. I think it might even be exponentially growing. But we haven't made our production of GPUs grow exponentially. That's like kind of a linear process. So as those lines intersect, there's going to be uh tightening.
25 May
22 May
13 May
-
Their words
because if you were to force AI to write a type annotation on everything, then it would probably get it wrong more often because now it has to keep track of all these types and and it and it has to just repeat itself over and over and over, right? And so, types are important where there's no context.
23 March
-
Their words
that was always illogical to me because inference is thinking, and I think thinking is hard. Thinking is way harder than reading.
13 March
From one piece Dylan Patel — The single biggest bottleneck to scaling AI compute 2 beliefs, in the piece's order there
-
Their words
They could release claw slow mode and have an increase in tokens per dollar by a significant amount. Um they could probably like reduce the price of Opus 46 by you know 4x 5x and reduce the speed by another by maybe just like 2x like the curve on inference throughput versus speed is there already just on hm um and yet they don't um because no one actually wants to use a slow model
-
Their words
So when you look at inference at let's say 100 tokens a second for deepseek and kimk 2.5 hopper versus blackwell the performance difference is on the order of 20x
13 February
-
korrents.com
Longer context is an engineering and inference problem, not a research problem — nothing prevents it from working.Their words
There's There's nothing preventing longer context from working. You just have to train at longer context and then learn to to serve them at inference. And both of those are engineering problems that we are working on and that I would assume others are working on as well.
10 February
-
Their words
The economics of orbital “datacenters” or essentially glorified Starlink satellites with a bunch of GPUs attached are likely to be even better than Starlink.
Space AI: I guess we’re doing Moon factories nowcaseyhandmer.wordpress.com
24 January
30 December 2025
4 December 2025
15 August 2025
-
Their words
actually it turns out that you can significantly decrease power consumption with a very small reduction in overall compute. So if you if you've got like three really bad days in a row or something, you can actually just like you can dial back your power usage quite a lot without compromising your inference or or um or training.
3 February 2025
-
Their words
OpenAI has a fantastic margin. When they're doing inference, their gross margins are north of 75%. So that's a four to five X factor right there of the cost difference, is that OpenAI is just making crazy amounts of money because they're the only one with the capability.
30 January 2025
-
Likedrcmnd.app
MacBook Pro M1 MaxTheir words
This allows me to run models locally on my MacBook Pro M1 Max. With the 64GB of RAM it has, it’s a pretty potent machine for basic inference despite it being three years old.
31 December 2024
-
Recommendsrcmnd.app
Is AI progress slowing down?Their words
To understand more about inference scaling I recommend Is AI progress slowing down?
6 August 2024
19 June 2024
-
Their words
I think if we can achieve that amount of inference compute, where it leads to a dramatically better answer as you apply more inference compute, I think that will be the beginning of real reasoning breakthroughs.
10 January 2023
Nothing matches.
What is a korrent?
A korrent is a belief a person has stated in their own words: one sentence stating the claim, backed by a quote and a source, kept at korrents.com.
Under a name here, the quoted block is what they actually said. The korrent beneath it is the claim those words support, in korrents' wording — tap it to see the record, its source, and who else holds it.
Nobody here wrote their own korrents. They are compiled from public statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they do, this site shows a machine translation beneath the post, in this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the serif above is exactly what the person published, and it is what to quote them on. A translation can be wrong in ways that matter, especially about tone.
Only the post's own words are translated. A quoted post, a linked article and a belief on korrents.com are left in their original language.