Margins will not climb to 100%. As inference gets cheaper, buyers will ask more of their agents, shifting the equilibrium over time in response to competition.
in this whole USA vs China thing OpenAI and Anthropic aren't relevant because they're positioned differently
them building better models doesn't hurt china at all
the competitor has to be
- american
- open source
- enough compute to do inference at scale
that can shift things
there are some people who have pure apathy, no depression like David. And there are other people who can be depressed, they can be sad, they can be hopeless about the future, but they don't necessarily lose motivation. They can have pure depression.
Structure and Interpretation of Computer Programs (SICP) by Gerald Jay Sussman and Hal Abelson. SICP is my all-time favourite programming book. It's also where I learned about the power of wishful thinking when coding. SICP uses wishful thinking as a design tool: write code as if the abstraction already existed. Pretend. Then, once the ideal abstraction has taken shape, you go ahead and implement those functions. This meta-level is what made SICP so valuable. Ultimately, it's a book that teaches how to think about code and problem solving. I'm tempted to even say that SICP is more a work of art than a pure coding book, but I won't go there. A beautiful book.
there are people who are really big on road maps and there are people who are really big on iteration. I I think the honest answer is you've got to do both. You can't over-index on either.
And um if you build a specialized chip for low precision dense linear algebra and can't do anything else that turns out to be really useful for machine learning inference uh even though it can't run Chrome or Word or whatever.
Um, and that's a very very useful general technique is you know inference time compute to perform search over plausible ways of solving the problem that can get much much higher performance or much more reliability in longunning agent flows.
Yeah, I mean it's a little different, but I think uh you're going to see more and more uh uh high performance and um low energy uh inference hardware systems because I think everyone is now realizing that inference is the key to making you know these agent-based systems be available to more and more people and that latency is really important and that specialization of the hardware is a really key way you can make uh things that are more energy efficient and lower latency than more general purpose uh computational devices like say GPUs or TPUs
Unfortunately, no, because politics is about people who disagree with you. If you’re working with computers, or robots, or pure math, you don’t have politics.
So if you do a pure chronological feed, the incentive for everybody is to just post as much as possible uh because it will always be at the top of everyone who follows you's feed as soon as you post.
cuz because we rent GPUs at scale to run the models and we still use middleman by the way. So we're not like going all the way down to the down to the floor. Even for us there are some models the sticker price and the cost to us there's like an 80% margin in there.
There's always negative sentiment that exists for any business that's getting hyped. They have no incentive to correct it. Um so again it's complicated because I know the training costs are a big part of it. Uh the R&D department is is hugely expensive but long-term inference makes sense as a business and I think it it always will.
The demand for inference is growing. So, like I don't think it's linearly growing. I think it might even be exponentially growing. But we haven't made our production of GPUs grow exponentially. That's like kind of a linear process. So as those lines intersect, there's going to be uh tightening.
because if you were to force AI to write a type annotation on everything, then it would probably get it wrong more often because now it has to keep track of all these types and and it and it has to just repeat itself over and over and over, right? And so, types are important where there's no context.
I mean, some problems have been basically solved by pure brute force. The four color theorem is is a famous example. Um, we have still not found a conceptually elegant proof of this theorem.
Um we are seeing a lot fewer sort of pure AI solutions now where um they are just one shots the problem. Um so so there was there was a month where that happened and and that has stopped.
players who are also commentators give better commentary than people who are just commentators. Stock analysts who never invest have zero skin in the game.
They could release claw slow mode and have an increase in tokens per dollar by a significant amount. Um they could probably like reduce the price of Opus 46 by you know 4x 5x and reduce the speed by another by maybe just like 2x like the curve on inference throughput versus speed is there already just on hm um and yet they don't um because no one actually wants to use a slow model
So when you look at inference at let's say 100 tokens a second for deepseek and kimk 2.5 hopper versus blackwell the performance difference is on the order of 20x
There's There's nothing preventing longer context from working. You just have to train at longer context and then learn to to serve them at inference. And both of those are engineering problems that we are working on and that I would assume others are working on as well.
The economics of orbital “datacenters” or essentially glorified Starlink satellites with a bunch of GPUs attached are likely to be even better than Starlink.
Jonathan Livingston Seagull by Richard Bach is a beautiful fable about a seagull who wants to fly for the pure joy of flight, rather than simply to find food, as most gulls do. Ostracized from his flock for his eccentric behavior, Jonathan embarks on a quest for aerodynamic perfection that doubles as a moving spiritual allegory. You've never read a book quite like this one before, and it'll stick with you long after you reach the end.
I want to give something that's good and pure and true and is like is like the thing that will once it's said there's nothing else to say. It's like done.
And when you link two together with the cable and if I got four, it would push yours forward. It was like the perfect game design. So from a pure puzzle perspective, nothing comes close.
actually it turns out that you can significantly decrease power consumption with a very small reduction in overall compute. So if you if you've got like three really bad days in a row or something, you can actually just like you can dial back your power usage quite a lot without compromising your inference or or um or training.
I’ve rarely seen a more capitalist society than China, from the pure economic side. I’ve rarely seen companies that are as competitive as Chinese companies. People as ambitious and obsessed with making money as Chinese people. Kind of ruthless actually. And look, consumers shop, firms invest. If you invest well financially, you’ll get great returns. What is not capitalistic about the Chinese economy? At the same time, the social fabric is highly socialist.
You should not try to create a ladder that functions as a pure checklist or scorecard that guarantees promotion if people check enough boxes; promotions are as much about the needs of the organization as the skills of the employees.
OpenAI has a fantastic margin. When they're doing inference, their gross margins are north of 75%. So that's a four to five X factor right there of the cost difference, is that OpenAI is just making crazy amounts of money because they're the only one with the capability.
This allows me to run models locally on my MacBook Pro M1 Max. With the 64GB of RAM it has, it’s a pretty potent machine for basic inference despite it being three years old.
And Friedman advanced a policy that was really abstract and really it was devoid of judgment. It was pure and beautiful in that way but utterly impractical.
but I'm very sorry to break it to all of you, but your problem is not unique. I'm 99% sure of that. Your problem has been felt by somebody somewhere and probably many, many places
I think if we can achieve that amount of inference compute, where it leads to a dramatically better answer as you apply more inference compute, I think that will be the beginning of real reasoning breakthroughs.
most macros don't actually do anything other than pure transformation of their arguments. It's almost always safe to run them during static analysis, except you don't know for sure
One of the reasons I've never cared for crypto currencies is that the associated utopia of trustless society had zero appeal to me. I don't think the world is better off by erasing the need to trust in our transactional counterparts, so turning these transactions into pure computing always struck me as a regression.
Their words now
But wherever this leads us next, it's clear to me now that I was too hasty to completely dismiss crypto on the basis of all the things wrong with it at the moment. Instead of appreciating the fundamental freedom to transact that it's currently our best shot at protecting.
If you're looking to understand the basics of building and scaling a sales process and team, this book is a great primer. It also goes into great detail in how the author built a better, more reliable method than pure cold-calling to generate leads beyond the marketing team at Salesforce's efforts.
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.