ppll

Related posts

Daniel Lemire x.com

What should be obvious is that inference, running a large language model, is the part that has to be cheap. Most of our hardware was not designed for that. LLM inference is limited by bandwidth. You do a great many matrix-vector multiplications against a huge weight matrix, and you barely reuse those weights. It is closer to streaming a video than to running a simulation or drawing a scene in a game. Graphics processors from Nvidia and others were built for something else. They pileed compute first, then bolted on high-bandwidth memory. It is expensive. An obvious answer is to invert the desi…

LLMsreusematrix

The subject this post names, from the same vocabulary the directory files beliefs under, and the words it uses that this site has seen least often elsewhere. Posts are matched on those words alone — nothing here is a summary of this one.

Nothing else here says these words yet. Try AI, or search the site for it.