ppll
About
Log in
Search
Search
1 result
Said and published
Daniel Lemire
x.com
11 Sept
What
should
be
obvious
is
that
inference
,
running
a large language model,
is
the part
that
has to
be
cheap. Most of our hardware was not designed for
that
. LLM
inference
is
limited by bandwidth. You do a great many…