yay for automatic prefix caching! just set a single cache control field at the top level of your request body. note that you do still need to structure your prompt templates to benefit from prefix caching though.
Filesystem reads in that case are 10 nanoseconds at most (the filesystem cache helps so much here) but doing any network roundtrip is 10 milliseconds at minimum. It's at least a million times slower because of how reality works.
That's a hell of a lot slower than knowing it in cognitive L1 cache. And you can have way more round trips in your brain than you can, you know, muttering through, you know, super whisper or typing it out or whatever.
As reasoning models and agent workflows keep more tokens around (for longer), KV-cache size, memory traffic, and attention cost quickly become the main constraints, and LLM developers are adding a growing number of architecture tricks to reduce those costs.
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.