Related posts
Eugene Yan
x.com
yay for automatic prefix caching! just set a single cache control field at the top level of your request body. note that you do still need to structure your prompt templates to benefit from prefix caching though.
caching templates cache
The the words it uses that this site has seen
least often elsewhere. Posts are matched on those words alone —
nothing here is a summary of this one.
Sources
All
Writing
GitHub Mastodon x.com Korrents Recommends Blog
Top people
Sebastian Raschka
Joe Hewitt
Brad Fitzpatrick
Mark Nottingham
Ray Dalio
Xe Iaso
Armin Ronacher
Ismail Ghallou
Wes Bos
Jason Liu
Manassarn "Noom" Manoonchai
Dillon Mulroy
Showing
Profile →
Show everything
Hiding
Show them again
Show them again
Further back ↓
Hiding
Show them again
19 September
Programmer known for LiveJournal, memcached, OpenID and gearman; worked on the Go programming language at Google and later joined Tailscale.
18 September
Web infrastructure engineer and standards editor; chaired the IETF HTTP working group and co-authored core HTTP and Atom specifications.
16 September
Founded Bridgewater Associates in 1975 and led it until stepping back from control in 2022. Author of Principles, The Changing World Order and How Countries Go Broke; publishes economic analysis at economicprinciples.org.
As I’ve explained here before, at this stage in my life, my main goal is to pass along the principles and templates I've learned with the hope that they can help others as much as they've helped me. One thing that has been incredibly valuable to me is learning that just because something hasn’t happened before in my lifetime doesn’t mean it won’t happen. So I needed to look not only at what the markets have been doing lately but to also study big events—like debt crises, internal political conflicts, and international wars—throughout history. For example, it was by studying the Great Depressi… wmi.edu.sg
Related
15 September
Technical educator, conference speaker and developer relations engineer based in Ottawa, Canada. Author of the Anubis bot filter and of over 400 articles at xeiaso.net.
Their words
Filesystem reads in that case are 10 nanoseconds at most (the filesystem cache helps so much here) but doing any network roundtrip is 10 milliseconds at minimum. It's at least a million times slower because of how reality works.
Show the whole quote
tigrisdata.com
Creator of Flask and Jinja. Writes about software at lucumr.pocoo.org.
Their words
I love it when tool calls take so long on Fable that afterwards you pay for it in cache misses.
Show the whole quote
x.com
13 September
Moroccan senior frontend developer writing at smakosh.com; builds side projects and client work under Smakosh LLC.
New on @llmgateway: Adaptive Cache-Aware Provider Selection Routing learns cache-hit rates and output-to-input proportions from recent project/model usage to compare providers for large prompts and sessions. Available automatically across plans, with workload defaults before enough history exists and explicit overrides on Enterprise. Learn more llmgateway.io
Related
11 September
Canadian full-stack web developer who makes JavaScript and CSS video courses and co-hosts the Syntax podcast.
Genie Garage doors enabled caching on their API resulting in everyones garage door being shared with everyone else using the API Cache Ruins Everything Around Me Quoting @IAmArcIvanov @TheGenieCompany Your AlladinConnect is either breached or malfunctioning where users using API integration (in my case via HASS) are seeing garage doors of other people's accounts. The number of the devices populated by the API endpoint keeps growing as of Sept 6th 1:42 AM. Related
9 September
Machine learning engineer and consultant focused on RAG and retrieval systems. He writes about applied AI engineering at jxnl.co and is the author of the instructor library.
instructor 1.17 is out cache isolation fixes, better Gemini retries, local PDF support, OpenAI SDK 3 compatibility, and refreshed model examples. lots of small fixes from the issues and PR backlog. thanks to everyone who sent one in github.com
OpenAI Google Related
7 September
Thai software engineer; publishes a public digital garden at garden.narze.live and builds developer and productivity tooling.
6 September
AI research engineer working on large language models. He writes the Ahead of AI newsletter and is the author of Build a Large Language Model (From Scratch).
Reasoning from scratch round 2: In this video, I cover the text generation process in LLMs and KV caching (to prepare the base model before adding reasoning techniques in the upcoming ones). 00:00 Introduction and reasoning model demo 01:55 How to work through the book 05:00 Chapter 2 overview 08:25 Checking PyTorch and hardware support 10:26 Apple silicon and MPS caveats 15:00 Cloud GPU options 16:08 Tokens and tokenization 18:20 Qwen3 and the Reasoning From Scratch package 23:05 Encoding and decoding text 26:24 Downloading weights and selecting a device 31:01 Loading the pretrained Qwen3 mo… LLMs Related
5 September
Principal engineer at Cloudflare as of 2026, on agents and developer tooling. Streams TypeScript, OCaml and Neovim on Twitch, and wrote the better-result library after writing the same Result type about a hundred times.
put cloudflare, workers, and workers cache in front and sleep well at night Quoting @RhysSullivan using github as a reference, budgeting for 10x-100x volume compounding how the hell is every service going to handle that? one option is everyone runs their own caches in front of the APIs they use, but then that becomes a data access and freshness nightmare Related
1 September
Member of technical staff at Anthropic. He has led ML/AI teams at Amazon, Alibaba and Lazada, and writes about LLMs, recommender systems and engineering at eugeneyan.com.
Fable 5.1 is a thoughtful collaborator, thinking hard about my requests, proactively patching my blindspots, and verifying the work's correct without being asked. And with cache reads now costing 75% less, to $0.25/M tokens, huge savings for long-running, agentic tasks! Quoting @claudeai We’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They're the world’s most advanced models for coding and knowledge work. Related
29 August
Founder and CEO of Get Lighthouse, a management-coaching software company, and previously a product lead at KISSmetrics. Writes about leadership, management and product at jasonevanish.com.
Common problem… Helps to have templates and skills you can share that either clean it up for you, or politely calls out gaps in what they were going to do. Quoting @lennysan A growing part of everyone’s job is cleaning up the AI slop from other people trying to do your job Related
Philosophy professor at UC Berkeley and creator of pandoc, the universal document converter, and of the CommonMark markdown specification.
24 August
Mac and iOS developer; created NetNewsWire and MarsEdit, co-created the JSON Feed format, and blogs at inessential.com.
31 July
Co-founder and CEO of Stripe and co-founder of the Arc Institute. Keeps a personal site of reading lists, open questions and notes on scientific progress and how research is funded.
Their words
That's a hell of a lot slower than knowing it in cognitive L1 cache. And you can have way more round trips in your brain than you can, you know, muttering through, you know, super whisper or typing it out or whatever.
Show the whole quote
youtube.com
25 July
Computer science professor who works on fast data processing; co-author of the simdjson parser and a weekly blogger about software performance since 2004.
Memory-level parallelism: AMD is the king When your program asks for memory that is not in cache, the processor has to go to RAM. That trip costs on the order of 100 nanoseconds. On a 3 GHz core, that is about 300 cycles of doing nothing. Memory latency has not…
Related
19 July
Chief Product and Technology Officer at Netflix, previously VP of Science at Lyft and an economist at Analysis Group.
Their words
I get really nervous about having different design languages or different types of user interactions and shipping Frankensteins, basically. So, designers need to then be the people we're hiring again for design systems thinking.
Show the whole quote
youtube.com
12 July
Sociologist at Duke University writing on markets, exchange and moral order, and on plain-text and data-visualisation tools for social science.
16 May
AI research engineer working on large language models. He writes the Ahead of AI newsletter and is the author of Build a Large Language Model (From Scratch).
5 December 2025
Scottish software developer on the Statamic core team; runs a one-person business maintaining Statamic addons such as Runway and writes at duncanmcclean.com.
9 November 2025
Egyptian front-end developer and course creator; writes for CSS-Tricks, publishes courses on Udemy and blogs at alialaa.dev.
24 April 2025
Novelist and writer, author of Mr. Penumbra's 24-Hour Bookstore and Sourdough; previously worked at Twitter and Current TV.
12 August 2024
Programmer and writer; co-wrote The Rust Programming Language and worked on Rust's documentation for years.
16 July 2024
Engineer and teacher; wrote Kubernetes the Hard Way, was a distinguished engineer at Google Cloud, and retired from full-time work in 2023.
8 June 2021
Front-end developer and writer; co-founder of CodePen, founder of CSS-Tricks, and co-host of the ShopTalk Show podcast.
27 December 2019
Engineer and writer on machine-learning systems; author of Designing Machine Learning Systems and AI Engineering. "I work to bring AI into production. I write about AI system design."
13 May 2014
Software developer; created Firebug, helped create Firefox on the Netscape browser team, and built the original Facebook iPhone app.
11 April 2012
Software developer; created Firebug, helped create Firefox on the Netscape browser team, and built the original Facebook iPhone app.
7 February 2012
Web developer and musician; co-creator of the Django web framework, founder of EveryBlock and of the music-notation site Soundslice.
Nothing matches. Show everything
What is a korrent?
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com .
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
Got it
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.
Got it