Related posts
Jason Liu
Site
Text Chunking Strategies for RAG Applications
chunking tips retrieval
The the words it uses that this site has seen
least often elsewhere. Posts are matched on those words alone —
nothing here is a summary of this one.
Sources
All
Writing
x.com Korrents Blog Newsletter Bluesky Site GitHub Recommends
Top people
Eugene Yan
Jason Liu
Kun Chen
Jason Zweig
Ryan Singer
Guillermo Rauch
Craig Hockenberry
Haley Nahman
Luke Wroblewski
Ben Vinegar
Steve Kaufmann
Jamie Brandon
Showing
Profile →
Show everything
Hiding
Show them again
Show them again
Further back ↓
Hiding
Show them again
14 September
Former L8 engineer at Meta, Microsoft and Atlassian, now writing and building solo on agentic engineering. Writes Kun's Field Notes and posts a lot about AI coding agents on X.
sharing an emotional milestone plus some tips for growing a social media presence if you came across my youtube channel, you might have seen i always had “ex-meta” or “L8 principal” mentioned in the title of my videos i hated doing it, and every time i published a video i always a/b tested removing that title unfortunately, every single time, removing it would destroy the reach and the videos simply wouldn’t be seen nearly as much the latest video was the first one i ever had where the “ex-Meta L8” title no longer made any difference. and i have now picked the variant that’s just “with Kun” i… Related
11 September
Investing columnist for The Wall Street Journal and editor of the revised edition of Benjamin Graham's The Intelligent Investor.
My latest @WSJ column: Treasury inflation-protected securities, or TIPS, can help investors in or near retirement to protect their nest egg wsj.com
inflation Related
7 September
Author of Shape Up; formerly head of strategy at Basecamp, where he worked on product design for seventeen years; now runs Felt Presence.
The 7+/-2 rule never fails. I thought I was ready to kickoff something we shaped. But putting it through my documentation process felt hard. Why? Oh yeah. Missing a level of chunking. Suddenly way easier to talk about w/ both the team and agents. documentation Related
3 September
CEO of Vercel; creator of Next.js and Socket.IO. Writes at rauchg.com.
JavaScript
Their words
Improving Next.js chunking can yield massive efficiency improvements at internet scale.
Show the whole quote
@rauchg on X x.com
31 August
Mac and iOS developer at the Iconfactory, where he has worked on apps including Twitterrific and Tot. He writes about development at furbo.org.
21 August
Writer living in Brooklyn and former Man Repeller features director; writes the weekly newsletter Maybe Baby on culture and consumption.
17 August
Product designer and entrepreneur; author of Mobile First and Web Form Design; formerly a product director at Google; now building AI products.
Ask LukeW: A New Retrieval System The Ask LukeW feature on my Web site has been answering people's product design questions using my writings, talks, images, and videos for over three years. During that time, I've seen people ask lots of different kinds…
Related
20 July
Co-founder of Modem and a founding engineer at Sentry, where he went on to be VP of Engineering. Co-author of Third-party JavaScript, and co-host of the State of Agentic Coding podcast with Armin Ronacher.
6 April
Polyglot who speaks twenty languages and co-founded LingQ; argues that comprehensible input, not study, is what does the work.
14 September 2025
Member of technical staff at Anthropic. He has led ML/AI teams at Amazon, Alibaba and Lazada, and writes about LLMs, recommender systems and engineering at eugeneyan.com.
11 September 2025
Machine learning engineer and consultant focused on RAG and retrieval systems. He writes about applied AI engineering at jxnl.co and is the author of the instructor library.
31 May 2025
Programmer based in Vancouver who writes at scattered-thoughts.net about databases, programming language design and developer tools; author of the Imp, Dida and Zest language experiments.
12 April 2025
Member of technical staff at Anthropic. He has led ML/AI teams at Amazon, Alibaba and Lazada, and writes about LLMs, recommender systems and engineering at eugeneyan.com.
Stumbled on the first(?) RAG in NarrativeQA from 2017. Because books & movies were too large for LSTMs to do Q&A on, they embedded 200-word chunks and retrieved similar snippets to answer questions. "Chunking and cosine similarity retrieval is so 2017." arxiv.org
Related
8 April 2025
Member of technical staff at Anthropic. He has led ML/AI teams at Amazon, Alibaba and Lazada, and writes about LLMs, recommender systems and engineering at eugeneyan.com.
Can't wait for when I can vibe code a production recommender system. Until then, here's some system designs: • Retrieval vs. Ranking: eugeneyan.com/writing/syst... • Real-time retrieval: eugeneyan.com/writing/real... • Personalization: eugeneyan.com
Related
21 March 2025
FreeBSD developer and former FreeBSD Security Officer; founder of the Tarsnap online backup service and author of the scrypt key derivation function.
20 January 2025
AI research engineer working on large language models. He writes the Ahead of AI newsletter and is the author of Build a Large Language Model (From Scratch).
16 December 2024
Writes The Pragmatic Engineer, the software-engineering newsletter, and wrote The Software Engineer's Guidebook. Formerly an engineering manager at Uber, and at Microsoft/Skype and Skyscanner before that.
14 December 2024
Coffee consultant and author of books on espresso, roasting and brewing, including The Professional Barista's Handbook and The Coffee Roaster's Companion; writes about brewing gear and roast control at scottrao.com.
ADVANCED HOOPING Now that I’ve used the Ceado Hoop brewer for several months, I have a few tips to share with readers. Although the Hoop produces beautiful extractions more easily than any other brewer, every brewer has a…
Related
29 July 2024
Machine learning engineer and independent AI consultant. He writes about LLM evaluation, tooling and applied ML at hamel.dev, and previously worked on machine learning at GitHub.
An Open Course on LLMs, Led by Practitioners Today, we are releasing Mastering LLMs, a set of workshops and talks from practitioners on topics like evals, retrieval-augmented-generation (RAG), fine-tuning and more. This course is unique because it is: Taught by 25…
LLMs Related
7 July 2024
Member of technical staff at Anthropic. He has led ML/AI teams at Amazon, Alibaba and Lazada, and writes about LLMs, recommender systems and engineering at eugeneyan.com.
19 June 2024
Co-founder and CEO of Perplexity, an AI answer engine; previously a research scientist at OpenAI and a PhD student at UC Berkeley.
Their words
There is an algorithm called BM25 precisely for this, which is a more sophisticated version of TF-IDF. TF-IDF is term frequency times inverse document frequency, a very old-school information retrieval system that just works actually really well even today. And BM25 is a more sophisticated version of that, that is still beating most embeddings on ranking.
Show the whole quote
youtube.com
15 April 2024
Research scientist at Meta in Berlin working on multilingual models and evaluation; led the multilingual team at Cohere and was a research scientist at Google DeepMind before that.
LLMs
Their words
Retrieval-augmented generation (RAG; Lewis et al., 2020), which conditions on the LLM's generation on retrieved documents is the most practical paradigm IMO.
Show the whole quote
Command R+ ruder.io
20 February 2024
Founder and chief executive of Pershing Square Capital Management, the activist investment firm he started in 2004.
Their words
It's really a self-perpetuating board that effectively elects its own members. Once the balance tips, politically, one way or another, it can be kept that way forever. There's no kind of rebalancing system.
Show the whole quote
youtube.com
5 December 2023
Software engineer and a long-time technical lead of the Go programming language; author of plan9port and of the research.swtch.com essays.
29 June 2023
Programmer; founded comma.ai and tinygrad, and was the first to unlock the iPhone.
LLMs
Their words
I think future LLMs are going to be smaller, but are going to run looping on themselves and are going to have retrieval systems. And the thing about using a retrieval system is you can cite sources, explicitly.
Show the whole quote
youtube.com
6 December 2022
Author of Fluent Forever and founder of the app of the same name; teaches pronunciation first and vocabulary through images rather than translation.
22 September 2022
Development economist; writes Global Developments, on growth, history and the numbers behind them.
5 January 2017
Game developer and moral philosopher; made QWOP, GIRP and Getting Over It, and teaches at the NYU Game Center.
7 September 2016
Founding member of OpenAI and former director of AI at Tesla; creator of nanoGPT and the term "vibe coding".
A Survival Guide to a PhD This guide is patterned after my “Doing well in your courses”, a post I wrote a long time ago on some of the tips/tricks I’ve developed during my undergrad. I’ve received nice comments about that guide, so in the same s…
Related
Nothing matches. Show everything
What is a korrent?
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com .
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
Got it
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.
Got it