Related posts
Adam Wiggins
Site
Email triage with an embedding-based classifier
labeled embeddings embedding
The the words it uses that this site has seen
least often elsewhere. Posts are matched on those words alone —
nothing here is a summary of this one.
Sources
All
Writing
x.com Mastodon Korrents GitHub Bluesky Site Blog
Top people
Adam Wiggins
Jason Liu
Lilian Weng
Vicki Boykis
Kun Chen
Kieran Healy
Adam Tooze
Kaspars Dambis
Sebastian Raschka
Norman Ohler
Eugene Yan
Drew DeVault
Showing
Profile →
Show everything
Hiding
Show them again
Show them again
Further back ↓
Hiding
Show them again
12 September
Former L8 engineer at Meta, Microsoft and Atlassian, now writing and building solo on agentic engineering. Writes Kun's Field Notes and posts a lot about AI coding agents on X.
every model provider should learn from meta's pricing model for muse spark, with fully transparent subsidization labeled for the contributor tier stop doing "in order to use our model at a good price, you must use our harness which secretly collects your data and/or upload your codebase" meta clearly proved there's a way to collect training data without relying on a harness. just be transparent about when you collect data vs not, clearly communicate how much the data is worth, and give the choice to the user and i really don't need your harness. i already have SO MANY state of the art harness… Related
Sociologist at Duke University writing on markets, exchange and moral order, and on plain-text and data-visualisation tools for social science.
@RecDiffs @siracusa @hotdogsladies That Italian guy pronounced “Samhain" and “Siobhán” really quite well. If you want some cheap entertainment, try a perfectly normal Irish word such as “tábhachtach”, which means "important”. Like your Northern and Southern Italians, the dictionary link provides three pronunciations labeled “C”, “M”, “U”, for Connacht, Munster, and Ulster Irish, the three main provincial accents/dialects. (I'm from Munster so I pronounce it properly.) focloir.ie
Related
7 September
Economic historian at Columbia; writes Chartbook on economics, geopolitics and history, and wrote Crashed and The Deluge.
Their words
People who act and speak in public as right-wing extremists and are labeled as such, establish a claim to authenticity.
Show the whole quote
adamtooze.substack.com
24 August
Latvian full-stack WordPress developer and course creator at WPElevator; blogs since 2007 about the open web, electronics, home automation and sustainable living.
During the weekend I built a contextual search tool for a photo library with 42k images across 2.8k galleries. The results are amazing! Turns out the AI models that generate embedding vectors from images directly are BAD at capturing the image contents, so I had to first generate captions and then generate the embeddings for the captions instead. Total cost: $30 for both captions (gpt-5-mini) and embeddings (text-embedding-3-large with 3072 dimensions). OpenAI Related
16 May
AI research engineer working on large language models. He writes the Ahead of AI newsletter and is the author of Build a Large Language Model (From Scratch).
12 February
Software developer and entrepreneur; co-founder of Heroku, author of The Twelve-Factor App, and a researcher at Ink & Switch.
Two weeks of real-world use on Firehose, my email triage prototype. • 572 emails automatically labeled • 77% coverage (confident labels) • 98.2% accuracy on confident ✨ Related
23 January
Software developer and entrepreneur; co-founder of Heroku, author of The Twelve-Factor App, and a researcher at Ink & Switch.
Can an algorithm learn your life priorities when sorting email? In the first phase of my “personal information firehose” reseach, I trained an embedding-based classifier: adamwiggins.com
Related
17 January
Machine learning engineer working on recommender systems, personalization and information retrieval. Previously worked on LLMs and LLM infrastructure at Mozilla.ai and on ML and recsys at Duo, Tumblr, Automattic and Comcast; wrote the 'What are embeddings?' text and ran the Normconf conference.
19 September 2025
German writer and historian; wrote Blitzed, on drug use in Nazi Germany, and Tripped.
Their words
They make beer. They labeled the beer, and the temple that would make the beer, the beer would be attributed to that temple. It would be sold, so that temple rises in status, makes money. That's how hierarchies started up.
Show the whole quote
youtube.com
17 September 2025
Member of technical staff at Anthropic. He has led ML/AI teams at Amazon, Alibaba and Lazada, and writes about LLMs, recommender systems and engineering at eugeneyan.com.
I've been nerdsniped by the idea of Semantic IDs. Here's the result of my training runs: • RQ-VAE to compress item embeddings into tokens • SASRec to predict the next item (i.e., 4-tokens) exactly • Qwen3-8B that can return recs and natural language! eugeneyan.com
Related
11 September 2025
Machine learning engineer and consultant focused on RAG and retrieval systems. He writes about applied AI engineering at jxnl.co and is the author of the instructor library.
1 September 2025
Machine learning engineer working on recommender systems, personalization and information retrieval. Previously worked on LLMs and LLM infrastructure at Mozilla.ai and on ML and recsys at Duo, Tumblr, Automattic and Comcast; wrote the 'What are embeddings?' text and ran the Normconf conference.
How big are our embeddings now and why? A few years ago, I wrote a paper on embeddings. At the time, I wrote that 200-300 dimension embeddings were fairly common in industry, and that adding more dimensions during training would create diminishing returns for…
Related
20 August 2025
Free software developer; maintainer of sway, wlroots, aerc and scdoc, and founder of the SourceHut (sr.ht) software forge.
3 July 2025
Founded PSPDFKit in 2011 and ran it for a decade. Came back from a break to work on AI agents — the OpenClaw project, and OpenAI, joined in February 2026. Writes at steipete.me.
5 June 2025
Senior economist at the Foundation for American Innovation; writes Second Best, and was the Niskanen Center's director of social policy before that.
19 June 2024
Co-founder and CEO of Perplexity, an AI answer engine; previously a research scientist at OpenAI and a PhD student at UC Berkeley.
Their words
There is an algorithm called BM25 precisely for this, which is a more sophisticated version of TF-IDF. TF-IDF is term frequency times inverse document frequency, a very old-school information retrieval system that just works actually really well even today. And BM25 is a more sophisticated version of that, that is still beating most embeddings on ranking.
Show the whole quote
youtube.com
20 February 2022
Machine-learning researcher; has written the Lil'Log survey posts on how a model technique works since 2017, and worked at OpenAI from 2018 to 2024, latterly leading its safety systems team.
5 December 2021
Machine-learning researcher; has written the Lil'Log survey posts on how a model technique works since 2017, and worked at OpenAI from 2018 to 2024, latterly leading its safety systems team.
31 May 2021
Machine-learning researcher; has written the Lil'Log survey posts on how a model technique works since 2017, and worked at OpenAI from 2018 to 2024, latterly leading its safety systems team.
27 February 2019
General partner at Benchmark Capital for over two decades; was the lead analyst on Amazon's IPO before moving into venture capital. Writes Above the Crowd on tech and marketplace economics.
Nothing matches. Show everything
What is a korrent?
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com .
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
Got it
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.
Got it