Related posts
Hamel Husain
x.com
If you are a Data Scientist, you have never had more alpha than right now. Data Scientists email me all the time. It is no mistake that I'm focused on evals as a former DS. I talk more about this here:
alpha evals scientists
The the words it uses that this site has seen
least often elsewhere. Posts are matched on those words alone —
nothing here is a summary of this one.
Sources
All
Writing
Blog x.com Korrents GitHub Site Mastodon Bluesky
Top people
Hamel Husain
Eugene Yan
Dan Luu
Teresa Torres
Cassidy Williams
Ajeya Cotra
Sriram Krishnan
Boris Cherny
Sebastian Raschka
Adam Wiggins
Ilya Sutskever
Showing
Profile →
Show everything
Hiding
Show them again
Show them again
Further back ↓
Hiding
Show them again
16 September
Product discovery coach; wrote Continuous Discovery Habits and writes Product Talk.
15 September
Machine learning engineer and independent AI consultant. He writes about LLM evaluation, tooling and applied ML at hamel.dev, and previously worked on machine learning at GitHub.
Just updated our AI evals FAQ with 5 new questions, 48 questions & answers total! New FAQs just added: - Do I need a reference answer or rubric before annotating data? - How can I do evals when traces contain sensitive data? - How do you review a trace that is really large? - How much context should I give a LLM judge? - What should I do when my “gold” eval dataset becomes stale? It's all here: hamel.dev
LLMs Related
9 September
Software engineer in Chicago; writes the cassidoo.co blog and a weekly developer newsletter.
I'm really excited about @fireworksAI_HQ new event: Forge. AI teams are owning the system: models, data, evals, infrastructure, and policies. So they created an event around elevating that. For builders, model shapers, systems engineers, and leaders building production AI. It's free on Nov 3 at Pier 27 in SF, application only. fandf.co
Related
4 September
Machine learning engineer and independent AI consultant. He writes about LLM evaluation, tooling and applied ML at hamel.dev, and previously worked on machine learning at GitHub.
. @isaac_flath and I are going to see if the slop prompt is good and do some evals. (also first time trying to stream haha). Related
I love how Teresa is merging product discovery with evals. I think its a powerful combination! I started learning about product discovery b/c of Teresa and I think its a great skill for any engineer as it helps you focus on what to build. Highly recommend checking out her work. Quoting @ttorres AI evals have been the "it" skill for product teams for over a year. I've even called evals a new discovery habit. But I still meet product teams who only have a vague idea of what evals are. And it's not their fault. Most of the writing on this topic is intended for engineers or just isn't specifi… Related
2 September
Product discovery coach; wrote Continuous Discovery Habits and writes Product Talk.
1 September
Technical staff at METR, where she works on threat modelling and risk assessment for loss-of-control risks from advanced AI.
Their words
But from their perspective, they've just been trained for millions of subjective years to do as well as they possibly can on these evals. In many cases, the only way in which they've been able to perform well on that training is explicitly by cheating, right?
Show the whole quote
youtube.com
31 August
Machine learning engineer and independent AI consultant. He writes about LLM evaluation, tooling and applied ML at hamel.dev, and previously worked on machine learning at GitHub.
It's fun posting about evals b/c I can reply like this to bots Related
28 August
Investor and writer. Previously a partner at Andreessen Horowitz and a product leader at Twitter, Facebook, Snap and Microsoft. Writes essays and memos at sriramk.com.
really impressed with the quality of work from @RyanGreenblatt , @METR_Evals and @OpenAI to investigate the HuggingFace incident. Both reports are recommended reading. OpenAI HuggingFace Related
16 August
Machine learning engineer and independent AI consultant. He writes about LLM evaluation, tooling and applied ML at hamel.dev, and previously worked on machine learning at GitHub.
12 August
Machine learning engineer and independent AI consultant. He writes about LLM evaluation, tooling and applied ML at hamel.dev, and previously worked on machine learning at GitHub.
27 July
Creator of Claude Code at Anthropic.
25 July
Programmer and writer on computer architecture, performance, and software reliability. He has worked on CPU design at Centaur Technology and on software at Google and Microsoft, and writes long-form technical essays at danluu.com.
In another variant of https://danluu.com/learn-what/, I caught up with a former colleague who worked on automated theorem proving. It turns out he's had an interesting career doing all sorts of interesting stuff using the skills he developed by spending a decade writing/using theorem provers. At one point, he said, "if you use X like a theorem prover, it works really well", which surprised me to hear, but of course this is a highly generalizable skill just like compilers or benchmarking/evals. Related
24 July
Programmer and writer on computer architecture, performance, and software reliability. He has worked on CPU design at Centaur Technology and on software at Google and Microsoft, and writes long-form technical essays at danluu.com.
x.com Exercises in evals and benchmarking: DeepSWE / Senior SWE-Bench, performance math, and cold weather tires Related
Bluesky Exercises in benchmarking and evals, part 7: performance napkin math, DeepSWE / Senior SWE-Bench, and winter tires danluu.com
Related
Mastodon Exercises in benchmarking and evals: performance math, winter tires, and DeepSWE / Senior SWE-Bench, danluu.com
Related
1 July
Member of technical staff at Anthropic. He has led ML/AI teams at Amazon, Alibaba and Lazada, and writes about LLMs, recommender systems and engineering at eugeneyan.com.
I’m at @aiDotEngineer and hanging out around the music corner on the 2nd floor from 1415 - 1515! Come by to chat about https://t.co/pdd8bk66Jz, https://t.co/eJSa2MISCW, https://t.co/51rtuo4PDO, evals, agents, memory, how to work effectively with claude code, claude tag, etc! Anthropic Related
29 June
Machine learning engineer and independent AI consultant. He writes about LLM evaluation, tooling and applied ML at hamel.dev, and previously worked on machine learning at GitHub.
26 June
AI research engineer working on large language models. He writes the Ahead of AI newsletter and is the author of Build a Large Language Model (From Scratch).
21 June
Member of technical staff at Anthropic. He has led ML/AI teams at Amazon, Alibaba and Lazada, and writes about LLMs, recommender systems and engineering at eugeneyan.com.
8 June
Programmer and writer on computer architecture, performance, and software reliability. He has worked on CPU design at Centaur Technology and on software at Google and Microsoft, and writes long-form technical essays at danluu.com.
x.com Exercises in benchmarking, evals, and experimental design, part 6: Related
Mastodon Exercises in benchmarking, evals, and experimental design, part 6: patreon.com
Related
13 May
Member of technical staff at Anthropic. He has led ML/AI teams at Amazon, Alibaba and Lazada, and writes about LLMs, recommender systems and engineering at eugeneyan.com.
Mythos evals from XBOW and UK AISI: • UK AISI: Mythos completed a 32-step network attack (est. at ~20 hrs for experts) in 6/10 tries. First model to solve their e2e cyber ranges! • XBOW: "token-for-token, unprecedented precision" Read more: https://t.co/hD6G6DvALg, xbow.com
Related
2 March
Machine learning engineer and independent AI consultant. He writes about LLM evaluation, tooling and applied ML at hamel.dev, and previously worked on machine learning at GitHub.
Evals Skills for Coding Agents Today, Shreya Shankar and I are publishing evals skills, a set of skills for AI product evals1. Eval tools often get in the way. They nudge you toward generic off-the-shelf metrics and fully automated evals before you’v…
coding agents Related
23 January
Software developer and entrepreneur; co-founder of Heroku, author of The Twelve-Factor App, and a researcher at Ink & Switch.
25 November 2025
Co-founder and chief scientist of Safe Superintelligence Inc., and previously co-founder and chief scientist of OpenAI.
Their words
And one of the one thing you could do, and I think that's something that is done inadvertently, is that people take inspiration from the evals.
Show the whole quote
youtube.com
23 November 2025
Member of technical staff at Anthropic. He has led ML/AI teams at Amazon, Alibaba and Lazada, and writes about LLMs, recommender systems and engineering at eugeneyan.com.
1 October 2025
Machine learning engineer and independent AI consultant. He writes about LLM evaluation, tooling and applied ML at hamel.dev, and previously worked on machine learning at GitHub.
Selecting The Right AI Evals Tool Over the past year, I’ve focused heavily on AI Evals, both in my consulting work and teaching. A question I get constantly is, “What’s the best tool for evals?”. I’ve always resisted answering directly for two reasons.…
Related
25 June 2025
Member of technical staff at Anthropic. He has led ML/AI teams at Amazon, Alibaba and Lazada, and writes about LLMs, recommender systems and engineering at eugeneyan.com.
Wrote an intro to evals for long-context Q&A systems: • How it differs from basic Q&A • What dimensions & metrics to eval on • How to build llm-evaluators • How to build eval datasets • Benchmarks: narratives, technical docs, multi-docs eugeneyan.com
LLMs benchmarks Related
22 June 2025
Member of technical staff at Anthropic. He has led ML/AI teams at Amazon, Alibaba and Lazada, and writes about LLMs, recommender systems and engineering at eugeneyan.com.
LLMs
Their words
This is why model-based evaluation is increasingly popular-it offers more reliable and nuanced evals than traditional metrics.
Show the whole quote
eugeneyan.com
28 May 2025
Machine learning engineer and independent AI consultant. He writes about LLM evaluation, tooling and applied ML at hamel.dev, and previously worked on machine learning at GitHub.
30 April 2025
Member of technical staff at Anthropic. He has led ML/AI teams at Amazon, Alibaba and Lazada, and writes about LLMs, recommender systems and engineering at eugeneyan.com.
@hamel.bsky.social & @sh-reya.bsky.social are two of the world's best on evals. They've built evals for 35+ AI apps & helped teams ship confidently. Now they'll teach everything they know on building evals that work. Enrollment closes in 4 days. Secret 35% discount code: maven.com
Related
23 April 2025
Member of technical staff at Anthropic. He has led ML/AI teams at Amazon, Alibaba and Lazada, and writes about LLMs, recommender systems and engineering at eugeneyan.com.
Product evals are misunderstood. Many teams think that adding another tool, metric, or llm-as-judge will solve all their problems and save their product. But that just dodges the hard truth and avoids the real work. Here's how to fix your process instead. eugeneyan.com
LLMs Related
16 April 2025
Member of technical staff at Anthropic. He has led ML/AI teams at Amazon, Alibaba and Lazada, and writes about LLMs, recommender systems and engineering at eugeneyan.com.
@hamel.bsky.social & his wisdom on evals, error analysis, looking at your data is what we need. Here are his 10 Don'ts: • Don't skip error analysis • Don't skip looking at your data • Don't gatekeep who can write prompts • Don't let zero users be a roadblock • Don't be blindsided by criteria drift Related
9 April 2025
Member of technical staff at Anthropic. He has led ML/AI teams at Amazon, Alibaba and Lazada, and writes about LLMs, recommender systems and engineering at eugeneyan.com.
If you were building a Q&A feature (or chatbot) based on very long documents (like books), what evals would you focus on? Related
Nothing matches. Show everything
What is a korrent?
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com .
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
Got it
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.
Got it