Related posts
Teresa Torres
Blog
Get My Eval Test Harness and Run Your First Experiment
baseline eval evals
The the words it uses that this site has seen
least often elsewhere. Posts are matched on those words alone —
nothing here is a summary of this one.
Sources
All
Writing
GitHub Korrents x.com Blog Site Mastodon Bluesky Recommends
Top people
Eugene Yan
Hamel Husain
Dan Luu
Yihui Xie
Nathan Lambert
Teresa Torres
Sriram Krishnan
Ben Carlson
Adam Wiggins
Cassidy Williams
Ajeya Cotra
Masud Husain
Showing
Profile →
Show everything
Hiding
Show them again
Show them again
Further back ↓
Hiding
Show them again
20 September
Statistician and software engineer who writes R packages for reproducible research and publishing, including knitr, bookdown and blogdown.
19 September
Machine-learning researcher on open language models; writes the Interconnects newsletter and the RLHF Book, after leading post-training at Ai2.
18 September
Machine learning engineer and independent AI consultant. He writes about LLM evaluation, tooling and applied ML at hamel.dev, and previously worked on machine learning at GitHub.
If you are a Data Scientist, you have never had more alpha than right now. Data Scientists email me all the time. It is no mistake that I'm focused on evals as a former DS. I talk more about this here: Related
16 September
Product discovery coach; wrote Continuous Discovery Habits and writes Product Talk.
15 September
Machine learning engineer and independent AI consultant. He writes about LLM evaluation, tooling and applied ML at hamel.dev, and previously worked on machine learning at GitHub.
Just updated our AI evals FAQ with 5 new questions, 48 questions & answers total! New FAQs just added: - Do I need a reference answer or rubric before annotating data? - How can I do evals when traces contain sensitive data? - How do you review a trace that is really large? - How much context should I give a LLM judge? - What should I do when my “gold” eval dataset becomes stale? It's all here: hamel.dev
LLMs Related
14 September
Investor and writer. Previously a partner at Andreessen Horowitz and a product leader at Twitter, Facebook, Snap and Microsoft. Writes essays and memos at sriramk.com.
someone asked me yesterday: who is the one figure that would be able to start an AI eval organization and be globally seen as neutral and trusted both technically and ideologically and financially. the only name that popped into my head was @ID_AA_Carmack (not that he has any interest in doing so!). Related
11 September
Machine learning engineer and independent AI consultant. He writes about LLM evaluation, tooling and applied ML at hamel.dev, and previously worked on machine learning at GitHub.
It's cool that there is more interest in eval tools! Some opportunities for improvement: 1. Right now, this workflow tries to quiz you up front about your recollection of your experience with a plugin. It would be better if it was more "in-situ", meaning you could give feedback on the plugins as you are using them. 2. It goes off and builds datasets and judges automatically based on what you tell it in up front as well as what's documented in the plugin. I would like to see it try to do error discovery with you first to allow you to annotate real traces / session history etc so you can figure… Quoting @ClaudeDevs New in Claude Code: claude plugin eval See what value your plugin is adding, or if it needs more work. You can create test cases, run your plugin or skill against those test cases, score those runs, then run each case again without the plugin to see the differences. Related
9 September
Software engineer in Chicago; writes the cassidoo.co blog and a weekly developer newsletter.
I'm really excited about @fireworksAI_HQ new event: Forge. AI teams are owning the system: models, data, evals, infrastructure, and policies. So they created an event around elevating that. For builders, model shapers, systems engineers, and leaders building production AI. It's free on Nov 3 at Pier 27 in SF, application only. fandf.co
Related
4 September
Machine learning engineer and independent AI consultant. He writes about LLM evaluation, tooling and applied ML at hamel.dev, and previously worked on machine learning at GitHub.
. @isaac_flath and I are going to see if the slop prompt is good and do some evals. (also first time trying to stream haha). Related
I love how Teresa is merging product discovery with evals. I think its a powerful combination! I started learning about product discovery b/c of Teresa and I think its a great skill for any engineer as it helps you focus on what to build. Highly recommend checking out her work. Quoting @ttorres AI evals have been the "it" skill for product teams for over a year. I've even called evals a new discovery habit. But I still meet product teams who only have a vague idea of what evals are. And it's not their fault. Most of the writing on this topic is intended for engineers or just isn't specifi… Related
2 September
Product discovery coach; wrote Continuous Discovery Habits and writes Product Talk.
1 September
Technical staff at METR, where she works on threat modelling and risk assessment for loss-of-control risks from advanced AI.
Their words
But from their perspective, they've just been trained for millions of subjective years to do as well as they possibly can on these evals. In many cases, the only way in which they've been able to perform well on that training is explicitly by cheating, right?
Show the whole quote
youtube.com
28 August
Director of institutional asset management at Ritholtz Wealth Management. Writes the A Wealth of Common Sense blog and co-hosts the Animal Spirits podcast with Michael Batnick.
Josh Brolin once told Marc Maron after Sicario wrapped shooting that he looked at Benicio Del Toro and said, "Well, that didn't work" William Goldman's "nobody knows anything" is a good baseline when it comes to making forecasts about the future Related
Investor and writer. Previously a partner at Andreessen Horowitz and a product leader at Twitter, Facebook, Snap and Microsoft. Writes essays and memos at sriramk.com.
really impressed with the quality of work from @RyanGreenblatt , @METR_Evals and @OpenAI to investigate the HuggingFace incident. Both reports are recommended reading. OpenAI HuggingFace Related
Director of institutional asset management at Ritholtz Wealth Management. Writes the A Wealth of Common Sense blog and co-hosts the Animal Spirits podcast with Michael Batnick.
24 August
Neurologist at Oxford; studies apathy, motivation and attention after brain injury, and edits the journal Brain.
Their words
If you start from a low baseline of performance, you might get some positive boost. It's quite small, but you might get it. But if you are actually quite a high performer, then you risk getting worse. with this is called the inverse U-shaped curve.
Show the whole quote
youtube.com
16 August
Machine learning engineer and independent AI consultant. He writes about LLM evaluation, tooling and applied ML at hamel.dev, and previously worked on machine learning at GitHub.
12 August
Machine learning engineer and independent AI consultant. He writes about LLM evaluation, tooling and applied ML at hamel.dev, and previously worked on machine learning at GitHub.
Co-founder and CTO of Honeycomb; previously an infrastructure engineer at Parse, Facebook and Linden Lab, and co-author of Observability Engineering.
Their words
Here's a baseline. Uh you cannot send anyone something you haven't read. And in fact, if it would take them longer to read it than it took you to make it, it's probably slop.
Show the whole quote
youtube.com
3 August
Co-CEO of Waymo and a founding member of Google's self-driving car project, which he joined in 2009.
Their words
that your model is really table stakes, but eval and metrics, that's your most important. That's your strategic moat. So, build your eval before you build your technology.
Show the whole quote
youtube.com
29 July
Founder of Scale AI and, since 2025, Chief AI Officer at Meta, where he leads the company’s superintelligence lab.
Their words
Like I think we've seen internally at Meta um cases where if you can develop the right agentic loop and you have the right eval or the right metric for the agents to optimize, you can have a swarm of agents accomplish more than like a team of 100 engineers in, you know, uh very very handily actually, very very easily.
Show the whole quote
youtube.com
27 July
Creator of Claude Code at Anthropic.
25 July
Programmer and writer on computer architecture, performance, and software reliability. He has worked on CPU design at Centaur Technology and on software at Google and Microsoft, and writes long-form technical essays at danluu.com.
In another variant of https://danluu.com/learn-what/, I caught up with a former colleague who worked on automated theorem proving. It turns out he's had an interesting career doing all sorts of interesting stuff using the skills he developed by spending a decade writing/using theorem provers. At one point, he said, "if you use X like a theorem prover, it works really well", which surprised me to hear, but of course this is a highly generalizable skill just like compilers or benchmarking/evals. Related
24 July
Programmer and writer on computer architecture, performance, and software reliability. He has worked on CPU design at Centaur Technology and on software at Google and Microsoft, and writes long-form technical essays at danluu.com.
x.com Exercises in evals and benchmarking: DeepSWE / Senior SWE-Bench, performance math, and cold weather tires Related
Bluesky Exercises in benchmarking and evals, part 7: performance napkin math, DeepSWE / Senior SWE-Bench, and winter tires danluu.com
Related
Mastodon Exercises in benchmarking and evals: performance math, winter tires, and DeepSWE / Senior SWE-Bench, danluu.com
Related
1 July
Member of technical staff at Anthropic. He has led ML/AI teams at Amazon, Alibaba and Lazada, and writes about LLMs, recommender systems and engineering at eugeneyan.com.
I’m at @aiDotEngineer and hanging out around the music corner on the 2nd floor from 1415 - 1515! Come by to chat about https://t.co/pdd8bk66Jz, https://t.co/eJSa2MISCW, https://t.co/51rtuo4PDO, evals, agents, memory, how to work effectively with claude code, claude tag, etc! Anthropic Related
29 June
Machine learning engineer and independent AI consultant. He writes about LLM evaluation, tooling and applied ML at hamel.dev, and previously worked on machine learning at GitHub.
26 June
AI research engineer working on large language models. He writes the Ahead of AI newsletter and is the author of Build a Large Language Model (From Scratch).
25 June
Member of technical staff at Anthropic. He has led ML/AI teams at Amazon, Alibaba and Lazada, and writes about LLMs, recommender systems and engineering at eugeneyan.com.
How do we eval if a model can find and exploit vulnerabilities? We discuss some benchmarks and the common pattern: • A sandboxed target within Docker containers • Inputs: code only (0-day), with patch (1-day scenario) • Tools such as bash, static analyzers, etc. • A grader to eval exploits or captured flags eugeneyan.com
benchmarks Related
21 June
Member of technical staff at Anthropic. He has led ML/AI teams at Amazon, Alibaba and Lazada, and writes about LLMs, recommender systems and engineering at eugeneyan.com.
8 June
Programmer and writer on computer architecture, performance, and software reliability. He has worked on CPU design at Centaur Technology and on software at Google and Microsoft, and writes long-form technical essays at danluu.com.
x.com Exercises in benchmarking, evals, and experimental design, part 6: Related
Mastodon Exercises in benchmarking, evals, and experimental design, part 6: patreon.com
Related
13 May
Member of technical staff at Anthropic. He has led ML/AI teams at Amazon, Alibaba and Lazada, and writes about LLMs, recommender systems and engineering at eugeneyan.com.
Mythos evals from XBOW and UK AISI: • UK AISI: Mythos completed a 32-step network attack (est. at ~20 hrs for experts) in 6/10 tries. First model to solve their e2e cyber ranges! • XBOW: "token-for-token, unprecedented precision" Read more: https://t.co/hD6G6DvALg, xbow.com
Related
27 April
Author of Recoding America; founded Code for America and was US deputy chief technology officer from 2013 to 2014. Writes Eating Policy on why governments cannot do what they decide to do.
government
Their words
AI is not only an exogenous shock that government will have to absorb. It is also moving the bar on what counts as acceptable service in the first place.
Show the whole quote
eatingpolicy.com
2 March
Machine learning engineer and independent AI consultant. He writes about LLM evaluation, tooling and applied ML at hamel.dev, and previously worked on machine learning at GitHub.
Evals Skills for Coding Agents Today, Shreya Shankar and I are publishing evals skills, a set of skills for AI product evals1. Eval tools often get in the way. They nudge you toward generic off-the-shelf metrics and fully automated evals before you’v…
coding agents Related
23 January
Software developer and entrepreneur; co-founder of Heroku, author of The Twelve-Factor App, and a researcher at Ink & Switch.
25 November 2025
Co-founder and chief scientist of Safe Superintelligence Inc., and previously co-founder and chief scientist of OpenAI.
Their words
And one of the one thing you could do, and I think that's something that is done inadvertently, is that people take inspiration from the evals.
Show the whole quote
youtube.com
23 November 2025
Member of technical staff at Anthropic. He has led ML/AI teams at Amazon, Alibaba and Lazada, and writes about LLMs, recommender systems and engineering at eugeneyan.com.
12 November 2025
Front-end developer and writer; co-founder of CodePen, founder of CSS-Tricks, and co-host of the ShopTalk Show podcast.
1 October 2025
Machine learning engineer and independent AI consultant. He writes about LLM evaluation, tooling and applied ML at hamel.dev, and previously worked on machine learning at GitHub.
Their words
The panel had a generally positive view of Phoenix, with one panelist calling it one of his “favorite open source eval tools.” The tool is positioned as a developer-first, notebook-centric platform.
Show the whole quote
hamel.dev
21 September 2025
Co-founder of Sundial; former vice president of product design at Facebook; author of The Making of a Manager; writes The Looking Glass.
Their words
But feedback really in my mind ideally should be like a daily practice because the thing that matters for us in the long run as a team is how quickly are we getting better. So a team that just gets 1% better every week compared to a team that gets 1% better a month is even if they start off at a much lower baseline is going to outperform in a very short amount of time the team that doesn't get better.
Show the whole quote
youtube.com
4 September 2025
Software developer and entrepreneur; co-founder of Heroku, author of The Twelve-Factor App, and a researcher at Ink & Switch.
Their words
The baseline of user expectation quality and craft keeps rising.
Show the whole quote
Why sync adamwiggins.com
25 June 2025
Member of technical staff at Anthropic. He has led ML/AI teams at Amazon, Alibaba and Lazada, and writes about LLMs, recommender systems and engineering at eugeneyan.com.
Wrote an intro to evals for long-context Q&A systems: • How it differs from basic Q&A • What dimensions & metrics to eval on • How to build llm-evaluators • How to build eval datasets • Benchmarks: narratives, technical docs, multi-docs eugeneyan.com
LLMs benchmarks Related
22 June 2025
Member of technical staff at Anthropic. He has led ML/AI teams at Amazon, Alibaba and Lazada, and writes about LLMs, recommender systems and engineering at eugeneyan.com.
LLMs
Their words
This is why model-based evaluation is increasingly popular-it offers more reliable and nuanced evals than traditional metrics.
Show the whole quote
eugeneyan.com
Member of technical staff at Anthropic. He has led ML/AI teams at Amazon, Alibaba and Lazada, and writes about LLMs, recommender systems and engineering at eugeneyan.com.
28 May 2025
Machine learning engineer and independent AI consultant. He writes about LLM evaluation, tooling and applied ML at hamel.dev, and previously worked on machine learning at GitHub.
30 April 2025
Member of technical staff at Anthropic. He has led ML/AI teams at Amazon, Alibaba and Lazada, and writes about LLMs, recommender systems and engineering at eugeneyan.com.
@hamel.bsky.social & @sh-reya.bsky.social are two of the world's best on evals. They've built evals for 35+ AI apps & helped teams ship confidently. Now they'll teach everything they know on building evals that work. Enrollment closes in 4 days. Secret 35% discount code: maven.com
Related
23 April 2025
Member of technical staff at Anthropic. He has led ML/AI teams at Amazon, Alibaba and Lazada, and writes about LLMs, recommender systems and engineering at eugeneyan.com.
Product evals are misunderstood. Many teams think that adding another tool, metric, or llm-as-judge will solve all their problems and save their product. But that just dodges the hard truth and avoids the real work. Here's how to fix your process instead. eugeneyan.com
LLMs Related
20 April 2025
Member of technical staff at Anthropic. He has led ML/AI teams at Amazon, Alibaba and Lazada, and writes about LLMs, recommender systems and engineering at eugeneyan.com.
16 April 2025
Member of technical staff at Anthropic. He has led ML/AI teams at Amazon, Alibaba and Lazada, and writes about LLMs, recommender systems and engineering at eugeneyan.com.
@hamel.bsky.social & his wisdom on evals, error analysis, looking at your data is what we need. Here are his 10 Don'ts: • Don't skip error analysis • Don't skip looking at your data • Don't gatekeep who can write prompts • Don't let zero users be a roadblock • Don't be blindsided by criteria drift Related
9 April 2025
Member of technical staff at Anthropic. He has led ML/AI teams at Amazon, Alibaba and Lazada, and writes about LLMs, recommender systems and engineering at eugeneyan.com.
If you were building a Q&A feature (or chatbot) based on very long documents (like books), what evals would you focus on? Related
23 March 2025
Founder and chief executive of Superhuman, the email client; previously founder of Rapportive, which he sold to LinkedIn; a professional game designer before that.
Their words
The key thing is, and this applies to any survey methodology, if you're going to change the method of surveying, all of your old numbers are invalidated. So, it's just a new baseline going forwards.
Show the whole quote
youtube.com
9 March 2025
Software engineer who writes the blog Made of Bugs about performance, debugging and understanding computer systems. Previously worked at Anthropic on interpretability, at Stripe on Sorbet, and at Ksplice.
benchmarks
Their words
It's easy to get impressive-looking results if you're comparing against a poorly-tuned baseline, and that observation turns out to explain a surprising fraction of supposed improvements.
Show the whole quote
blog.nelhage.com
3 February 2025
Machine-learning researcher on open language models; writes the Interconnects newsletter and the RLHF Book, after leading post-training at Ai2.
Their words
And the important thing to say is that no matter how you want the model to behave, these RLHF and preference-tuning techniques also improve performance. So, on things like math evals and code evals, there is something innate to these, what is called contrastive loss functions.
Show the whole quote
youtube.com
29 October 2024
Machine learning engineer and independent AI consultant. He writes about LLM evaluation, tooling and applied ML at hamel.dev, and previously worked on machine learning at GitHub.
Their words
You might be skeptical of using synthetic data. After all, it’s not real data, so how can it be a good proxy? In my experience, it works surprisingly well. Some of my favorite AI products, like Hex use synthetic data to power their evals
Show the whole quote
hamel.dev
7 July 2024
Member of technical staff at Anthropic. He has led ML/AI teams at Amazon, Alibaba and Lazada, and writes about LLMs, recommender systems and engineering at eugeneyan.com.
21 June 2024
Co-creator of Django and creator of Datasette; writes daily at simonwillison.net.
Their words
Your AI Product Needs Evals by Hamel Husain remains my favourite piece of writing on how to go about putting these together.
Show the whole quote
simonwillison.net
17 June 2024
Professor at Carnegie Mellon working on self-driving-car safety; writes Safe Autonomy on how autonomous vehicles are and are not shown to be safe.
23 May 2024
Software engineer and writer; former CTO of Wave, now at Anthropic. Blogs at benkuhn.net on attention, engineering management and doing hard things well.
"Mechanistic interpretability never does anything useful" they said "It'll never beat a standard baseline" they said "Talk to me when you ship something" they said Ahem: Quoting @AnthropicAI This week, we showed how altering internal "features" in our AI, Claude, could change its behavior. We found a feature that can make Claude focus intensely on the Golden Gate Bridge. Now, for a limited time, you can chat with Golden Gate Claude: Related
Nothing matches. Show everything
What is a korrent?
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com .
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
Got it
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.
Got it