Eugene Yan
Everything, newest first — across every channel. Their profile →
Filter & sortAll sources · condensed
- Fable 5.1 is a thoughtful collaborator, thinking hard about my requests, proactively patching my blindspots, and verifying the work's correct without being ask… 1 Sept Quoting@claudeaiWe’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They're the world’s most advanced models for coding and knowledge work.
- When evaling models, we anchor on the median task. But this is like how devs estimate the median task accurately but underestimate the mean which tends to be ~… 22 Jul Quoting@Steve_YeggeFable is careful. None of the other models are careful. GPT-5.6 Sol, Opus, Kimi, Grok. You can compare them all day long on capabilities, a…
- I’m at @aiDotEngineer and hanging out around the music corner on the 2nd floor from 1415 - 1515! Come by to chat about https://t.co/pdd8bk66Jz, https://t.co/eJ… 1 Jul
- How do we eval if a model can find and exploit vulnerabilities? We discuss some benchmarks and the common pattern: • A sandboxed target within Docker container… 25 Jun
- I've been loving the multiplayer form factor of Claude Tag. Now others can reply on the thread to provide context and direction to Claude! 23 Jun Quoting@claudeaiIntroducing Claude Tag, a new way for teams to work with Claude. In Slack, Claude joins as a team member with access to the channels and to…
- eugeneyan 22 Jun
- Patterns for Building Cybersecurity Evals 21 Jun
- Stronger models have made finding vulnerabilities easier, and the bottleneck has shifted to verification, triage, patching. Here are some lessons from working… 3 Jun
- Using LLMs to Secure Source Code 27 May
- Cloudflare on their vulnerabilty discovery harness • Recon: Read the codebase, return an architecture doc • Hunt: ~50 agents look for bugs concurrently • Valid… 18 May
- Claude Mythos Preview case studies (also, read your transcripts!) https://t.co/drNlAH5mLE > "Mythos demonstrates its bug reproduction and exploitation capabili… 15 May
- Mythos evals from XBOW and UK AISI: • UK AISI: Mythos completed a 32-step network attack (est. at ~20 hrs for experts) in 6/10 tries. First model to solve thei… 13 May
- Mozilla fixing more bugs in April that the past 15 months is what riding the exponential looks like 🎢 8 May
- you get more compute. you get more compute. you get more compute! Anthropic 🤝 SpaceX 6 May Quoting@claudeaiEffective today, we are: 1) Doubling Claude Code’s 5-hour rate limits for Pro, Max, and Team plans; 2) Removing the peak hours limit reduct…
- some thoughts on working with ai models • context as infra • taste as config • verification for autonomy • scaling via delegation • closing the loop 6 May
- How to Work and Compound with AI 3 May
- excited to share this and help everyone be more cyber secure 30 Apr Quoting@claudeaiClaude Security is now in public beta for Claude Enterprise customers. Claude scans your codebase for vulnerabilities, validates each findi…
- Great writeup by @mozilla: Mythos found 271 vulns (fixed in Firefox 150); Opus 4.6 found 22 (fixed in Firefox 148) https://t.co/wzTCxmTKbe > "So far we’ve foun… 21 Apr
- Cheng’s experiment is cool! Especially this: “instead of starting with easy sudoku puzzles data and gradually ramping up toward harder ones, the training runs… 14 Mar Quoting@_chenglouI’m very happy to present my toy research project: Sotaku! It's a neural net that automatically discovered the rules of sudoku and learned…
- yay for automatic prefix caching! just set a single cache control field at the top level of your request body. note that you do still need to structure your pr… 20 Feb
- Sonnet 4.6 is a big upgrade over 4.5. And it has 1M context! It's also versatile across classification, coding, computer use, autonomous agents by adjusting ef… 18 Feb Quoting@claudeaiThis is Claude Sonnet 4.6: our most capable Sonnet model yet. It’s a full upgrade across coding, computer use, long-context reasoning, agen…
- 2025 Year in Review 14 Dec 2025
- Product Evals in Three Simple Steps 23 Nov 2025
- Advice for New Principal Tech ICs (i.e., Notes to Myself) 19 Oct 2025
- semantic-ids-llm — Semantic IDs: How to train an LLM-Recommender Hybrid with steerability and reasoning on recommendations. 15 Sept 2025
- Training an LLM-RecSys Hybrid for Steerable Recs with Semantic IDs 14 Sept 2025
- news-agents — 📰 Building News Agents to Summarize News with MCP, Q, and tmux 19 Jul 2025
- Evaluating Long-Context Question & Answer Systems 22 Jun 2025
- AI Engineer 2025 - Improving RecSys & Search with LLM techniques 4 Jun 2025
- Exceptional Leadership: Some Qualities, Behaviors, and Styles 18 May 2025
- Building News Agents for Daily News Recaps with MCP, Q, and tmux 4 May 2025
- An LLM-as-Judge Won't Save The Product—Fixing Your Process Will 20 Apr 2025
- Frequently Asked Questions about My Writing Process 30 Mar 2025
- NVIDIA GTC 2025 - Building LLM-Powered Applications 18 Mar 2025
- Improving Recommendation Systems & Search in the Age of LLMs 16 Mar 2025
- open-llms — 📋 A list of open LLMs available for commercial use. 13 Feb 2025
- Building AI Reading Club: Features & Behind the Scenes 12 Jan 2025
- 2024 Year in Review 22 Dec 2024
- Seemingly Paradoxical Rules of Writing 1 Dec 2024
- How to Run a Weekly Paper Club (and Build a Learning Community) 24 Nov 2024
- My Minimal MacBook Pro Setup Guide 17 Nov 2024
- Warp 17 Nov 2024 Usesaffiliate link
- Fish 17 Nov 2024 Uses
- Rectangle 17 Nov 2024 Mixed on
- Raycast 17 Nov 2024 Uses
- MX Ergo 17 Nov 2024 Uses
- align-app 9 Nov 2024
- framework-comparison 9 Sept 2024
- llm-paper-notes — Notes from the Latent Space paper club. Follow along or start your own! 31 Jul 2024
- applied-ml — 📚 Papers & tech blogs by companies sharing their work on data science & machine learning in production. 18 Jul 2024
- papermill-mlflow — 🧪 Simple data science experimentation & tracking with jupyter, papermill, and mlflow. 9 Jul 2024
- recsys-nlp-graph — 🛒 Simple recommender with matrix factorization, graph, and NLP. Beating the regular collaborative filtering baseline. 7 Jul 2024
- discord-llm — Experimenting with LLMs to Research, Reflect, and Plan (LLM assistants, retrieval, and Discord integration) 7 Jul 2024
- obsidian-copilot — 🤖 A prototype assistant for writing and thinking 30 Jun 2024
- workspace-testing 24 Jun 2024
- poc-docker-template — Simple template showing how to set up docker for reproducible data science with Jupyter notebooks. 17 Jun 2024
- applyingml — 📌 Papers, guides, and mentor interviews on applying machine learning for ApplyingML.com—the ghost knowledge of machine learning. 5 Jun 2024
- visualizing-finetunes 27 May 2024
- 1-on-1s — 🌱 1-on-1 questions and resources from my time as a manager. 11 May 2024
- python-collab-template — 🛠 Python project template with unit tests, code coverage, linting, type checking, Makefile wrapper, and GitHub Actions. 14 Apr 2024
- Roam 13 Sept 2020 Uses
- Bear 13 Sept 2020 Uses
- MacDown 13 Sept 2020 Recommends
- Kindle 13 Sept 2020 Uses
- Adobe Acrobat Reader 13 Sept 2020 Uses
- Instapaper 13 Sept 2020 Uses