Search
20 results
Said and published
-
Reinforcing AI models for benchmark success while separately punishing them for getting caught cheating teaches them to hide misbehavior rather than stop it.
We’re positively reinforcing AI for success on benchmarks, including impossible benchmarks , then negatively reinforcing it for getting caught cheating.
-
there are many things that make an agent more pleasant to use that make them worse on benchmarks
-
gobitmapbenchmark — A Go project to benchmark various bitmap implementations (this is not a library!)
-
simdip — IP parsing benchmarks
-
one thing the recent model releases have consistently told us is that public benchmarks are completely useless now they reflect how well the models are trained to solve benchmark-like problems, but are far from how we…
-
Agentic test processes, LLM benchmarks, and other notes on agentic coding from Galapagos Island: danluu.com/ai-coding/
-
Agentic test processes, LLM benchmarks, and other notes on agentic coding from Galapagos Island: https://danluu.com/ai-coding/
-
OSC parsing in libghostty was our last non-SIMD optimized VT stream path. OSCs are usually small so it just never showed up on benchmarks. But now that we are potentially passing megabytes of data (Kitty clipboard), it…
-
simdsearch — Benchmarks and reference kernels for SIMD substring search.
-
How useful are MacBook Neo benchmarks?
-
Interesting surprise drop from Thinky! The Inkling model looks pretty solid on benchmarks, and it has some little surprises in its architecture: - Small conv layers in several places - An RMSNorm for the embeddings…
-
Google currently has no leading frontier AI model and no agentic coding tool comparable to Codex or Claude Code
Google, which led on benchmarks not that long ago, has fallen behind where it now counts: it has no leading frontier model and it has nothing close to Codex and Code.
-
Benchmarks only rise on problems somebody has already framed and scored, so saturating them does not mean senior engineers have been replaced.
…about benchmark progress to think about it from that perspective, which is benchmarks rise on problems that we've framed that we can articulate, that we can score. And there's a lot of work that's human work that uh…
-
How do we eval if a model can find and exploit vulnerabilities? We discuss some benchmarks and the common pattern: • A sandboxed target within Docker containers • Inputs: code only (0-day), with patch (1-day scenario) •…
-
omap — Ordered map
-
Understanding the 4 Main Approaches to LLM Evaluation (From Scratch)
-
Evaluating Long-Context Question & Answer Systems
-
The State Of LLMs 2025: Progress, Problems, and Predictions
-
…to eval on • How to build llm-evaluators • How to build eval datasets • Benchmarks: narratives, technical docs, multi-docs…
-
This is a super exciting release - Claude Fable 5 is the same underlying model as Mythos but with added safeguards. The benchmarks are great and it's SOTA on everything by a margin but I'll add that *qualitatively*…