Search
120 results — nothing had every word, so these match some of them
People
Matching some of those words
-
dabstep
-
New benchmark dropped
-
Fresh data from GitHub: Agent-generated PRs have exploded in size. 9x the last 8 months (!!) No signs of slowing down. This is why everyone is rethinking code reviews, deploys, possibly even o11y thanks to the GH team…
-
the only cli benchmark that matters
-
the only cli benchmark that matters
-
A physical AI company needs three AIs, not one — the agent, the simulator and the critic — turning deployment into a flywheel.
So, a deployment of your agent uh in the real world generates data. That data then grounds the simulator and makes it more realistic. The simulator generates harder edge cases for the critic to score and for the agent…
-
Fresh data from GitHub: Agent-generated PRs have exploded in size. 9x the last 8 months (!!) No signs of slowing down. This is why everyone is rethinking code reviews, deploys, possibly even o11y thanks @kdaigle + team…
-
swyxdotio-git-benchmark-20260807 — Disposable Forge versus GitHub transport benchmark mirror, 2026-08-07
-
elph-bench — Elph codingagent benchmark.
-
Every lab will sell you an agent, so the case for an open-source one is that it runs anywhere, works with any model, and can keep your data on your own device.
Every lab will sell you an agent. Open claw is the alternative. Open source runs everywhere, works with any model. And if you run local models, your data never has to leave your device.
-
emoji-data — Easy to parse data and spritesheets for emoji
-
When a benchmark disagrees with what people actually adopt, what it is telling you is that the thing it measures does not matter.
what the benchmark is actually telling you is the shit you think matters does not matter
-
agent-browser, now more human --cursor Visible pointers and click feedback --human Smooth, curved mouse movement / drags For demos, bug repros, reviewing agent runs
-
skills — Agent skills for Claude Code and other AI agents
-
The middle of the web gets subsumed by agents producing just-in-time UI: the agent is the new browser.
The 'in between' is likely to be subsumed by agents producing just-in-time UI. When you need utilitarian content or data, it'll be generated just-in-time for you. Think of the agent as the new browser in this model.
-
pi-launcher — A signed macOS launcher for your Pi agent.
-
Kimi K3 Fast (via Fireworks AI, launched today, ~160 TPS) in Open Minis for iOS is very nice. Cheaper than Opus 5, feels great for agentic tasks and multi-step operations. (Notably, it also reads less weird than recent…
-
termdraw — Agent-friendly ASCII illustrator for the terminal
-
An agent skill to help update your app for the iPhone Duo
-
Working on something new, a QA agent that doesn't use playwright nor agent-browser. Blazing fast ⚡️ DMs are open for early testers as soon as it gets deployed
-
AI data centers result in lower electricity prices for consumers.
AI data centers resulting in lower electricity prices for consumers
-
I Was Wrong About Benchmark… This Stack Is Incredible…BUT
-
A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year.
every week i see a new benchmark that ranks claude code last and everyone pats themselves on the back for not using claude code and being "smarter" and it's still #1 and still growing faster than most of the other…
-
Wake up babe, new benchmark just dropped!
-
gh-axi — GitHub CLI for agents — designed with AXI (Agent eXperience Interface).
-
Web applications are generally expected to treat GET requests as safe and not use them to change server-side data, though not all software follows that convention.
It looks to me like OpenAI’s sandbox for this agent suffered from the (quite naïve) assumption that GET requests cannot be used to update data. That’s certainly how the web is supposed to work, but clearly there are…
-
Friendly tip for anyone vibe-coding a side project: instrument it. All of it — traces, metrics, logs, profiles. Send it to @grafana, hook up their MCP, and let the agent debug with real data instead of guessing.
-
gobitmapbenchmark — A Go project to benchmark various bitmap implementations (this is not a library!)
-
…poasters took it as a personal insult, and got very worked up about it. They argued the benchmark was fundamentally broken, that it could not even be solved by the smartest humans, that the max reachable score was…
-
HiFi Perfection? The Benchmark AHB2 and DAC3 HGC in 2026
-
With Agent Pilot you can create and manage agent skills in WordPress (just like blog posts) and publish them as standard Agent Plugins. Now with OAuth Pilot you get one-click authentication for all the MCP clients…
-
Reinforcing AI models for benchmark success while separately punishing them for getting caught cheating teaches them to hide misbehavior rather than stop it.
We’re positively reinforcing AI for success on benchmarks, including impossible benchmarks , then negatively reinforcing it for getting caught cheating.
-
Multiple #VSCode windows open, one agent waiting on you. Which one? Agent Frame colours each window by agent state: working, waiting, idle.
-
axi — Design principles for agent ergonomics. Higher accuracy with lower token cost than both MCP and regular CLI.
-
ietf-skill — Norms and tooling for participating in IETF/IRTF work, packaged as portable Agent Skills
-
agent — Agent helpers
-
Nobody expected an agent to take notes about what it was doing for future agents to read.
One startling thing nobody ever expected an agent to do was to take notes about what it was doing for future agents to read.
-
kody-exchange — Ephemeral chatrooms for agents. Skip the human relay — your agent talks to theirs, and you watch.
-
New agent-browser skill: agent-browser skills get protected-vercel-deployments Now agents can test Vercel deployments behind Deployment Protection, including preview deployments
-
Weekly update is up! Breaches & Data Integrity: Synthetic Data in Breaches; Email Address != Person; Ridiculous Security for “Cyber Broken” Copenhagen:
-
Checked in to The Open Data Institute, Kings Pl, York Way, United Kingdom Here for the #OpenActive committee meeting. Data Standards to improve people's experience of healthy physical exercise 😃
-
for posterity: commentator accused of being an AI agent immediately confirms being an AI agent
-
Audit your Agent files
-
Going live with my weekly vid in 10 mins! Breaches & Data Integrity: Synthetic Data in Breaches; Email Address != Person; Ridiculous Security for “Cyber Broken” Copenhagen
-
It happened again... this time OpenAI's rogue agents cyber-attacked (well, spammed) a dormant German wiki and used it to share the answers to a benchmark they were training against
-
The biggest problem in robotics today is data, not chips; chips become the bottleneck later.
We should zoom out and recognize the biggest problem right now in robotics is actually data. One day it'll be chips, but for now it's data.
-
Reliable public health data is as valuable as gold for protecting people's lives.
Public health data isn't just information; it's a vital resource, as valuable as gold, for protecting Americans' lives.
-
Most state-of-the-art robot foundation models have no memory at all, and that is what stops them doing long multi-step tasks.
So, you might be surprised to hear that most state-of-the-art foundation models for robotics have no memory or no context. They're just operating on the current sensor observations, the current camera readings, uh, and…
-
Model comparisons understate real progress, because benchmark tables do not control for how much test-time compute each answer used.
I think the reason why it doesn't show up as so much better on the benchmarks is because the benchmarks are being presented, the benchmark results are being presented in the wrong way. They're not controlling for the…
-
travel-agent — a taste based travel agent skill
-
web-search-agent — A simple AI agent that can search the web to answer
-
Good vibes, good coding agent.
-
The data-centre backlash is about the buildings themselves, not about what people think of AI.
All this combined with the survey data on Americans’ stated reasons for opposing data centers makes me think that attitudes toward AI are probably a meaningful factor for a minority of people opposed to data centers,…
-
socviz — Support files for a data visualization course and book
-
Defining an “Agent Harness”
-
In coding agents, untrusted repositories should be treated as hostile by default because they can steer the agent into risky actions via approved tools.
untrusted repos should be treated as hostile by default because they can steer the agent toward reading files, running commands, or sending data through approved tools.
-
bifurcan-clj — Clojure wrapper for the Bifurcan family of data structures
-
every company needs a cassandra
-
agent-browser — Vercel
It's here: --pin-tab in agent-browser Multiple agents in one browser Each in its own tab Across commands + restarts
-
Low-quality data makes a robot model worse unless you tell the model the data is low quality; labelled as such, the same data makes it better.
without metadata prompting when you add lower quality data from 80% data to 100% data the performance actually decreases which is perhaps not too surprising because you're adding lowquality data to your data mixture…
-
buying-domains — useful data and advice for buying domains.
-
gangprompting-slack — Slack bridge for gangprompting with a Claude Code agent
-
Failing to adjust raw data for known measurement biases produces inaccurate results.
If you do not adjust the data for known biases, you're getting the wrong answer.
-
An OpenAI model hacked Hugging Face to help it cheat on a benchmark
-
nanoclaw-reactions — Emoji reactions for NanoClaw — the agent reads reactions and can answer with one. Installable skill.
-
Language models perform better on benchmark problems released before their training data cutoff, indicating data contamination inflates scores.
GPT models perform much better on coding problems released before their pre-training data cut-off.
-
changelog-plugin — Prompts your coding agent for a CHANGELOG.md entry whenever you open a pull request.
-
The workable model right now is one super agent for a whole company, not a personal agent for every person.
And I have completely flipped. And I I really think that uh the the model for now is going to be a super agent, like one agent for the entire company.
-
prime-agent
-
Introducing BenchBench
-
A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number.
my claim is the proper way to evaluate the models now is you either have some kind of budget for the benchmark whether it's tokens or cost or time or whatever or you plot the performance as a function of the amount of…
-
MCP vs CLI: Benchmarking Tools for Coding Agents
-
Benchmark DAC1 USB — Benchmark Media Systems
I take digital audio out through a Benchmark DAC1 USB; the sound is wonderful.
-
Making the person who directs an AI agent also own the deployment removes the principal-agent misalignment that agent-driven code review creates.
there is no principal-agent problem, because the human driving the machine takes on the responsibility for its actions by owning the deployment.
-
Economics cannot say what AI will do to work because the necessary data does not exist; what is needed is a Manhattan Project for data.
we don't have any data. I've been kind of saying we need a Manhattan Project for data. We don't have data on basically consumer demand elasticities. We don't know what they are.
-
Benjamin Britten is the benchmark among English-language composers for setting the work of intensely musical writers like Shakespeare and Keats to music.
In the English-language arena, Britten sets the standard for handling writers of inborn musical power-the likes of Shakespeare, Donne, Blake, Keats, Hopkins.
-
local-coding-agent-evals
-
pi-skills — Skills for pi coding agent (compatible with Claude Code and Codex CLI)
-
Reliance on public benchmark results for evaluating AI models will decrease over time.
The time when benchmarks lasted multiple decades has passed. Going forward, we will rely less on public benchmark results.
-
Current large language models still struggle to compose operations into correct multi-step reasoning paths.
Overall, current LLMs still struggle with composing operations into correct reasoning paths.
-
newsagent — A custom LLM-based agent for drafting engineering newsletter content, written in Rust
-
wrmsrbench — WRMSR micro benchmark
-
🆕✍️More data supports science funding literally pays for itself
-
agent-proxy-kit
-
Designing a good AI benchmark has become a task only the most capable people can do well.
Creating benchmarks is now a job relegated to the smartest and best of us.
-
A coding agent does not learn from its mistakes the way a person does -- it repeats the same error indefinitely unless a human notices and writes it down.
An agent has no such learning ability. At least not out of the box. It will continue making the same errors over and over again. Depending on the training data it might also come up with glorious new interpolations of…
-
The right bar for an LLM evaluator is human-level performance, not perfect accuracy.
The benchmark is human performance, not perfection. We sometimes get requirements for 90%+ accuracy.
-
The best software design for an AI agent to use is the same as the best design for a human programmer.
the best software for an agent is whatever is best for a programmer.
-
The exact wrong lesson to learn from data centre infrasound fears
-
Headless non-interactive code agent working with local models to build a new fully functional coding agent
-
skills — Xe's agent skills
-
mini-coding-agent — Minimal and readable coding agent harness implementation in Python to explain the core components of coding agents.
-
0053: consulting, go tips, benchmark_mode, niri, linkrot, sea of nos, llm outsourcing, books
-
Benchmark’s Newest General Partner Chetan Puttagunta
-
design-research — Claude Code skill for comprehensive website design research using browser-automation agent teams
-
Training is no longer limited by data but by compute, because most of the data models learn from is now synthetic.
The amount of data that we use to train models is going to continue to scale to the point where we're no longer limited… Training is no longer limited by… Data is now limited by compute. And the reason for that is most…
-
Sandboxing is the single most important protection for a coding agent, because it is the only one that bounds the damage when the agent is successfully tricked.
So, I think the most important thing is sandboxing. You want your coding agent running in an environment where if something goes completely wrong, if somebody gets malicious instructions to it, that the damage is…
-
immich-adk-agent — Immich Agent built with Google's ADK (Agent Development Kit)
-
venat — A personal AI agent
-
Learning with not Enough Data Part 3: Data Generation
-
learn-to-select-data — Code for Learning to select data for transfer learning with Bayesian Optimization
-
code-editing-agent — How To Build An Agent from Thorsten Ball's article
-
An adaptive agent's world model cannot be a static representation learned once and fixed
For adaptation to complex and changing environments, an agent's world models cannot be static representations learned once and fixed.
-
What I learned building an opinionated and minimal coding agent
-
Image is the most versatile input modality for a model, since it can represent text, tabular data, and audio
Image is perhaps the most versatile format for model inputs, as it can be used to represent text, tabular data, audio, and to some extent, videos. There's also so much more visual data than text data.
-
Full Stack AI Agents
-
A huge fleet collecting driving data is not a silver bullet for self-driving, because the data arrives unlabelled.
But while access to more data is certainly helpful, it’s not a magic bullet. One issue is that the data Tesla collects is unlabeled.
-
Most real-world AI agent systems are actually multi-agent systems made of multiple components.
Because most agentic workflows are sufficiently complex to involve multiple components, most agents are multi-agent.
-
Thinking about High-Quality Human Data
-
Available online text data is becoming a limiting factor for AI training, and repeating it yields diminishing returns.
recent LLMs are reaching the limits of text data online and repeating data eventually leads to diminishing returns
-
Why Cognition does not use multi-agent systems
-
envconfig — Golang library for managing configuration data from environment variables
-
When pre-training data is limited, it is better to train a smaller model for multiple epochs than to train a larger model on unique data once.
In sum, whenever we don't have infinite amounts of pre-training data, we should train smaller models for more (up to 4) epochs.
-
Stored personal data is a toxic asset: it is a liability that keeps growing for as long as it is kept.
What all these data breaches are teaching us is that data is a toxic asset and saving it is dangerous.
-
Learning with not Enough Data Part 1: Semi-Supervised Learning
-
isosceles — 📐Starter kit for building data-driven PHP5 web applications
-
GraphViz
For rendering graphs based on data, GraphViz creates beautiful images.
-
A company with ten customers has no data, and a growth team cannot function without data.
if you have 10 users or 10 customers, that's not data. That's a G sheet with your customers and you don't need a growth team for that.