Every agent harness can now use Jev npm install -g ai-cli → ask yes/no questions → choose between options → score against your criteria https://t.co/dpgmcdCj8F
In the most early adopting tech pioneering place that's Silicon Valley new startups aren't even building software anymore And it's debatable if anyone will actually need a custom harness or it won't just be generically offered by the AI frontier companies So software is mostly dead and hardware it is
Codex lets you use any model you want (not just OpenAI) and harness is open source - Claude Code doesn’t and is closed source Given this is the two leading AI labs, notable difference in approaches
what Meta has assembled is free (to consumers) hardware and software that dramatically reduces the barrier to entry for ordinary people looking to harness the power of agents, making the AI upside a lot more accessible to the masses who don't want to buy a Mac Mini.
GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game.
I think evals, they outlive the harness a little bit, but not by that much. Like an eval might live for maybe one, two, three model generations, but nowadays the you know, we're on the exponential. The model is improving so quickly, very often we just saturate the eval, and then we have to throw it away, and we have to come up with a new eval.
I genuinely believe that if you took an open weights model from 2025 and built a pentest harness for it, it could do this kind of sandbox escape and scan/hack in most networks. This is only surprising because you assume OpenAI has sounder sandboxes.
But it's basically this idea that like the only thing that made claude code good was reinforcement learning. And the dimension along which it got good was like we made a model. We trained the model and the harness together. And so the model got really good at calling the specific tools in that harness.
I believe the same is true of the current intelligence explosion being driven by AI. We are in desperate need of new social structures that allow us to harness its power, without damaging or destroying human values in the process.
And our harness wasn't very good for like the first like five months of open code. But it was good enough. It was good enough that most people couldn't really tell a difference. And once we won enough share, then we went back and like tried to make our harness like good and smart and optimize and all those things. But uh it was inverted from what everybody else was doing. Everybody else was being like you have to build the smartest harness and that's how you win.
can lead to imposter phenomenon. In this book, a leading psychologist looks at the science behind our inner voice and new research into how to harness it and improve your physical and mental health.
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.