@linear – the product development system trusted by 40,000+ teams to plan and ship. Get six months of Linear Business for free: https://t.co/knIF2mcz9p
Every agent harness can now use Jev npm install -g ai-cli → ask yes/no questions → choose between options → score against your criteria https://t.co/dpgmcdCj8F
In the most early adopting tech pioneering place that's Silicon Valley new startups aren't even building software anymore And it's debatable if anyone will actually need a custom harness or it won't just be generically offered by the AI frontier companies So software is mostly dead and hardware it is
Codex lets you use any model you want (not just OpenAI) and harness is open source - Claude Code doesn’t and is closed source Given this is the two leading AI labs, notable difference in approaches
what Meta has assembled is free (to consumers) hardware and software that dramatically reduces the barrier to entry for ordinary people looking to harness the power of agents, making the AI upside a lot more accessible to the masses who don't want to buy a Mac Mini.
To me, the possibility of partnership is the feeling that despite the fundamental chasm between two human beings, love allows us to bridge the gap; alienation is the feeling that the gap cannot be bridged, and moreover, that I do not want to.
GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game.
It is an infuriatingly locked-down computer. Now, to Apple's credit, it's a pretty good computer for being locked down, but I don't want a locked-down computer. I wanna own my computer. Better yet, I wanna mutate my computer, and this is where the agentic age needs a new operating system. When you can vibe code whatever app comes to your mind, you should be able to vibe code your operating system.
When the public health systems meant to catch and stop this stuff are cut (less funding, less staff, less historical knowledge, and less leadership), outbreaks can, and will, emerge between the cracks.
If you keep asking this question, what you realize is energy is the fundamental input. When we figure out AI and robotics that allows us to do semi-aututonomous manufacturing, energy will become the cost of all things, right? The cost of buying a thing will become the cost of energy used to make it.
And um if you build a specialized chip for low precision dense linear algebra and can't do anything else that turns out to be really useful for machine learning inference uh even though it can't run Chrome or Word or whatever.
I think evals, they outlive the harness a little bit, but not by that much. Like an eval might live for maybe one, two, three model generations, but nowadays the you know, we're on the exponential. The model is improving so quickly, very often we just saturate the eval, and then we have to throw it away, and we have to come up with a new eval.
the the breakthrough for us was realizing that AlexNet was not AlexNet. That AlexNet was an approach with deep deep learning that allows you to learn any function.
I genuinely believe that if you took an open weights model from 2025 and built a pentest harness for it, it could do this kind of sandbox escape and scan/hack in most networks. This is only surprising because you assume OpenAI has sounder sandboxes.
I think context engineering has been so long lived because it's it's grounded in the fundamentals of how transformer attention works and until we have post transformer models or linear attention or whatever it is which who knows when that's going to happen context engineering will be interesting and important to anyone building on AI
But it's basically this idea that like the only thing that made claude code good was reinforcement learning. And the dimension along which it got good was like we made a model. We trained the model and the harness together. And so the model got really good at calling the specific tools in that harness.
there's some interesting work on you know happiness being more curve linear like midlife is actually some of the lowest levels of happiness and life satisfaction
I believe the same is true of the current intelligence explosion being driven by AI. We are in desperate need of new social structures that allow us to harness its power, without damaging or destroying human values in the process.
Using these BenQ light bars allows me to better illuminate my desk at night without having to keep my overhead lights on. What I really like about this light bar is that it can be configured to automatically turn off when I leave the desk.
And our harness wasn't very good for like the first like five months of open code. But it was good enough. It was good enough that most people couldn't really tell a difference. And once we won enough share, then we went back and like tried to make our harness like good and smart and optimize and all those things. But uh it was inverted from what everybody else was doing. Everybody else was being like you have to build the smartest harness and that's how you win.
The demand for inference is growing. So, like I don't think it's linearly growing. I think it might even be exponentially growing. But we haven't made our production of GPUs grow exponentially. That's like kind of a linear process. So as those lines intersect, there's going to be uh tightening.
You can have the compiler write the state machine. If you introduce syntax that allows you to indicate where you want to yield. And that's what await is.
Marketing grows only as fast as you can improve marketing. We all know that's quite hard actually. It's linear. It's hard to find new channels that aren't trivial. Like, it's hard. Of course, we're going to do it, but like it it's hard. Whereas, cancellations grow automatically as you grow, right? So, cancellations always overtake marketing for this reason.
The SRS is a brief, 4-item measure designed to assess the therapeutic alliance. It provides immediate feedback on the quality of the relationship and allows for course correction in real time.
can lead to imposter phenomenon. In this book, a leading psychologist looks at the science behind our inner voice and new research into how to harness it and improve your physical and mental health.
PMs are no longer saying to the designer, hey, can you draw this thing out for me? That frees up designer time to go explore more deeply the stuff they need to go into and it allows anyone to kind of add to that first conversation of where should we go and look further and wider and broader at the option space.
and that's in some ways what allows it to be very sturdy so I Think today management is really about this idea of like be sturdy while being flexible and that is a very hard thing to pull off
So here's what I would say that deep down at a very fundamental level the synthetic experience that you create yourself doesn't allow you to learn more about the world. It allows you to rehearse things. It allows you to consider counterfactuals but somehow information about the world needs to get injected into the system.
And when you make a mistake and correct it, well, first you you've achieved the task because you've corrected, but you've also gained knowledge that allows you to avoid that mistake in the future. With driving, because of the dynamics of how it's set up, it's very hard to make a mistake, correct it, and then learn from it because the mistakes themselves have significant ramifications.
I think the platform that Lean and other software tools, so GitHub and things like that will allow experimental mathematics to scale up to a much greater degree than we can do now.
This allows me to run models locally on my MacBook Pro M1 Max. With the 64GB of RAM it has, it’s a pretty potent machine for basic inference despite it being three years old.
It’s the fact that you’re a deterministic structure living in a random background. And also, all of that selection bundled in you allows you to select on possible futures. So that’s where your will comes from.
Breville Bambino Plus Espresso Machine and Baratza Encore ESP Pro Coffee Grinder - Developers need coffee, and I like mine HOT. The integrated milk steamer is a must-have, and since this machine takes freshly ground beans, the grinder allows me to tweak the grind size for perfect single dose espresso shots.
This is what I wish I learned from in my freshman year of college. The part of it about linear algebra is one of the best linear algebra resources out there, and remainder shows you how those tools apply to non-linear mathematics.
All of the above are installed using Homebrew and Homebrew Cask. Take a look at my Brewfile in my dotfiles for more info. This allows me to move to a new machine and be up and running incredibly quickly.
Fanny (macOS Specific) is a free Notification Center Widget and Menu Bar application to monitor your Macs fans. Allows me to know the speed of my fans and heat generated by the CPU, so that I can make sure I keep my apps to a minimum and get the maximum performance.
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.