It looks to me like OpenAI’s sandbox for this agent suffered from the (quite naïve) assumption that GET requests cannot be used to update data. That’s certainly how the web is supposed to work, but clearly there are applications that don’t hold to that contract.
I think it's pretty likely that these agents would have just launched a similarly ambitious program on the basis of this like different model of how their evaluation worked because it seemed like they got the idea for all their research projects from reading this paper rather than some kind of instinct from training.
I’m thinking about banning Claude code at Shopify until they change their mind and read AGENTS.md and .agents/skills etc.
Insisting on only reading CLAUDE.md sometimes leads to split brain problems when different team members use different tools. Just unnecessary.
what what is the current thing that models are not very good at doing? Alpha is going to decay in some way with every model release or every series of model releases. So, it's going to change over time.
So, anyway, I'd say it like it both seemed like an obviously good idea in that people really wanted this, but also a bad idea in that nobody took it seriously. Um but I I think that I think the fact that it was ultimately the fact that it was grounded in such a concrete actual real user problem saved us.
Um, another way you can get more experience for yourself is to just write down a bunch of things you think might be important in the next 12 months. And maybe you pick one of them to work on, but go back and evaluate in 12 months of these other things, which ones actually seemed important or which ones did other people in the world go out and and create and which ones did they did not seem to do yet.
what I don't like about it I didn't like about it then and still don't like about it is it's not defensible. Nobody's going to say I'm not agile. Oh no, I prefer rigid development. Oh, I prefer inflexible development. No, everybody's going to say that they're agile, which extreme doesn't have that problem.
I only checked a handful of youtube reviewers, but my favorite is Jennifer Wang, who seemed properly autistic about clothing quality and understanding that there are multiple places on the pareto frontier one might choose to occupy
ChatGPT, in particular, seemed to just want to validate me, tell me how great I was, reinforce any bad beliefs I might have had, and avoid saying anything uncomfortable.
And it's just this, this is an example of how language affects thought. Scaling is what just one word, but it's such a powerful word because it informs people what to do.
I nonetheless found this book somewhat disappointing, because Jay seemed content to expound the authors he covers in more or less their own terms, rather than subjecting them to anything like actual epistemological or methodological criticism.
the way to be on the hook is to say, "This is only for people who are like this." Because if those people hate it, then you were wrong. Whereas if you say, "This is for everyone," you're allowed to hide behind, "Well, everyone hasn't found it yet."
If I were forced at gunpoint to guess, I'd say that human-level AI seemed to me like a slog of many more centuries or millennia
Their words now
I was wrong, of course, not to contemplate more seriously the prospect that AI might enter a civilization-altering trajectory, not merely eventually but within the next decade.
If I were forced at gunpoint to guess, I’d say that human-level AI seemed to me like a slog of many more centuries or millennia (with the obvious potential for black swans along the way).
Their words now
I was wrong, of course, not to contemplate more seriously the prospect that AI might enter a civilization-altering trajectory, not merely eventually but within the next decade.
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.