It looks to me like OpenAI’s sandbox for this agent suffered from the (quite naïve) assumption that GET requests cannot be used to update data. That’s certainly how the web is supposed to work, but clearly there are applications that don’t hold to that contract.
it's not that they lack ambition, it's more like they've been trained not to show it. Because when you're young you're often in situations where you're supposed to be obedient. You're not supposed to like go off and do what you want.
in the future, uh, we would be very concerned about investigator agents and like monitor agents colluding with the agents they're supposed to investigate or monitor.
I think the agents were right to presume causal grading. It turned out to be wrong, but it’s a mistake you are clearly supposed to make here, in response to a mistake by OpenAI where they failed to implement properly.
So, a deployment of your agent uh in the real world generates data. That data then grounds the simulator and makes it more realistic. The simulator generates harder edge cases for the critic to score and for the agent to learn from.
And often I found with my clients I have to tell them like it's doing a good job at generating the actual design, but in actually expressing what the design is supposed to do, it cannot do that yet. You have to do that part yourself.
when you start talking about like most interesting domain problems, you have to pull in so much context that basically even writing what the function is supposed to do becomes a nightmare. The imperative program you write that will get correct 99% of the time is probably good enough to use in almost all cases.
In 1969, we landed on the moon for the first time and the same year we flew Concord, the faster than the speed of sound airliner. The future was supposed to be faster and better. where we're supposed to look forward to innovation in air and in space. And yet, half a century later, we can't go to the moon and we can't fly faster than the speed of sound.
So if you do a pure chronological feed, the incentive for everybody is to just post as much as possible uh because it will always be at the top of everyone who follows you's feed as soon as you post.
Taste isn't consensus. It's not judgment first, either. It's easy to be a critic. Curiosity has to come first, because without curiosity you never broaden your horizons.
I also read critic Michael Dirda's On Conan Doyle, a lovely short book I'd never heard of that I spotted on the shelf at the library — hooray for the serendipity of the stacks!
it doesn't matter how good the drink is if people think you've bought it for $8.95 it's not doing the job it's supposed to do which is to signal generosity to signal hospitality or to mark a special occasion
You know, I’ve been in games for 29 years, and all the time, the piece of tech that’s going to make making games much easier and much cheaper is about to turn up, and all that’s happened is the games have got much better and way more expensive.
And you're supposed to be really connected, but the conclusion you reach very early is that the more connected and accessible you are, the less productive you are.
So there's no ground truth. You can't have prior knowledge if you don't have ground truth because the prior knowledge is supposed to be a hint or an initial belief about what the truth is.
It's just that Zig, C, C++, all those languages that were being tested, they're all LLVM backends, right? That's the one that actually turns the thing into the executable part. And if there's a variation in speed, it just means in one language you didn't quite express what you are supposed to correctly.
It's easy to get impressive-looking results if you're comparing against a poorly-tuned baseline, and that observation turns out to explain a surprising fraction of supposed improvements.
But the underlying ideology is familiar from Silicon Valley, Mark Zuckerberg's dictum to "move fast and break things" and replace the messily human with frictionless technology, built by elite engineers who are supposed to know better than anyone else.
One is ByteDance, arguably is the largest smuggler of GPUs for China. China's not supposed to have GPUs. ByteDance has over 500,000 GPUs. Why? Because they're all rented from companies around the world.
And as responsible scientists, we’re trying to disprove our theories. We are not supposed to be trying to prove our theories. That’s one more foot out of the science box that archeology often steps.
The principle in Perplexity is you’re not supposed to say anything that you don’t retrieve, which is even more powerful than RAG because RAG just says, “Okay, use this additional context and write an answer.” But we say, “Don’t use anything more than that too.” That way we ensure a factual grounding.
It’s one of the many factors I think that went into the fact that the Large Hadron Collider became the only machine in town, and the Superconducting Supercollider would’ve just been a much… If it had really had achieved what it was supposed to, would’ve been a much more robust test of the space.
The short answer is that elevating the fulfilment of humanity’s supposed potential above all else could nontrivially increase the probability that actual people – those alive today and in the near future – suffer extreme harms, even death.
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.