Reported a security issue to aws-security@amazon.com. Got back a canned email encouraging me to consider an AWS Support plan. That is not how this is supposed to work, people.
It looks to me like OpenAI’s sandbox for this agent suffered from the (quite naïve) assumption that GET requests cannot be used to update data. That’s certainly how the web is supposed to work, but clearly there are applications that don’t hold to that contract.
it's not that they lack ambition, it's more like they've been trained not to show it. Because when you're young you're often in situations where you're supposed to be obedient. You're not supposed to like go off and do what you want.
in the future, uh, we would be very concerned about investigator agents and like monitor agents colluding with the agents they're supposed to investigate or monitor.
I think the agents were right to presume causal grading. It turned out to be wrong, but it’s a mistake you are clearly supposed to make here, in response to a mistake by OpenAI where they failed to implement properly.
What we find is that those individuals who score highly on just self-reported questionnaires for motivation, if they score highly for apathy, they have twice the risk of developing Alzheimer's disease than people who don't.
But when they looked at the happiness levels of these people, it was the average person who actually would self-report the highest levels of happiness. So ambition doesn't necessarily equate with happiness, which I think many people would agree with with that.
And often I found with my clients I have to tell them like it's doing a good job at generating the actual design, but in actually expressing what the design is supposed to do, it cannot do that yet. You have to do that part yourself.
when you start talking about like most interesting domain problems, you have to pull in so much context that basically even writing what the function is supposed to do becomes a nightmare. The imperative program you write that will get correct 99% of the time is probably good enough to use in almost all cases.
In 1969, we landed on the moon for the first time and the same year we flew Concord, the faster than the speed of sound airliner. The future was supposed to be faster and better. where we're supposed to look forward to innovation in air and in space. And yet, half a century later, we can't go to the moon and we can't fly faster than the speed of sound.
So if you do a pure chronological feed, the incentive for everybody is to just post as much as possible uh because it will always be at the top of everyone who follows you's feed as soon as you post.
my claim is the proper way to evaluate the models now is you either have some kind of budget for the benchmark whether it's tokens or cost or time or whatever or you plot the performance as a function of the amount of test time compute that's going into the model
Being a great listener when people speak. Deep insights about understanding, connection, helping people express themselves, overcoming assumptions, the ethics of gossip, and more. Specific techniques for the support response, encouraging elaboration, and keeping it balanced. You can’t be ethical without being a good listener. When people say, “I can’t talk right now,” what they really mean is “…
it doesn't matter how good the drink is if people think you've bought it for $8.95 it's not doing the job it's supposed to do which is to signal generosity to signal hospitality or to mark a special occasion
You know, I’ve been in games for 29 years, and all the time, the piece of tech that’s going to make making games much easier and much cheaper is about to turn up, and all that’s happened is the games have got much better and way more expensive.
When you lose a game now as opposed to surfacing your blunders and your like horrible stuff that you did, we flip it on its head. And so we show you your brilliant moves, your best moves. And we have coach say something encouraging. You know, losing just part of learning like keep it up, that type of thing. That change alone was pretty dramatic for us. It grew game reviews by 25%, subscriptions by 20%
And you're supposed to be really connected, but the conclusion you reach very early is that the more connected and accessible you are, the less productive you are.
So there's no ground truth. You can't have prior knowledge if you don't have ground truth because the prior knowledge is supposed to be a hint or an initial belief about what the truth is.
It's just that Zig, C, C++, all those languages that were being tested, they're all LLVM backends, right? That's the one that actually turns the thing into the executable part. And if there's a variation in speed, it just means in one language you didn't quite express what you are supposed to correctly.
It's easy to get impressive-looking results if you're comparing against a poorly-tuned baseline, and that observation turns out to explain a surprising fraction of supposed improvements.
But the underlying ideology is familiar from Silicon Valley, Mark Zuckerberg's dictum to "move fast and break things" and replace the messily human with frictionless technology, built by elite engineers who are supposed to know better than anyone else.
One is ByteDance, arguably is the largest smuggler of GPUs for China. China's not supposed to have GPUs. ByteDance has over 500,000 GPUs. Why? Because they're all rented from companies around the world.
And as responsible scientists, we’re trying to disprove our theories. We are not supposed to be trying to prove our theories. That’s one more foot out of the science box that archeology often steps.
The principle in Perplexity is you’re not supposed to say anything that you don’t retrieve, which is even more powerful than RAG because RAG just says, “Okay, use this additional context and write an answer.” But we say, “Don’t use anything more than that too.” That way we ensure a factual grounding.
Indeed, as I've reported in previous posts, research not only shows therapists do not improve with time and experience in the field, but on average become less effective (1, 2, 3, 4).
It’s one of the many factors I think that went into the fact that the Large Hadron Collider became the only machine in town, and the Superconducting Supercollider would’ve just been a much… If it had really had achieved what it was supposed to, would’ve been a much more robust test of the space.
The short answer is that elevating the fulfilment of humanity’s supposed potential above all else could nontrivially increase the probability that actual people – those alive today and in the near future – suffer extreme harms, even death.
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.