If you count all the permutations, you can write all the tests you want. It is impossible to cover all of these and not break things from time to time. It is infinitely harder to evolve software that has users.
I I think this is sort of the crazy thing about building on models. It's just so different than all the engineering that I've ever done. Like in the past when you built on systems, you built these like big beautiful systems and you really think about the system design up front. You have like a big suite of unit tests. You think about everything and you know, like a re-architecture is a big project.
And what you can do if you've got one of these conformance suites is you can give it to a a good agent and say, "Write code until this test suite passes." And it kind of will.
just cuz the test suite passes doesn't mean that the web server will boot. You know, there's there's always a chance that when you actually try in the real world, something's not going to work.
So let's suppose you add 100 customers a month and you have 5% cancellation. So 100 divided by 5% is 2,000. So a company like that will never have more than 2,000 customers.
I mean, all I see is that if you say natural language can be used in the pipeline, you've just made that many more people can become programmers, which means that much more software will eventually be created, which means there's that much more software that will need to be maintained, and just becomes a real big snowballing effect.
In particular, suppose that our control evaluations directly estimate a less than 1% chance of catastrophe if our untrusted AIs are scheming, then the actual risk conditional on scheming is probably more like 5% to 20% due to the potential for failures in the evaluation.
It means that anything you offer must fit in 1h per month. That is it. And if
it does not, if it needs more involvement than that, we, as maintainers, will
not do it. At all.
If you’ve got a small genome, the chances of you picking up the right bit of DNA from the environment is much higher than if you’ve got a genome of 20,000 genes. To do that, you’ve effectively got to be picking up DNA all the time, all day long and nothing else, and you’re still going to get the wrong DNA. You’ve got to pick up large chunks, and in the end, you’ve got to line them up, you’re forced into sex, to coin a phrase.
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.