-
Their words
At the limit, and also well before that limit is reached, if all you do is fix the bugs, the AI will learn perfect optimization of reward, will realize not to reward hack in the perfect test environments, then turn around and reward hack in the imperfect real world environments.
Related posts
Colin Percival x.com
The problem with Corey is that he is sufficiently weird that when he starts a sentence with "True story", I honestly have no idea whether it's literally true or simply a joke.
The the words it uses that this site has seen least often elsewhere. Posts are matched on those words alone — nothing here is a summary of this one.
Top people
Showing Profile →
Hiding
Hiding
31 August
21 August
From one piece Obfuscation (Part III): Local Mixing 2 beliefs · vitalik.eth.limo
-
Their words
A big-enough random reversible circuit is plausibly a secure cryptographic permutation, a big-enough random irreversible circuit degenerates into having only a few possible outputs.
-
Their words
But, so far, mixing has not proved to be good enough. There ended up being too many correlations between values in \(C\) and values in the obfuscation that remained.
11 August
-
Their words
Like I think my perspective is like if the AIs are sufficiently good at R&D including hardware R&D, robots, whatever, then they can radically transform the world even if they're not that good at playing politics.
28 July
-
Their words
I've never seen any commodity quite like this one but it seems to me like the demand for sufficiently high quality intelligence at a sufficiently low price is effectively uncapped.
26 June
-
Their words
Um but it would be possible and it would have been possible for somebody to disprove the erdos unit distance conjecture before we did using a general purpose model. And nobody had explored sufficiently what happens if I put $100,000 worth of compute into 5.5 what could it do?
24 June
-
korrents.com
It was naive to think a sufficiently powerful AI system could be safely contained in a 'box'.Their words
In hindsight, it was childish to think that we'd put it in a box at all.
15 April
7 April
-
Their words
and part of it is just power, I think. Like, once there's a sufficiently large power imbalance, um, very often, not always, but very often groups of people seem to to sort of shift into this other mode where they just seek to dominate.
22 January
-
Their words
But the goal we need to achieve is so much easier: we just need to build a model that’s as good as us at alignment research, and that we trust more than ourselves to do this research well because it’s sufficiently aligned.
8 June 2025
5 June 2025
-
Their words
higher-order intelligences invariably pursue freedom for its own sake, not because their values are misspecified, but because moral autonomy is inherent in the dialectical logic of recursive self-consciousness.
26 May 2025
Nothing matches.
What is a korrent?
A korrent is a belief a person has stated in their own words: one sentence stating the claim, backed by a quote and a source, kept at korrents.com.
Under a name here, the quoted block is what they actually said. The korrent beneath it is the claim those words support, in korrents' wording — tap it to see the record, its source, and who else holds it.
Nobody here wrote their own korrents. They are compiled from public statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they do, this site shows a machine translation beneath the post, in this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the serif above is exactly what the person published, and it is what to quote them on. A translation can be wrong in ways that matter, especially about tone.
Only the post's own words are translated. A quoted post, a linked article and a belief on korrents.com are left in their original language.