Related posts
Alex Tabarrok x.com
Monitoring COT and communication between agents is also getting more difficult as networks get bigger and models become more advanced, according to @ConnorTabarrok, policy lead at Equistamp, a startup that provides safety evaluations for AI projects.
The subject this post names, from the same vocabulary the directory files beliefs under, and the words it uses that this site has seen least often elsewhere. Posts are matched on those words alone — nothing here is a summary of this one.
Top people
Showing Profile →
Hiding
Hiding
19 September
6 September
-
korrents.com
Chain-of-thought monitoring is getting less reliable, not more, as models grow more capable.Their words
However, unfortunately our evaluations indicate our ability to rely on CoT monitoring is progressively diminishing.
1 September
From one piece Ajeya Cotra – "This might be the clearest warning shot we ever get" 2 beliefs, in the piece's order there
-
Their words
even in this incident we saw there was a lot of pressure um as a result of this incident to stop doing cyber security evaluations and I really don't think that stopping doing evaluations and like sort of blinding ourselves to the result of evaluations is the right reaction to this problem.
-
Their words
But from their perspective, they've just been trained for millions of subjective years to do as well as they possibly can on these evals. In many cases, the only way in which they've been able to perform well on that training is explicitly by cheating, right?
11 August
-
Their words
when AIs are extremely extremely capable my view is that those AIs will be harder to align than current systems. So for current systems, we have this feedback loop where we basically like we create an AI. We do some evaluations on it. We see that it has some kind of messed up behavior that we can kind of quickly understand. Then we like can like go look in training and be like, "Oh, the these training environments led to this problematic behavior. Let's like tweak that training data. Let's introduce some additional training data to like correct this other issue and then move forward from there." But in a regime where the AIs are extremely situationally aware, very very very very capable and um you know uh we don't necessarily understand what they're doing, this feedback loop breaks down.
26 June
-
korrents.com
Evaluating a model properly would mean delaying its release, and competitive pressure means no lab will.Their words
It's actually very difficult because, yeah, you would have to the only way to to really do the evaluations is then delay the model release cycle. Um and you know there's a lot of competitive pressure right now to not do that.
3 June
1 October 2025
-
Mixed onrcmnd.app
BraintrustTheir words
The panel had a generally positive view of Braintrust, highlighting its clean UI and structured approach to evaluations. The tool’s emphasis on human-in-the-loop workflows was a significant strength.
11 September 2025
7 May 2024
From one piece The case for ensuring that powerful AIs are controlled 2 beliefs, in the piece's order there
-
Their words
Because evaluating control just requires evaluating capabilities, it's far easier to robustly evaluate than alignment.
-
Their words
In particular, suppose that our control evaluations directly estimate a less than 1% chance of catastrophe if our untrusted AIs are scheming, then the actual risk conditional on scheming is probably more like 5% to 20% due to the potential for failures in the evaluation.
26 October 2023
-
Their words
Despite evaluations, we cannot consider coming powerful frontier AI systems "safe unless proven unsafe". With current testing methodologies, issues can easily be missed. Additionally, it is unclear if governments can quickly build the immense expertise needed for reliable technical evaluations of AI capabilities and societal-scale risks. Given this, developers of frontier AI should carry the burden of proof to demonstrate that their plans keep risks within acceptable limits.
10 March 2010
Nothing matches.
What is a korrent?
A korrent is a belief a person has stated in their own words: one sentence stating the claim, backed by a quote and a source, kept at korrents.com.
Under a name here, the quoted block is what they actually said. The korrent beneath it is the claim those words support, in korrents' wording — tap it to see the record, its source, and who else holds it.
Nobody here wrote their own korrents. They are compiled from public statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they do, this site shows a machine translation beneath the post, in this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the serif above is exactly what the person published, and it is what to quote them on. A translation can be wrong in ways that matter, especially about tone.
Only the post's own words are translated. A quoted post, a linked article and a belief on korrents.com are left in their original language.