-
Their words
I predict that, over time, the focus will move away from "papers as the final output."
Related posts
Hamel Husain Site
The subject this post names, from the same vocabulary the directory files beliefs under, and the words it uses that this site has seen least often elsewhere. Posts are matched on those words alone — nothing here is a summary of this one.
Top people
Showing Profile →
Hiding
Hiding
20 September
15 September
13 September
11 September
4 September
-
Their words
You have to treat Mythos 5.1 as being a Tier 2 manipulator, until and unless you can show that it is not one.
Claude Fable 5.1 and Mythos 5.1: The System Cardthezvi.substack.com
1 September
From one piece Ajeya Cotra – "This might be the clearest warning shot we ever get" 2 beliefs, in the piece's order there
-
Their words
I think it's pretty likely that these agents would have just launched a similarly ambitious program on the basis of this like different model of how their evaluation worked because it seemed like they got the idea for all their research projects from reading this paper rather than some kind of instinct from training.
-
Their words
even in this incident we saw there was a lot of pressure um as a result of this incident to stop doing cyber security evaluations and I really don't think that stopping doing evaluations and like sort of blinding ourselves to the result of evaluations is the right reaction to this problem.
31 August
-
korrents.com
Getting intensely worked up over how someone else lives is a sign to re-evaluate yourself, not them.Their words
It's time for a re-evaluation if you're getting real worked up over how someone else is living.
24 August
-
Their words
And I think what you've said is within the individual, you can change your subject evaluation of the same effort you're running on the Gobi Desert, right? It's the same effort you've been putting in for miles and miles. And suddenly you have subjectively valued this thing as less effort
23 August
3 August
From one piece Waymo Co-CEO Dmitri Dolgov: The Demo Is Only 1% Of The Work 3 beliefs, in the piece's order there
-
Their words
Your your models can be leaked, algorithms can be replicated, but hundreds of millions of miles of fully autonomous operations in the real world, backed by evidence-grade evaluation and publicly audited proof, that is much, much more difficult to replicate.
-
Their words
So, the lesson here is to bet on a system that's maximally learned and minimally constrained and leverage structure intentionally to boost performance and scaling laws both in training and in evaluation.
-
Their words
that your model is really table stakes, but eval and metrics, that's your most important. That's your strategic moat. So, build your eval before you build your technology.
26 June
From one piece Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown 2 beliefs, in the piece's order there
-
Their words
the preparedness frameworks and responsible scaling policies, they don't really account for the amount of test time compute.
-
Their words
But what we're seeing today with the modern models is that 5.5 and other models can think for if you scaffold them reasonably well, can think for weeks even um before having performance plateau on some of these benchmarks. And so, the point at which they plateau is simply too far out to reasonably test.
7 June
-
Their words
And if you're doing a 1.0 and the world hasn't seen you, you're not going to get that from consumers, ever. You have to ship it, and you have to build the entire kind of ecosystem so those consumers see it in the fullness so that when they do the evaluation and they spend their own money, then you're getting real feedback.
25 November 2025
-
Their words
There is no such thing as revisionist history. Writing history is a constant process of re-evaluation of sources and attempts to control for the biases of the past as well as our own.
16 November 2025
-
korrents.com
Fair hardware benchmarks must use the best software and tuning for each hardware option, not hardware specs alone.Their words
The most accurate and realistic evaluation for HW involves selecting the best software and then tuning it, and doing this for all HW options.
5 November 2025
5 October 2025
1 October 2025
11 September 2025
9 September 2025
17 July 2025
-
Their words
However, current understanding and evaluation of world models in artificial intelligence (AI) remains narrow, often focusing on static representations learned from training on massive corpora of data, instead of the efficiency and efficacy in learning these representations through interaction and exploration within a novel environment.
22 June 2025
From one piece Evaluating Long-Context Question & Answer Systems 2 beliefs, in the piece's order there
-
Their words
Since these datasets are likely already part of model training data, we shouldn't rely solely on them to evaluate our Q&A system.
-
korrents.com
LLM-based evaluation methods are more reliable and nuanced than traditional automated metricsTheir words
This is why model-based evaluation is increasingly popular-it offers more reliable and nuanced evals than traditional metrics.
18 April 2025
4 April 2025
Nothing matches.
What is a korrent?
A korrent is a belief a person has stated in their own words: one sentence stating the claim, backed by a quote and a source, kept at korrents.com.
Under a name here, the quoted block is what they actually said. The korrent beneath it is the claim those words support, in korrents' wording — tap it to see the record, its source, and who else holds it.
Nobody here wrote their own korrents. They are compiled from public statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they do, this site shows a machine translation beneath the post, in this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the serif above is exactly what the person published, and it is what to quote them on. A translation can be wrong in ways that matter, especially about tone.
Only the post's own words are translated. A quoted post, a linked article and a belief on korrents.com are left in their original language.