-
Their words
I predict that, over time, the focus will move away from "papers as the final output."
Related posts
Gary Marcus Newsletter
BREAKING: Secret US AI evaluation framework has been partly revealed
The the words it uses that this site has seen least often elsewhere. Posts are matched on those words alone — nothing here is a summary of this one.
Top people
Showing Profile →
Hiding
Hiding
20 September
19 September
18 September
-
korrents.com
Freedom of navigation is a global public good that decays once its guarantor is seen as unwilling to bear its cost.Their words
Freedom of navigation is a global public good, and public goods decay the instant their provider is revealed to be unwilling to pay for them.
16 September
15 September
From one piece @levelsio on X 2 beliefs · x.com
-
Their words
I'm increasingly convinced gray hair is at least partly due to oxidative stress from the average diet now that's mostly processed food, as well as lack of heavy exercise and obesity
-
korrents.com
Gray hair can be at least partly avoided by eating clean, exercising heavily, and sleeping well.Their words
you can at least partly avoid it by eating clean, exercising heavy and sleeping well I think
-
Their words
it was a chance to stand up to the oligarchs and this technology they've staked their future on.
The Two Horsemen, Riding in Formationbillmckibben.substack.com
13 September
12 September
-
Their words
I suspect that some of the labs' willingness to put money into the social sciences stems from the realization of senior people that the practical build out of AI is highly unpopular in the U.S and their urgent desire to figure out how to sugarcoat the pill so that it will get swallowed
How should social science think about AI?programmablemutter.com
11 September
7 September
-
Their words
Once the entire plot is revealed in shows like this one, it can be a letdown, flattening what has happened previously.
My Frustration With Puzzlebox Televisiondanieldrezner.substack.com
4 September
-
Their words
You have to treat Mythos 5.1 as being a Tier 2 manipulator, until and unless you can show that it is not one.
Claude Fable 5.1 and Mythos 5.1: The System Cardthezvi.substack.com
-
Their words
Worse, U.S. officials now openly support and encourage politicians from far-right anti-European political parties, partly because they are ideologically aligned, but partly because Americans now want to weaken the European Union so that U.S. tech companies can evade any regulation, or even taxation.
The Myth of the “Censorship Industrial Complex”anneapplebaum.substack.com
1 September
-
korrents.com
AI companies are disrupting entire industries, causing a partial collapse of the traditional SaaS business model.Their words
I think you have the real economic effects of AI companies sucking up entire industries now (a lot where indie hackers operate) I do think the SaaSpocalypse is at least partly real
From one piece Ajeya Cotra – "This might be the clearest warning shot we ever get" 2 beliefs, in the piece's order there
-
Their words
I think it's pretty likely that these agents would have just launched a similarly ambitious program on the basis of this like different model of how their evaluation worked because it seemed like they got the idea for all their research projects from reading this paper rather than some kind of instinct from training.
-
Their words
even in this incident we saw there was a lot of pressure um as a result of this incident to stop doing cyber security evaluations and I really don't think that stopping doing evaluations and like sort of blinding ourselves to the result of evaluations is the right reaction to this problem.
31 August
-
korrents.com
Getting intensely worked up over how someone else lives is a sign to re-evaluate yourself, not them.Their words
It's time for a re-evaluation if you're getting real worked up over how someone else is living.
24 August
-
Their words
And I think what you've said is within the individual, you can change your subject evaluation of the same effort you're running on the Gobi Desert, right? It's the same effort you've been putting in for miles and miles. And suddenly you have subjectively valued this thing as less effort
23 August
9 August
6 August
-
korrents.com
The only style worth having is the one you cannot help; setting out to develop a personal style is the wrong goal.Their words
And that there might come a moment, while reading a new, just-edited-for-the-thousandth-time draft, where that true self, or voice, pops out at you, suddenly revealed, as if someone else, not you, had done it.
On the Perils of Self-Assessing...georgesaunders.substack.com
3 August
From one piece Waymo Co-CEO Dmitri Dolgov: The Demo Is Only 1% Of The Work 3 beliefs, in the piece's order there
-
Their words
Your your models can be leaked, algorithms can be replicated, but hundreds of millions of miles of fully autonomous operations in the real world, backed by evidence-grade evaluation and publicly audited proof, that is much, much more difficult to replicate.
-
Their words
So, the lesson here is to bet on a system that's maximally learned and minimally constrained and leverage structure intentionally to boost performance and scaling laws both in training and in evaluation.
-
Their words
that your model is really table stakes, but eval and metrics, that's your most important. That's your strategic moat. So, build your eval before you build your technology.
23 July
8 July
-
korrents.com
Recorded crime partly measures policing: a neighbourhood policed harder produces more crime in the statistics.Their words
Particularly with gun and drug crimes, communities with more aggressive policing tend to produce more crimes in part because law enforcement polices them more aggressively.
A guide to racism in the criminal justice systemradleybalko.substack.com
1 July
-
Their words
And then there were people who used it as a moral cudgel. Like you should be if you're not using TDD, you're not professional. And that's just such People can write very good software with a wide variety of workflows.
30 June
-
Lovedrcmnd.app
Forty Signs of RainTheir words
I loved this, but partly this was the pleasant shock of recognition: much of the action revolves around the National Science Foundation and, physically, its old headquarters in Arlington, which Robinson captures extremely well (even the atmosphere of review panels).
26 June
From one piece Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown 2 beliefs, in the piece's order there
-
Their words
the preparedness frameworks and responsible scaling policies, they don't really account for the amount of test time compute.
-
Their words
But what we're seeing today with the modern models is that 5.5 and other models can think for if you scaffold them reasonably well, can think for weeks even um before having performance plateau on some of these benchmarks. And so, the point at which they plateau is simply too far out to reasonably test.
7 June
-
Their words
And if you're doing a 1.0 and the world hasn't seen you, you're not going to get that from consumers, ever. You have to ship it, and you have to build the entire kind of ecosystem so those consumers see it in the fullness so that when they do the evaluation and they spend their own money, then you're getting real feedback.
29 May
-
Likedrcmnd.app
MacBook NeoTheir words
I had already put both laptops through my benchmark gauntlet, which revealed one theme: the Mac is faster (in most cases), more efficient, quieter, built better, has a much nicer display, and costs much less.
7 April
-
korrents.com
Once an explanation is revealed, people tend to believe they already knew it, even when they did not.Their words
Once you see the answer it's easy to believe that you already knew
25 March
4 February
12 December 2025
-
Their words
And it was in no one’s interest whatsoever. Nobody would ever concede any interest in the idea of literacy for all.
25 November 2025
-
Their words
There is no such thing as revisionist history. Writing history is a constant process of re-evaluation of sources and attempts to control for the biases of the past as well as our own.
16 November 2025
-
korrents.com
Fair hardware benchmarks must use the best software and tuning for each hardware option, not hardware specs alone.Their words
The most accurate and realistic evaluation for HW involves selecting the best software and then tuning it, and doing this for all HW options.
5 November 2025
5 October 2025
1 October 2025
17 September 2025
16 September 2025
-
Their words
Reading is attentional, not just informational.
11 September 2025
9 September 2025
17 July 2025
-
Their words
However, current understanding and evaluation of world models in artificial intelligence (AI) remains narrow, often focusing on static representations learned from training on massive corpora of data, instead of the efficiency and efficacy in learning these representations through interaction and exploration within a novel environment.
22 June 2025
From one piece Evaluating Long-Context Question & Answer Systems 2 beliefs, in the piece's order there
-
Their words
Since these datasets are likely already part of model training data, we shouldn't rely solely on them to evaluate our Q&A system.
-
korrents.com
LLM-based evaluation methods are more reliable and nuanced than traditional automated metricsTheir words
This is why model-based evaluation is increasingly popular-it offers more reliable and nuanced evals than traditional metrics.
18 April 2025
4 April 2025
13 March 2025
17 December 2024
-
Lovedrcmnd.app
Every Hand RevealedTheir words
A hand-by-hand narrative of Professional poker player Gus Hansen winning the 2007 Aussie Millions tournament.
29 October 2024
16 October 2024
-
Their words
I think it’s partly territorial. I cannot speak of all archeologists, but some archeologists feel very territorial about their profession. They do not feel happy about outsiders entering their realm, especially if those outsiders have a large platform.
19 June 2024
-
Their words
So most people, when they advertise new context window increase, they talk a lot about finding the needle in the haystack sort of evaluation metrics and less about whether there’s any degradation in the instruction following performance. So I think that’s where you need to make sure that throwing more information at a model doesn’t actually make it more confused.
18 May 2024
7 May 2024
-
Their words
In particular, suppose that our control evaluations directly estimate a less than 1% chance of catastrophe if our untrusted AIs are scheming, then the actual risk conditional on scheming is probably more like 5% to 20% due to the potential for failures in the evaluation.
6 November 2023
28 May 2023
-
Their words
The answer is not more fields, which means destroying even more wild ecosystems. It is partly better, more compact, cruelty-free and pollution-free factories.
16 May 2023
From one piece I wanted to be a teacher but they made me a cop 3 beliefs, in the piece's order there
-
korrents.com
Evaluation should be taken out of the classroom and given to people whose whole job it is.Their words
One solution is to separate instruction and evaluation. Let teachers be teachers and cops be cops.
I wanted to be a teacher but they made me a copexperimental-history.com
-
Their words
Instruction is collaborative: students want to learn, and I want them to learn, too. Evaluation is adversarial: students want the points, and I have to make sure they don’t get too many.
I wanted to be a teacher but they made me a copexperimental-history.com
-
korrents.com
Grading flattens whatever is taught into what can be tested, and teaches students to ignore everything else.Their words
Evaluation forces me to flatten everything I teach into something that can be tested, and it encourages students to ignore everything that isn’t on the test.
I wanted to be a teacher but they made me a copexperimental-history.com
1 May 2023
-
Their words
Good tools let the user choose when to switch between implementation and evaluation. When I work with a chatbot, I'm forced to frequently switch between the two modes.
1 January 2022
-
Their words
The US focus on efficiency has revealed the brittleness of its economy, which has neither the manufacturing capability to scale up domestic production of goods nor the logistics capacity to handle greater imports.
1 August 2021
-
Their words
Information revealed about “who sent what” in an E2E system is not the same as metadata.
Thinking about “traceability”blog.cryptographyengineering.com
1 January 2021
Nothing matches.
What is a korrent?
A korrent is a belief a person has stated in their own words: one sentence stating the claim, backed by a quote and a source, kept at korrents.com.
Under a name here, the quoted block is what they actually said. The korrent beneath it is the claim those words support, in korrents' wording — tap it to see the record, its source, and who else holds it.
Nobody here wrote their own korrents. They are compiled from public statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they do, this site shows a machine translation beneath the post, in this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the serif above is exactly what the person published, and it is what to quote them on. A translation can be wrong in ways that matter, especially about tone.
Only the post's own words are translated. A quoted post, a linked article and a belief on korrents.com are left in their original language.