Related posts
Kun Chen Newsletter
Evaluating the Effectiveness of Programming Languages for Agents
effectivenessevaluatinglanguages
The the words it uses that this site has seen least often elsewhere. Posts are matched on those words alone — nothing here is a summary of this one.
Top people
Showing Profile →
Hiding
Hiding
8 September
9 August
-
Their words
we say we want disagreement, but every disagreement costs the person making it a little social/political capital. do it often enough and you become the negative person. people stop evaluating the argument and start evaluating the source.
26 June
-
korrents.com
Evaluating a model properly would mean delaying its release, and competitive pressure means no lab will.Their words
It's actually very difficult because, yeah, you would have to the only way to to really do the evaluations is then delay the model release cycle. Um and you know there's a lot of competitive pressure right now to not do that.
24 June
-
Their words
And DSA interviews were never the best for that. Well, thinking, sure. But in terms of like does that skill translate to what you're doing on the job? It never really translated to that. It was more about evaluating like does somebody think?
3 April
23 March
-
korrents.com
AI language models are capable of evaluating good fiction even though they cannot yet write it themselves.Their words
An LLM may not be able to write a good story yet, but it can already evaluate them.
20 March
-
Their words
So, we're now in a situation where suddenly people can generate thousands of theories for a given scientific problem. And now we have to to verify them, evaluate them and this is something which we we have to to change our structures of science to actually sort this out.
9 December 2025
-
Their words
We don't really have much evolved experience in evaluating postal efficiency, do we? Okay, we have quart of a million half a million years of evolved experience in deciding who to like and trust because for most of our evolutionary, you know, existence, that was one of the most fi five most important questions to get right.
25 November 2025
-
Their words
There is no such thing as revisionist history. Writing history is a constant process of re-evaluation of sources and attempts to control for the biases of the past as well as our own.
20 August 2025
17 July 2025
From one piece Assessing Adaptive World Models in Machines with Novel Games (with 13 co-authors) 2 beliefs, in the piece's order there
-
Their words
We expect that building and evaluating AI systems capable of this kind of rapid world model induction will be critical for achieving robust, general AI capable of functioning effectively in the complex and fast-changing real world, and especially in human worlds - the environments that human beings have evolved in, created, and are continually changing and re-creating.
-
Their words
We contend that games provide particularly rich and controlled environments uniquely well-suited for systematically evaluating rapid model adaptation and the process of world model induction.
22 June 2025
13 May 2024
-
Their words
The time when benchmarks lasted multiple decades has passed. Going forward, we will rely less on public benchmark results.
7 May 2024
From one piece The case for ensuring that powerful AIs are controlled 3 beliefs, in the piece's order there
-
Their words
Because evaluating control just requires evaluating capabilities, it's far easier to robustly evaluate than alignment.
-
Their words
The basic problem with evaluating alignment is that no matter what behaviors you observe, you have to worry that your model is just acting that way in order to make you think that it is aligned.
-
Their words
So when evaluating control, we should count catching an AI red-handed as a win condition.
1 May 2023
-
Their words
Good tools let the user choose when to switch between implementation and evaluation. When I work with a chatbot, I'm forced to frequently switch between the two modes.
5 November 2019
-
Their words
To make deliberate progress towards more intelligent and more human-like artificial systems, we need to be following an appropriate feedback signal: we need to be able to define and evaluate intelligence in a way that enables comparisons between two systems, as well as comparisons with humans.
13 June 2018
17 April 2009
Nothing matches.
What is a korrent?
A korrent is a belief a person has stated in their own words: one sentence stating the claim, backed by a quote and a source, kept at korrents.com.
Under a name here, the quoted block is what they actually said. The korrent beneath it is the claim those words support, in korrents' wording — tap it to see the record, its source, and who else holds it.
Nobody here wrote their own korrents. They are compiled from public statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they do, this site shows a machine translation beneath the post, in this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the serif above is exactly what the person published, and it is what to quote them on. A translation can be wrong in ways that matter, especially about tone.
Only the post's own words are translated. A quoted post, a linked article and a belief on korrents.com are left in their original language.