Related posts
Dillon Mulroy x.com
“And I'm getting really sick of having to re-evaluate this stuff every two months... TBH a "pause" might be nice for that reason alone” co-signed
The the words it uses that this site has seen least often elsewhere. Posts are matched on those words alone — nothing here is a summary of this one.
Top people
Showing Profile →
Hiding
Hiding
3 September
31 August
-
korrents.com
Getting intensely worked up over how someone else lives is a sign to re-evaluate yourself, not them.Their words
It's time for a re-evaluation if you're getting real worked up over how someone else is living.
28 August
17 August
-
Their words
so I think as long as you have some kind some concept which need not be put into words some way to think about and differentiate the different emotions that are going on in you that gives you a purchase on your ability to to some extent evaluate them and say this is an emotion I want or this is an emotion that I don't want
12 August
-
Their words
kind of going back to this reliability question, we took this policy and we ran it not just once, but we ran it for 13 hours straight. Uh and we basically wanted to evaluate is this policy not only good at making a latte once, but can it do so reliably to the extent that it would be needed to be useful in the real world?
30 July
-
Their words
Um, another way you can get more experience for yourself is to just write down a bunch of things you think might be important in the next 12 months. And maybe you pick one of them to work on, but go back and evaluate in 12 months of these other things, which ones actually seemed important or which ones did other people in the world go out and and create and which ones did they did not seem to do yet.
26 June
From one piece Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown 3 beliefs, in the piece's order there
-
Their words
The problem is we're in a world now where the capability of the model is a function of how much money you put into it. Basically, if you give it a budget of $10,000, it can do a lot more than what it can do with a budget of $10. Give it a budget of $10 million, you can do even more. At what budget should you evaluate these models? The policies that exist today don't really address that question.
-
Their words
my claim is the proper way to evaluate the models now is you either have some kind of budget for the benchmark whether it's tokens or cost or time or whatever or you plot the performance as a function of the amount of test time compute that's going into the model
-
korrents.com
The only way to know what a model can do after running for a month is to actually run it for a month.Their words
And the problem is if you want to evaluate the capabilities of a model, what it can do after running for a month, the only way to be fully sure is to actually run it for a month. And if you want to know after 6 months, the only way to know fully is to run it for six months.
24 June
-
Their words
And the second reason that's stayed sticky is that companies just have no idea how to evaluate. Like And they probably never did.
16 June
-
Their words
They fell because half of what happens in the world is never in our control. And you can do everything right and it's out of your control, but we have to evaluate what would have happened and therefore we should imitate them, because everything they did was right.
23 March
From one piece 🌻 why LLMs are bad writers but good editors 2 beliefs, in the piece's order there
-
korrents.com
AI language models are capable of evaluating good fiction even though they cannot yet write it themselves.Their words
An LLM may not be able to write a good story yet, but it can already evaluate them.
-
korrents.com
Good writing is hard to train into AI models because writing quality is hard to evaluate.Their words
"Good writing" is hard to train because it's hard to evaluate.
20 March
-
Their words
So, we're now in a situation where suddenly people can generate thousands of theories for a given scientific problem. And now we have to to verify them, evaluate them and this is something which we we have to to change our structures of science to actually sort this out.
7 January
-
Recommendsaffiliate linkrcmnd.app
Crossing the ChasmTheir words
The classic technology marketing book. Moore was the first to evaluate the role of early adopters.
22 June 2025
-
Their words
Since these datasets are likely already part of model training data, we shouldn't rely solely on them to evaluate our Q&A system.
19 January 2025
-
Their words
And my contrarian opinion is that full-time jobs are not the best way to monetize the skill that you have. It's one of the packages that everybody should evaluate and take advantage of, but too many people blindly default to that package
16 January 2025
-
korrents.com
AI judges used to evaluate AI outputs need ongoing validation and iteration, just like any other AI system.Their words
AI judges must be evaluated and iterated over time, just like all other AI applications.
7 May 2024
-
Their words
Because evaluating control just requires evaluating capabilities, it's far easier to robustly evaluate than alignment.
1 May 2023
-
Their words
I want to see more tools and fewer operated machines - we should be embracing our humanity instead of blindly improving efficiency. And that involves using our new AI technology in more deft ways than generating more content for humans to evaluate. I believe the real game changers are going to have very little to do with plain content generation.
26 April 2022
5 November 2019
-
Their words
To make deliberate progress towards more intelligent and more human-like artificial systems, we need to be following an appropriate feedback signal: we need to be able to define and evaluate intelligence in a way that enables comparisons between two systems, as well as comparisons with humans.
10 March 2010
Nothing matches.
What is a korrent?
A korrent is a belief a person has stated in their own words: one sentence stating the claim, backed by a quote and a source, kept at korrents.com.
Under a name here, the quoted block is what they actually said. The korrent beneath it is the claim those words support, in korrents' wording — tap it to see the record, its source, and who else holds it.
Nobody here wrote their own korrents. They are compiled from public statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they do, this site shows a machine translation beneath the post, in this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the serif above is exactly what the person published, and it is what to quote them on. A translation can be wrong in ways that matter, especially about tone.
Only the post's own words are translated. A quoted post, a linked article and a belief on korrents.com are left in their original language.