so I think as long as you have some kind some concept which need not be put into words some way to think about and differentiate the different emotions that are going on in you that gives you a purchase on your ability to to some extent evaluate them and say this is an emotion I want or this is an emotion that I don't want
kind of going back to this reliability question, we took this policy and we ran it not just once, but we ran it for 13 hours straight. Uh and we basically wanted to evaluate is this policy not only good at making a latte once, but can it do so reliably to the extent that it would be needed to be useful in the real world?
when the machine says I'm sorry you're so sick today it's very different from how your friend says it to you because the machine said that because it has learned through pattern when someone tells it I'm sick you should say I'm sorry you're sick instead of I'm so glad you're sick because that data exists.
Um, another way you can get more experience for yourself is to just write down a bunch of things you think might be important in the next 12 months. And maybe you pick one of them to work on, but go back and evaluate in 12 months of these other things, which ones actually seemed important or which ones did other people in the world go out and and create and which ones did they did not seem to do yet.
The problem is we're in a world now where the capability of the model is a function of how much money you put into it. Basically, if you give it a budget of $10,000, it can do a lot more than what it can do with a budget of $10. Give it a budget of $10 million, you can do even more. At what budget should you evaluate these models? The policies that exist today don't really address that question.
my claim is the proper way to evaluate the models now is you either have some kind of budget for the benchmark whether it's tokens or cost or time or whatever or you plot the performance as a function of the amount of test time compute that's going into the model
And the problem is if you want to evaluate the capabilities of a model, what it can do after running for a month, the only way to be fully sure is to actually run it for a month. And if you want to know after 6 months, the only way to know fully is to run it for six months.
There are things that we naturally take pleasure in—such as friendship, sex, and Italian food—and while we can become sick of anything if continually exposed to it in large doses, we don’t get tired of these things if they are properly distributed in time.
They fell because half of what happens in the world is never in our control. And you can do everything right and it's out of your control, but we have to evaluate what would have happened and therefore we should imitate them, because everything they did was right.
Another part of my recent office upgrade was the Sanodesk 72x30 Standing Desk. Overall, I love this desk so far. The bamboo top looks and feels great, and I have plenty of room on the desk to set it up how I want to. Being able to raise it to a standing position is also nice too when I get sick of sitting.
Whenever humans experience situations where we are closer to our animal selves than usual (needing to use the bathroom, having periods, giving birth, raising a baby, being sick), we usually behave as animals do: instinctively, urgently, and impolitely.
So, we're now in a situation where suddenly people can generate thousands of theories for a given scientific problem. And now we have to to verify them, evaluate them and this is something which we we have to to change our structures of science to actually sort this out.
But right now like there's no way to press pause on their biological time. So like what if you had an ambulance of the future, right? Like what if you could take someone who is on their deathbed and um you know find some way to you know just sort of hibernate them basically until um the sort of critical cure for their disease comes online.
And my contrarian opinion is that full-time jobs are not the best way to monetize the skill that you have. It's one of the packages that everybody should evaluate and take advantage of, but too many people blindly default to that package
But then we have the interesting case of in the Postclassic, they shed the idea of kings. They don’t like kings anymore. That’s probably a big part of why the Classic disappearance and the abandonment of all those cities happened. People just got sick of kings.
It's perfectly balanced right on the edge of too much—Incarnate is powerful and lush but you somehow never get sick of it. At once excessive and refined.
I want to see more tools and fewer operated machines - we should be embracing our humanity instead of blindly improving efficiency. And that involves using our new AI technology in more deft ways than generating more content for humans to evaluate. I believe the real game changers are going to have very little to do with plain content generation.
To make deliberate progress towards more intelligent and more human-like artificial systems, we need to be following an appropriate feedback signal: we need to be able to define and evaluate intelligence in a way that enables comparisons between two systems, as well as comparisons with humans.
Presentism is sometimes conscious, but often unconscious, so mindful historians will pause whenever we see something that feels revolutionary, or progressive, or proto-modern, or too comfortable, to check for other readings, and make triple sure we have real evidence.
He is important to read because he's simultaneously critical of his characters and sympathetic to them. It's very difficult, now, to think back into that time, when people had very limited choices. If you chose to oppose the regime, that might mean you couldn't get medicine for your sick mother, your children couldn't go to school and you might get kicked out of your apartment. We don't have to face those choices. Miłosz is very good at explaining them.
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.