Every agent harness can now use Jev npm install -g ai-cli → ask yes/no questions → choose between options → score against your criteria https://t.co/dpgmcdCj8F
We know that arteries with a calcium (CAC) score of zero can be highly inflamed, and that high CAC scores are not necessarily high-risk for cardiovascular events.
Astra is now the best cybersecurity model in the world on @vercel DeepsecBench. It pulls off in 49 minutes what took Sol ~4 hours, with a better score, at nearly the same cost.
the issue that I take with some of the language around "the body keeps the score" is that it maybe doesn't also recognize that the body also carries so much wisdom and so much intelligence and kind of helps guide us towards healing
What we find is that those individuals who score highly on just self-reported questionnaires for motivation, if they score highly for apathy, they have twice the risk of developing Alzheimer's disease than people who don't.
If I just take the general population on average it's the case that those people who would have a diagnosis of autism or who score higher on autistic traits even if they don't have a diagnosis look less at faces.
the score is the classifier's estimated probability for the AI-generated class based on its training distribution. However, we shouldn't interpreted it as a general probability that the text was written by AI.
Roast quality and color can have a big impact on perceived cup score and buying decisions, so if you are serious about choosing the right green for your company, it is a good idea to do your own sample roasting, using a system that yields your preferred roast level consistently.
So, a deployment of your agent uh in the real world generates data. That data then grounds the simulator and makes it more realistic. The simulator generates harder edge cases for the critic to score and for the agent to learn from.
In other words, the absence of calcium is not the same as the absence of disease—it may simply mean the disease hasn’t reached the stage CAC is designed to detect.
Um, another way you can get more experience for yourself is to just write down a bunch of things you think might be important in the next 12 months. And maybe you pick one of them to work on, but go back and evaluate in 12 months of these other things, which ones actually seemed important or which ones did other people in the world go out and and create and which ones did they did not seem to do yet.
Uh but I don't think it's easy but I do think it's something you would have to do in order to uh reward the gawwa like instinct rather than just rewarding have you solved a problem.
Um so if you say okay well we're going to instead of just running this model once we're going to run it five times and take the best of the five responses or like ask a judge which one it thinks is best then you can get much higher scores than that model. And so it's really easy to make something that looks a lot better on paper but is actually not better once you control for the amount of test time compute.
And so I think it's it's really important uh when when we think about benchmark progress to think about it from that perspective, which is benchmarks rise on problems that we've framed that we can articulate, that we can score. And there's a lot of work that's human work that uh it it can't be scored until you write it down
Trustworthiness is the most underrated asset in all of business. And the things that create trustworthiness by definition stack rank to the bottom if we do it by ROI because doing the right thing has intangible rewards but tangible costs.
First of all, I think we are speaking from the point of view of a society which intensely values this particular trait, you know, ability to score well on IQ tests or things like them or to go to school for a long time or whatever it is. And I think this is unprecedented in human history that we live in a time like this.
The ORS is a brief, 4-item measure designed to assess areas of life functioning known to change as a result of therapeutic intervention. It takes less than one minute to complete and score.
Jurassic Park by Michael Crichton is a surprisingly weird and weirdly underrated novel given how many copies it sold and the popularity of the blockbuster franchise it spawned. The story weaves together many apparently disparate threads and there are extensive speculative digressions into the biotechnology, business interests, and institutional dynamics that make the park possible and its dissolution inevitable. If the movie is supremely entertaining, the book is supremely thought-provoking.
And my hunch is that the archaeological evidence that we have on this score is ultimately misleading, because I think this: that probably for a very, very long time before the Sumerians, people in the world, the world of what we call the Middle East, were in contact.
You see this for the music now. You can generate new music. It could generate new stuff that you wouldn't know. For an LLM variance is bad. originality is is is a lower score. So what's the feedback loop for original but good?
What he found is that the companies that struggled to grow almost always had less than 40% very disappointed. Whereas the companies that grew the fastest almost always had more than 40% very disappointed. And this question, this metric is way more predictive of success than something for example like net promoter score.
This is a worthy entry in the genre. It suffers a bit from extensive plot not about banking and characters who barely speak to dragons. But all of that irrelevant fluff makes it one of the most underrated fantasy series I’ve ever read.
If you’re interested in learning more about the history of Africa, its central role in the world, and the ways in which both have been misunderstood in the West, I highly recommend Born in Blackness.
in quote-unquote real life, right? Which is starting businesses or you know, writing books, right? Or composing music or playing basketball or playing poker, right? Or basically doing anything interesting we're in a probabilistic domain, right?
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.