Incredible new investment tool via @ckaiwu ETF screener Fund comparisons Compare a fund to its own history and my favorite...analyze a random fund! RIP to dozens of websites
I think it’s entirely plausible that screens have a real negative effect on teenagers, but studies can’t find them because (a) the effects need not be dose-dependent (maybe two hours on your phone isn’t worse than one hour), and (b) given that everyone else is on social media, giving it up can have a negative effect.
Performance tests are highly specific and dependent on the exact machine it runs on, the exact third party libraries and their versions that are used, the other components involved in the tests, such as the servers, and more.
You can't get to product market fit on a supersonic jet by building it and seeing if anybody likes it. You have to analyze very carefully the art of the possible and also what will maximize the market opportunity.
More data for better comparisons are good. Now everybody has to do it, and regulators and the public have to learn to look at only the data, and not individual incidents
You could imagine systematizing that or like having multiple different agents deliberately given different pieces of context and try to like compare and contrast there. Like we we don't have the same level of manipulation on our own context.
I think the reason why it doesn't show up as so much better on the benchmarks is because the benchmarks are being presented, the benchmark results are being presented in the wrong way. They're not controlling for the amount of test time compute that is being used on that benchmark question.
but but if you have good locality where you you clearly stating what you're importing and whatever and and you can analyze just a single source file and from that extract its protocol to the outside world without having to to know anything deeper. Do you do you know what I mean? I think those are important aspects just simply to reduce the size of the of the of the context window and also make it easier to summarize each module in a program, right?
The short, and very much un-sweet, vision of a tariff-bound economy is that it is stochastic. We can observe it, measure it, analyze it, but we cannot really predict it.
In hindsight, the hardest comparisons are in fact the least useful - if I do the fourth most important task before the third most important, that's no big deal, but doing the tenth before the first is a screw-up!
WinMerge just gets better and better. It's free, it's open source and it'll compare files and folders and help you merge your conflicted source code files like a champ.
Ultimately, however, one-shot, or even sometimes zero-shot, seem like the fairest comparisons to human performance, and are important targets for future work.
Holy shit what a ride! Absolutely amazing story of perseverance and leadership, a must read. Whatever struggles you think you are going through simply cannot compare.
We argue that ARC can be used to measure a human-like form of general fluid intelligence and that it enables fair general intelligence comparisons between AI systems and humans.
To make deliberate progress towards more intelligent and more human-like artificial systems, we need to be following an appropriate feedback signal: we need to be able to define and evaluate intelligence in a way that enables comparisons between two systems, as well as comparisons with humans.
This book gave me the tools to analyze a text and identify the reasons why it doesn't work, for example stating the topic of a paragraph only in the middle of it. I found it very useful to consciously analyze my own writing.
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.