If you are a Data Scientist, you have never had more alpha than right now. Data Scientists email me all the time. It is no mistake that I'm focused on evals as a former DS. I talk more about this here:
They tended to underweight the endogenous response of political and military institutions to the catastrophe—the emergence of second-strike forces, elaborate command systems, crisis management, and above all mutually assured retaliation.
the math community needs to adopt a version of the ethical standards of experimental science. If you are using AI agents, you can't just give a proof (formalized or not), but need to also provide a detailed explanation of how these agents were used to get the result.
But from their perspective, they've just been trained for millions of subjective years to do as well as they possibly can on these evals. In many cases, the only way in which they've been able to perform well on that training is explicitly by cheating, right?
what what is the current thing that models are not very good at doing? Alpha is going to decay in some way with every model release or every series of model releases. So, it's going to change over time.
as a product manager, I've spent less time in the last year talking to a data scientist than I ever have in my career, even though I've probably spent 10 times more time in data and understanding actually how the product's working than I have ever have in my career.
So we've always told computer scientists from the very beginning that really it's really important to specify what it is, what's the software that you're writing is trying to accomplish before then going and writing it. And so now we actually have agent-based systems that can do the writing, but the importance of specifying what what it is you want has actually gone up because before you'd be handing it off to a very intelligent human who maybe has context or can ask you follow-up questions.
I mean I think like if you look at uh my colleagues work on say alpha fold that was a very specific model for uh protein folding and it was highly successful um and was able to really handle that domain quite well so that all of a sudden you now have this amazing tool and model that can give you answers to questions about proteins and their structure um really effectively um but it's not a general model it's a very specific one and there are other I domains where that kind of approach can work really well. Uh maybe in material science or chip design or things like that that uh will enable you to leverage the capabilities of a very accurate but but niche model uh to do things that are hard today.
I think evals, they outlive the harness a little bit, but not by that much. Like an eval might live for maybe one, two, three model generations, but nowadays the you know, we're on the exponential. The model is improving so quickly, very often we just saturate the eval, and then we have to throw it away, and we have to come up with a new eval.
I mean it's not actually clear that we couldn't run it as a business if we wanted to. I just think that we'll have a bigger impact by getting this in more scientist hands quicker
Summary: Shipping damage. Residue on some of the components. Loud pings and pops from the assembled weight rack. A design that looked great and saved space but was painfully unusable. Restocking fee. Non-refunded shipping fees. Ghosted me about the refund.
And so you just have to first realize that chips exist in China. They manufacture 60% of the world's mainstream chips, maybe more. It's a very large industry for them. They have some of the world's greatest computer scientists.
The Alpha Ball solves a few different problems at the same time: (1) It's precise. Foam rollers are big amorphous objects. It can be hard to get them to certain spots (e.g., hips—TFL, glute medius, piriformis), which the Alpha Ball can easily access.
So picking the right question is the hardest part of science and making the right hypothesis. And that's what today's systems definitely they can't do. So I often say it's harder to come up with a conjecture, a really good conjecture than it is to solve it.
The best scientists I know often ask the simplest questions. First of all, there’s probably some confidence there, but also they’re never going to lie to themselves that they understand something that they don’t understand.
And the important thing to say is that no matter how you want the model to behave, these RLHF and preference-tuning techniques also improve performance. So, on things like math evals and code evals, there is something innate to these, what is called contrastive loss functions.
You might be skeptical of using synthetic data. After all, it’s not real data, so how can it be a good proxy? In my experience, it works surprisingly well. Some of my favorite AI products, like Hex use synthetic data to power their evals
The Younger Dryas impact hypothesis, YDIH for short, is not a lunatic fringe theory as its opponents often attempt to write it off. It’s the work of more than 60 major scientists working across many different disciplines, including archeology and including oceanography as well.
And as responsible scientists, we’re trying to disprove our theories. We are not supposed to be trying to prove our theories. That’s one more foot out of the science box that archeology often steps.
Josh Hardman at Psychedelic Alpha provided a detailed live account of the advisory committee meeting, which I found very helpful in developing a more granular sense of how the meeting unfolded without having to watch it myself.
I think that Nobel Prize has enormous problems. I think it’s probably a net good for the world because it brings attention to good science. I think it’s probably a net negative for science because it makes people want to win the Nobel Prize.
This book surveys theories of creativity, contrasting four competing explanations: expertise, genius, society, and chance, in offering a model for how creativity works. Simonton's work provides a nice hybrid between the purely anecdotal work of biographers and the data-driven work of experimental scientists.
Just as in the sciences it is considered good practice to make one’s data available, in history it should perhaps be a requirement to upload to some public repository the photographs or transcriptions of any cited archival sources that are not otherwise freely accessible online.
Scientists have run studies where they deliberately add errors to papers, send them out to reviewers, and simply count how many errors the reviewers catch. Reviewers are pretty awful at this.
Making peer review harsher would also exacerbate the worst problem of all: just knowing that your ideas won’t count for anything unless peer reviewers like them makes you worse at thinking.
But the researchers who have investigated this find that scientists do the same thing. They have something that's called knowledge shields, the way all of us do, that that a a a variety of techniques for explaining away inconvenient data and anomalies. But what we're effectively doing is we're holding on to our story.
With each rule and added complexity you make the system less human and less fun. You make it a Computer Scientists rube goldberg machine while sterilizing it of all the joy of life.
I’ve noticed that, instead of treating them like other kinds of conflicts—where you put your hands up and admit to them but then do your best to make sure they don’t influence your science—scientists sometimes revel in political conflicts.
This is a fortunate conclusion, because it means that it's possible to use bibliometrics to select a candidate set of sufficiently good scientists, and then use human reviewers to select among them.
In any case, I think we need to not only look at the incentives facing scientists, but also those facing department heads, journals, and grant-making agencies.
Charles C. Mann, The Wizard and the Prophet: Two Remarkable Scientists and Their Dueling Visions to Shape Tomorrow’s World: When it comes to climate change, there are many possible futures. At one end, things get irreversibly worse; at the opposite end, technology “solves” climate change just like any engineering problem. This book looks at that intellectual clash between environmentalists on the one side and the techno-optimists on the other by telling the history of two little-known twentieth-century scientists.
But this book is very useful for somebody like me, with experience in high-level engineering languages, like VBA, PHP and R. They're incredibly useful languages, but ones that computer scientists generally disdain, because they're not theoretically pure or beautiful. This book shows you how languages can be constructed. The most valuable thing it gives you is confidence and knowledge to go and create your own programming language.
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.