Being a computer scientist who refuses to find anything about LLMs interesting right now is a bit like being a geneticist who refuses to find anything interesting about the recently opened Jurassic Park.
The idea that you can have autonomous cloud labs running experiments on a loop without human intervention, testing hypotheses, running the experiments, seeing the results, creating a new hypothesis, testing it and running it, and ultimately getting to a conclusion, that is in our sights.
Like the thing that I think is most likely to be sort of the bottleneck in terms of like the AI are really good at verifiable domains but not not at doing the actual thing is just like big experiments. You only get a few tries um well a few is maybe a bit understated but like basically like historically R&D has been driven by doing near frontier scale experiments and that has been pretty important and like actually doing the one big training run where you decide exactly what to include in that.
And so if nobody knows how much an experiment costs, like every time you like grow up a new cell line or every time we make a new a new probe in in the fab, how much does that loop cost? Nobody knows. Therefore, experiments are free. Um, it doesn't cost dollars. It costs media. And media comes from the fridge.
Um but I think there's no uh you know real impediment to making that be a much more automated loop where the model itself decides it's going to explore or maybe with a nudge from some people uh at the various highest level like oh why don't you try some new ideas around model architectures that incorporate this and then it will go run lots of experiments uh see which ones work and then those will get incorporated at a much more rapid rate
One of my favorite lecture series he has given is The Evidence for Modern Physics where he breaks down the experiments that validate some of the weird laws of physics we have, and what it would take to validate even the weirder ones.
so many of the best practices that your lawyers, your bankers, whoever advisers you have, they're going to be pushing best practices on you that are younger than the trees in your local park.
But okay, they shouldn't actually be enacting these ideas. There is a queue of ideas and there's maybe an automated scientist that comes up with ideas based on all the archive papers and GitHub repos and it funnels ideas in or researchers can contribute ideas, but it's a single queue and there is workers that pull items and they try them out.
you're going to run out of things that you're going to do purely in the digital space. At some point you have to go to the universe and you have to ask it questions. Um you have to run an experiment and see what the universe tells you to get back to learn something.
Trials serve two distinct functions: validation — confirming whether a drug works and is safe — and learning, or generating biological data to refine our understanding of a disease, a compound, and the relationship between the two.
Try a number of things, right? So, experiment with a number of things. Run a number of small experiments. Once you find something that works, double down on it and then keep doing it until it stops working, which is the step that a lot of people skip. Um, and then once it stops working, go back to the start and try a lot of small experiments again.
Jurassic Park by Michael Crichton is a surprisingly weird and weirdly underrated novel given how many copies it sold and the popularity of the blockbuster franchise it spawned. The story weaves together many apparently disparate threads and there are extensive speculative digressions into the biotechnology, business interests, and institutional dynamics that make the park possible and its dissolution inevitable. If the movie is supremely entertaining, the book is supremely thought-provoking.
And then the top down belief is the thing that sustains you when the experiments contradict you. Because if you just trust the data all the time, well, sometimes you can be doing a correct thing, but there's a bug.
But even more important than efficiency, designed experiments can inform about causality, which is very difficult to determine from collected observed data.
I mean in my experience the typical win rate and I hate to use that term for experiments is is often something like 30 to 50%. Like usually you're not actually get like you're trying a bunch of things a lot of hypotheses turn out not to be true. you know, consumer products are very unpredictable like that.
And if I'm starting to see like more and more experiments that are not statistically significant, that may be a signal to me to say, okay, we might have kind of tried to exploit a little bit too far. Like there might not be as much juice to squeeze.
The system matters just as much as any given experiment. Probably even more. Right. I think starting with a growth model so you have an understanding of how your company grows in the first place and which channels you're going to leverage is critical. You need to make sure that you are instrumenting your product in and out otherwise you're going to run experiments and have wonky results.
And a thousand experiments by itself, like if you just did that, but you didn't learn, you didn't make an impact, and that's kind of a waste of time, right? The whole point of setting a goal is that you can have conversations about what would need to be true to actually hit that goal.
The overall workflow is intuitive, especially for those new to formal evaluation processes. The UI guides you through creating datasets, running experiments, and annotating results.
A more how-to pragmatic version of carving your own path. Includes some personal story of reinvention but more on experimentation that challenges our default scripts of success and ambition.
Accepted practice is that for any given model that is a notable advancement, you're going to do two to 4x compute of the full training run in experiments alone.
This is why you want to work in post-training because the GPU cost for training is lower. So you can make a higher percentage of your training runs YOLO runs.
I really believe that the founder led growth is not being popularized enough that you do not need growth teams until you actually can start running experiments on your user base
And so, what we found was basically that there have been 10 billion trillion habitable zone planets in the universe. And what that means is that those are 10 billion trillion experiments that have been run. And the only way that we're this whole process from a biogenesis to a civilization has occurred is if every one of those experiments failed.
On the other hand, the manager-free experiments I’m aware of (e.g. holacracy at Medium and GitHub, or “Choose Your Own Work” at Linden Lab) have all been quietly abandoned or outgrown.
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.