I've said like the licensable engine thing kind of was our AI transition already unfortunately and uh I regret to inform you that the news is not probably that positive.
The idea that you can have autonomous cloud labs running experiments on a loop without human intervention, testing hypotheses, running the experiments, seeing the results, creating a new hypothesis, testing it and running it, and ultimately getting to a conclusion, that is in our sights.
we find that the performance on held out tasks decreases dramatically. Whereas if we um just take out a random 20% of the data that's less diverse than the most diverse subset, the performance um only decreases a little bit. And so this suggests that actually having really diverse data plays an important role in enabling it to generalize to new tasks.
And so if nobody knows how much an experiment costs, like every time you like grow up a new cell line or every time we make a new a new probe in in the fab, how much does that loop cost? Nobody knows. Therefore, experiments are free. Um, it doesn't cost dollars. It costs media. And media comes from the fridge.
there are definitely some ideas that are worth funding with $50 million or zero dollars, but not $5 million. Um you won't run the experiment. It'll be really frustrating experience. You'll get an ambiguous outcome.
Always try your stupid ideas if you can do it cheaply and reversibly. Jumping off a bridge is not a reversible decision. Not talking about that. I'm talking about stuff like this where you're just like, "Here's a stupid idea." 99 times out of a 100 it'll fail. But that one time you won't have any competition cuz nobody else is stupid enough to try this idea.
Yeah, what what's the smallest experiment I can run to verify to my own satisfaction and everybody's level of satisfaction is going to be different whether or not this claim is true. That's That's the skill that is suddenly in the last year become a thousand times more valuable is that skill of saying, "What's the least I can do to validate for my to my own satisfaction whether this claim is true?
you're going to run out of things that you're going to do purely in the digital space. At some point you have to go to the universe and you have to ask it questions. Um you have to run an experiment and see what the universe tells you to get back to learn something.
the goal is not to teach the model every possible skill within RL just as we don't do that within pre-training, right? Within pre-training, we're not trying to expose the model to, you know, every every possible you know, way that words could be put together, right? You know, we're it's it's rather that the model trains on a lot of things and then and then it reaches generalization across pre-training, right?
Booth's formally trained (and her grandparents are watercolor artists!) but her use of color is just so free and unexpected, it makes you want to experiment yourself and join along in the fun.
Try a number of things, right? So, experiment with a number of things. Run a number of small experiments. Once you find something that works, double down on it and then keep doing it until it stops working, which is the step that a lot of people skip. Um, and then once it stops working, go back to the start and try a lot of small experiments again.
The system matters just as much as any given experiment. Probably even more. Right. I think starting with a growth model so you have an understanding of how your company grows in the first place and which channels you're going to leverage is critical. You need to make sure that you are instrumenting your product in and out otherwise you're going to run experiments and have wonky results.
And a thousand experiments by itself, like if you just did that, but you didn't learn, you didn't make an impact, and that's kind of a waste of time, right? The whole point of setting a goal is that you can have conversations about what would need to be true to actually hit that goal.
Yeah. So there's a subtlety here. Emerging capabilities don't just come from the fact that internet data has a lot of stuff in it. They also come from the fact that generalization once it reaches a certain level becomes compositional.
but um to make robotic foundation models really work it's not just a laboratory science uh kind of experiment. It's also uh it also requires kind of industrial scale uh building effort like it's it's like it's more like the Apollo program than it is like a science experiment
The general view of science, I think, is that we’re accidents of evolution. When we die, the light blinks out. There’s no more of us. There’s no such thing as the soul. But that’s not a proven point. There’s no experiment that proves that’s the case.
I think there’s an interesting thought experiment here that, we’ve been spending a lot of compute in the pre-training to acquire general common sense, but that seems brute force and inefficient. What you want is a system that can learn like an open book exam.
This experiment started from the observation that despite being critical for the functioning of the Internet—and, by extension, the economy—the role of open-source maintainer has not yet found a sustainable manifestation.
This economics thought experiment is a classic book originally written in 1946. Like most books that follow Nassim Taleb's "all books worth reading are at least 20 years old" rule, it has timeless lessons throughout. It tackles all kinds of ongoing political/economic concepts like rent control, minimum wage, tariffs, "saving industries", and more.
Selling is hard. Building a repeatable model with a team that you attempt to rapidly scale is even harder. Fortunately, Mark Roberge, who was one of the very first employees at HubSpot shares the processes and frameworks he used to build a sales machine. If you like the principles behind Lean Startups, then his data- and experiment-driven approach to sales will be very appealing.
If you ever needed convinced of the importance of why starting with a small experiment is really important this is the book to convince. With interesting stories from Pixar and amongst others, it makes a compelling case.
We argue that solely measuring skill at any given task falls short of measuring intelligence, because skill is heavily modulated by prior knowledge and experience: unlimited priors or unlimited training data allow experimenters to "buy" arbitrary levels of skills for a system, in a way that masks the system's own generalization power.
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.