From one piece Why Fei-Fei Li Is Betting on Spatial Intelligence 13 beliefs, in the piece's order there
-
Their words
One thing that's under appreciated on the website of the demos is the Stanford demo where Ben showed uh anywhere between 3 to 25 images you can reconstruct that entire Stanford quad. But the thing is we had to show it from aerial view. But every single input image is been standing on the ground taking a picture from the ground. So, everything you see are generated but according to the laws of reconstruction.
-
korrents.com
New view prediction is to spatial intelligence what next token prediction is to language.Their words
we do believe very strongly that next viewpoint prediction is is the equivalent of next token prediction.
-
Their words
So okay, so to take a evolutionary view, right? That new viewpoint prediction is exactly evolution had to solve by making animals move. You you nature give animals eyes. But nature didn't give trees eye. Eyes. Why? Because when you move, you see a new viewpoint.
+ 10 more
-
Their words
I think for me, let's go back to the first principle of intelligence. Intelligence is not sitting there stuck and just seeing something or interpreting something when it comes to space and physical space, right? It's really this uh closing the loop between seeing and experiencing and interaction.
-
Their words
cuz we don't yet have a a frontier foundation model that's robust enough for for uh robotics.
-
Their words
There is also a very important step called randomization. Is that you have to take the same environment and then randomize the conditions. So, the cable doesn't literally only, you know, uh bend this way. It can bend a different way or the box can have different sizes, colors, different lids, and all that.
-
korrents.com
The biggest problem in robotics today is data, not chips; chips become the bottleneck later.Their words
We should zoom out and recognize the biggest problem right now in robotics is actually data. One day it'll be chips, but for now it's data.
-
Their words
I think three of us have total conviction about the scaling law. That that I think we do. I do think the exact architecture choices and data mixtures is where the the devils are in the details.
-
Their words
And I do believe Atlas is a significant step forward because now with every single frame, you have a you can generate an estimate a important piece of information, which is the the view viewpoint, the camera pose. And that is the most critical information one needs about the geometry of the of the space.
-
Their words
So in the on the path to spatial intelligence, generating pixels is definitely a a early step, which we have seen with what you call it gazillions of models. But generating pixels that are truly spatially contextualized and grounded is absolutely another major step. And that is the very hard step that Atlas has taken.
-
Their words
But to do that, a fundamental problem to solve is to understand the geometry and structure and the physics of the space.
-
Their words
Well, spatial intelligence eventually must enable us to both generate what the space is, reason within it, and being able to edit and interact within it.
-
Their words
It's the first time we have a unification of pixel generation and pixel reconstruction. In the world of computer vision this field has been around for more than half a century.