From FSD to Humanoid Robots

Even without an official declaration that “FSD is solved,” FSD has become massively successful. Whether you like Tesla or not, or whether you want to admit it or not, it is difficult to deny the achievement. The remaining obstacles to much broader deployment are increasingly regulatory, legal, and liability-related rather than simply whether a neural network can drive.

That success makes me think more deeply about robot development.

Led by companies such as Figure in the US and Unitree in China, many companies are now investing enormous resources into humanoid robots. At the same time, there are plenty of naysayers who believe general-purpose robots will never really work. So who is right?

To understand this, I think it is useful to look at how autonomous driving itself has evolved.

In traditional autonomous-driving systems, the pipeline was roughly:

camera
↓
feature extraction / matching
↓
3D position
↓
SLAM
↓
SE(3) / Lie-group geometry
↓
planning
↓
steering / braking / acceleration

The machine explicitly reconstructed geometry, estimated its position, built a representation of the environment, and then used mathematical planning and control to decide how to move.

Then came the Transformer.

“Attention Is All You Need” changed the way we think about learning representations. Images can be converted into visual tokens, and Transformers can learn relationships between those tokens. With enough data and compute, the model can learn spatial and temporal relationships directly from video instead of requiring every aspect of the world to be explicitly programmed.

The pipeline starts to look more like:

camera/video
↓
visual tokens
↓
Transformer
↓
learned world representation
↓
learned driving policy
↓
desired trajectory
↓
steering / brake / throttle

The important shift is that the system does not necessarily have to first construct a traditional SLAM map and then reason on top of it. Instead, the neural network can learn a representation of the world that is useful for prediction and action.

And this is where the evolution becomes interesting:

LLM
↓
VLM
↓
VLA
↓
world model
↓
policy
↓
embodied intelligence

A language model learns relationships between words. A vision-language model extends this to visual information. A vision-language-action model connects perception and language to action. A world model represents aspects of the physical environment. A policy maps observations and goals into actions. Put them together and you start getting what we call embodied intelligence: intelligence that can perceive and act in the physical world.

This is why I find humanoid robotics much more interesting today than I did several years ago.

Figure is now showing evidence that the scaling approach is beginning to work beyond carefully scripted demonstrations. In September 2026, Figure reported Helix 2.5 performing tasks such as bed-making, towel folding, and room tidying across 30 previously unseen homes, without environment-specific data collection or fine-tuning. Figure also reported that pretraining on its human-behavior dataset improved zero-shot success from 9% to 56% in its comparison.

More importantly, Figure says its Index dataset is generating roughly 35 minutes of new human experience every second, and it has committed $3.5 billion of compute to Helix training.

This is the part I find most important. The robotics problem may increasingly become a data-and-scaling problem, rather than a problem of manually programming every individual household task.

Figure’s Helix 02 already demonstrates the direction: a unified neural system takes vision, touch, and proprioception and uses them to control the entire body. The manufacturing scale matters too. Figure has reported producing more than 350 Figure 03 robots and increasing production capability toward one robot per hour. More robots mean more physical interaction, more failures, more successful behaviors, and ultimately more training data.

But there is still a huge question. Can a robot operate for thousands of hours, across thousands of environments, with very little human intervention, while being economically useful and safe?

That is a much higher bar than a four-minute autonomous dishwasher demonstration.

And this is where I think FSD provides an important precedent. Autonomous driving was also an extremely difficult physical-control problem. Yet enormous amounts of real-world data, neural networks, and compute have fundamentally changed what is possible. Tesla now describes FSD as an end-to-end foundation-model approach trained on customer and Robotaxi data.

Humanoid robotics is trying to push the same idea much further.

A car has a relatively constrained action space: steering, braking, and acceleration. A humanoid has an enormous action space involving the head, torso, arms, hands, fingers, legs, feet, balance, and contact forces. The environment is also far less structured.

That makes the humanoid problem much harder than FSD. But it also makes the potential breakthrough much larger.

The real breakthrough would not be a robot with thousands of individually programmed skills. It would be a general physical foundation model that learns the relationship between perception, world representation, intention, whole-body action, feedback, and correction.

perception
↓
world model
↓
intention
↓
whole-body action
↓
feedback
↓
correction
↺

If physical intelligence scales in the same way that language and vision intelligence have scaled, then piling up data, robots, compute, and training time could eventually produce something very different from today’s robots.

Leave a Reply