EN / 中文

Shared Physical Foundation Model: Why Autonomous Driving and Embodied AI Can Tap One Unified Model

by zhijiazuiqianyan·September 28, 2026

Many enterprises are now exploring the use of similar foundation models for autonomous driving and embodied AI. XPeng has incorporated VLA 2.0, Robotaxi, and humanoid robots into its Physical AI layout, proposing to support different intelligent carriers with a foundation model for the physical world. Academic research is also investigating the use of a single model to complete navigation tasks across different carriers such as vehicles, wheeled robots, and drones. Given that cars and robots face different environments, sensors, and modes of motion, why attempt to share a single foundation model?

01. Why Are Cars and Robots Starting to Need Similar Foundation Models?

In the past, autonomous driving and robotics basically developed along their respective technological paths. Autonomous driving deals with roads, vehicles, pedestrians, traffic signals, and complex traffic relationships, requiring continuous environmental perception, scene understanding, behavior prediction, path planning, and vehicle control. The environments robots face are more diverse; in addition to mobility, they involve tasks such as target recognition, grasping, carrying, and manipulation. The two seem vastly different, but from a more fundamental perspective, they are both solving a similar problem: how to enable machines to understand the real world and take actions based on that understanding.

The emergence of Vision-Language-Action (VLA) models has precisely pushed this problem further forward. VLA does not just recognize objects in images; it attempts to link visual information, task goals, and actions, enabling the model to move from what it sees to what it should do next. For cars, the model needs to understand the road environment, surrounding traffic participants, and the current driving task, and form driving behaviors accordingly. For robots, the model similarly needs to understand the surrounding space, task requirements, and target objects before deciding to move, grasp, or manipulate.

Therefore, what is truly worth sharing is not a complete set of car models or robot models, but more fundamental capabilities such as visual understanding, spatial relationship understanding, task reasoning, and the ability to predict environmental changes. This cross-carrier sharing has begun to appear in specific research. For instance, the Embodied Navigation Foundation Model (NavFoM), accepted at ICLR 2026, attempts to use a unified architecture to handle navigation tasks across different carriers such as quadruped robots, drones, wheeled robots, and vehicles, including autonomous driving scenarios.

02. Understanding the Same World Does Not Mean Executing the Same Actions

Although both cars and robots need to perceive, understand, and act, their action spaces are completely different. The motion of a car is subject to multiple constraints, including vehicle size, steering characteristics, dynamics, road structure, and traffic rules. For systems adopting end-to-end or VLA architectures, the model can directly generate the vehicle trajectory for a future period, or it can first form driving intentions, leaving trajectory generation and control to subsequent modules.

Robots, on the other hand, are different. Humanoid robots can control body, arm, and hand joints; robotic arms need to handle the spatial position of the end effector; and wheeled robots mainly solve mobility and navigation problems. Even when facing the same object, the action methods of different robots may be completely different. Therefore, research on unified models should not be simply understood as whether a single model can handle both cars and robots simultaneously, but rather as investigating exactly which layer the model can be unified to.

Semantic information such as the presence of obstacles ahead, targets located on the left, or current paths being blocked has strong commonality across different physical carriers. For a car, this may mean decelerating, stopping, changing lanes, or detouring, requiring the generation of trajectories that comply with vehicle motion constraints. For a wheeled robot, it may mean replanning the movement path. For a robot with a robotic arm, it may also mean changing the manipulation method or directly moving the obstacle.

Therefore, what large models unify is the part close to world understanding and task reasoning. The closer it gets to the action generation and control layer of specific carriers, the more it needs to be adapted according to its own physical structure, dynamics, and action space. In the field of robotics, this approach has already been practically implemented. For example, X-VLA introduces information on different robot hardware, sensor configurations, and data domains through soft prompts, allowing the same foundation model to adapt to different robots, rather than requiring all robots to have exactly the same action space. This also highlights an important issue: foundation models can be unified as much as possible, but different physical bodies cannot be ignored.

03. The Real Challenge Is Enabling the Model to Understand What It Can Do

The most difficult part of sharing foundation models between cars and robots is not environmental understanding, but embodied differences. The model needs to know not only what the external world is like, but also its own state, mobility capabilities, and the limitations on its actions. For instance, when both see an obstacle ahead, cars and robots process this information differently. A car driving on a road needs to combine road boundaries, traffic participants, the vehicle's own size, speed, and vehicle dynamics to determine whether decelerating, stopping, or detouring is feasible.

It cannot simply interpret avoiding an obstacle as moving to the side, because there might be other vehicles nearby, or there might be no drivable space at all. Some robots, however, may have a richer set of action choices. It can stop, turn around, or change its path, and some robots with manipulation capabilities can even use robotic arms to handle obstacles. But these capabilities also depend on the robot's specific structure, joint configuration, payload capacity, and task requirements.

This also means that so-called physical intelligence cannot be understood solely as understanding the external world. It also needs to answer multiple questions such as: What is my current state? What capabilities do I possess? Which actions are executable? What changes might occur in the environment after executing this action?

The same command to "move forward" does not represent exactly the same action for a car, a wheeled robot, and a humanoid robot. They have different physical structures, mobility capabilities, sensor configurations, and ways of interacting with the environment. Therefore, if a foundation model serving both autonomous driving and embodied AI truly emerges in the future, it needs to handle not only a unified world representation but also incorporate the states, capabilities, and action spaces of different physical carriers into the model system.

This is also why current research and industrial exploration in Physical AI focus simultaneously on foundation models, VLA, and world models. Models need to gradually move from understanding the current environment to understanding the relationship between actions and environmental changes.

04. Will a Single Foundation Model Ultimately Do It All?

If "doing it all" is understood as using a completely identical model without any adaptation to directly control cars, robots, or even other intelligent devices, this route is not realistic in the short term. However, if "doing it all" is understood as different physical agents sharing some core capabilities of a foundation model, then this route has already seen relatively clear research and industrial exploration. Cars still need their own vehicle trajectory generation and control capabilities, and robots still need to adapt to their own mobility, manipulation, and execution capabilities. The foundation model may be responsible for higher-level understanding and reasoning, while different physical bodies are responsible for translating these capabilities into specific actions that comply with their own constraints.

In other words, what may truly be unified in the future is not necessarily how to drive and how to walk, but the more fundamental intelligent capabilities behind them. This is also why an increasing number of autonomous driving companies are entering the field of embodied AI. The perception, spatial understanding, prediction, planning, and real-time decision-making capabilities accumulated over the long term in autonomous driving are themselves a set of technological accumulations oriented towards the physical world; robots face the same problems of perception, reasoning, and action in real environments. A single foundation model doing it all does not mean one model controlling all machines, but rather that different physical agents such as cars and robots begin to share a foundation model base for understanding the world and reasoning, while retaining their respective adaptation layers related to physical structure, action space, and control capabilities.

#AutonomousDriving #EmbodiedAI #LargeModels