End-to-end (E2E) autonomous driving is widely recognized as the optimal path to achieving Level 3 (L3) autonomy. When E2E first emerged, many in the autonomous driving industry believed it was the most promising route to realizing L3 or even Level 4 (L4) autonomy. However, its black-box nature leaves it merely guessing when facing edge cases. When confronted with extreme situations such as construction zones or overturned vehicles, this shortcoming can be fatal. As E2E technology has evolved, it has given rise to Vision-Language-Action (VLA) and World Model technologies. With the support of these two approaches, will the ability of autonomous driving systems to handle edge cases be improved?
01. VLA Equips Autonomous Driving with the Ability to Understand
VLA, or Vision-Language-Action model, introduces natural language reasoning capabilities based on the end-to-end vision architecture. The direction of VLA emerged because pure vision end-to-end models possess only perception capabilities without understanding; thus, a module capable of thinking was added.
The VLA model consists of three core components: a vision encoder, a Large Language Model (LLM) backbone network, and an action decoder. The vision encoder converts camera feeds into high-dimensional feature vectors, the LLM performs logical processing on these features, and the action decoder translates the reasoning results into physical actions such as steering and acceleration.
Integrating all three within a unified Transformer framework enables the alignment of perception, reasoning, and execution within the same semantic space. This architecture allows the system to do more than just see obstacles and lane lines; it can understand the semantic logic behind these elements. For instance, at a city intersection with mixed pedestrian and vehicle traffic, a traditional end-to-end model might only know that an object is moving somewhere, whereas VLA can understand the intention that a pedestrian is likely about to cross the street.
Li Auto's MindVLA-o1, released at GTC 2026, is an attempt focused on 3D spatial understanding, multimodal thinking, and unified behavior generation. Similarly, NVIDIA's 32-billion-parameter Alpamayo 2 Super also positions reasoning capability as the core selling point of its VLA model.
However, the VLA approach is not without its shortcomings. Zhan Kun, Head of Foundation Models at Li Auto, pointed out at GTC 2026 that current industry VLA solutions face three key pain points: suboptimal alignment efficiency between 3D spatial understanding and semantic reasoning, decision-making delays caused by an overly long transmission pathway for vision-language-action, and the difficulty of covering long-tail scenarios solely by scaling up real-world data.
Li Chuanhai, CTO of Geely Auto Group, also proposed three major limitations of VLA: it can only match standard answers while lacking cognitive understanding of patterns; it relies on limited driving operation data rather than massive internet videos, making it difficult to model the operating laws of the physical world. Weak 3D spatial perception is another issue VLA needs to address. Many early solutions directly adopted 2D vision-language models, which are inherently limited in spatial perception; even when 3D spatial representations are introduced to enhance spatial understanding, it faces the challenge of balancing this with the language model's original reasoning capabilities.
Reasoning latency is equally tricky. The autoregressive reasoning process accounts for most of the latency, and in high-speed scenarios, millisecond-level delays can affect driving behavior. Although architectural innovations from 2025 to 2026 are alleviating this issue, latency remains one of the core obstacles to deploying VLA in vehicles. Furthermore, the high cost of data coverage for long-tail scenarios makes it difficult to exhaust all possibilities through real-world road testing. The massive parameter count of VLA puts pressure on in-vehicle computing power and system costs. Some academic research also indicates that these models are highly sensitive to input perturbations, and their reasoning reliability remains to be verified. These issues demonstrate that while VLA has opened up new directions, there are still many problems to be solved between theory and large-scale reliable deployment.
02. World Models Teach Systems to Anticipate
World models take a different path, bypassing the language intermediate layer to directly model and predict in 3D space. They enable the system to possess proactive deduction capabilities, meaning they internally construct a dynamic, physics-compliant virtual traffic scene. Based on physical causality such as inertia, friction, and motion trajectories, the system can predict the future trajectories of surrounding vehicles, pedestrians, and obstacles, thereby planning the optimal path a few seconds in advance. This capability is particularly crucial for edge cases.
For example, if an overturned truck suddenly appears ahead, a system equipped with a world model can start deducing how this obstacle will move next and what the outcomes of applying different braking forces would be, right at the moment of identification, thus making more rational decisions. In contrast, when facing such scenarios, the VLA approach relies more on behavioral patterns learned from human driving data, lacking the ability to directly model and deduce physical processes. This is precisely the fundamental difference in the reasoning paths of the two when dealing with unknown scenarios.
XPeng showcased the X-World model at CVPR 2026. Its inputs include historical multi-view videos and future ego-vehicle actions, while its output is the visual scene the vehicle might see over a future period. This allows the autonomous driving system to anticipate what the surrounding world will look like if the vehicle executes a specific action next.
Waymo's World Model, released in early 2026, is based on Google DeepMind's Genie 3 architecture and can generate simulation environments containing rare situations such as tornadoes, flooded roads, and even elephants straying onto the road. This generative capability allows the system to encounter in the virtual world those extreme scenarios that are almost impossible to meet in real-world road testing. However, world models are not omnipotent. Their strength lies in low-level physical deduction—predicting how objects will move—but they lack high-level semantic reasoning and an understanding of social rules, meaning they cannot determine how they should react. It knows the vehicle ahead might decelerate, but it is not entirely sure what it should do according to traffic rules in such a situation.
03. The Two Paths Are Converging into One
In the second half of 2026, the integration path of VLA and world models has overshadowed the path of the two competing against each other. XPeng explicitly proposed a dual-pillar architecture at CVPR 2026, where its second-generation VLA and world model jointly constitute the two pillars of the physical world foundation model. The logic of VLA is to learn from humans, namely learning human decision-making habits in complex road conditions from driving videos and instructions; the logic of the world model is to learn from the world, which means performing frame-by-frame predictions on massive unlabeled videos to gradually acquire the dynamics and causal structure of the physical world. The former provides sparse but high-density behavioral supervision, while the latter provides dense physical prediction signals, perfectly complementing each other.
Li Auto, on the other hand, takes a different integration path. In 2025, Li Auto unified spatial understanding, language understanding, and action decision-making into a single model framework, building the VLA Driver Large Model based on three major technology stacks: VLA, world model, and reinforcement learning. The MindVLA-o1 released at GTC 2026 was further upgraded. On the basis of the language model handling semantic understanding, it introduced a predictive latent world model to efficiently simulate future scene changes in latent space. Li Auto defines this architecture as a general-purpose agent for the physical world, where the same set of VLA models can simultaneously control vehicles and robots.
From a technical principle perspective, this fusion indeed enhances the system's ability to cope with edge cases. Traditional end-to-end models can only guess based on probabilities when facing unseen scenarios, whereas the combination of VLA and world models provides the system with two independent sources of information: one from semantic understanding of human driving data, and the other from causal deduction based on physical laws.
When the conclusions from the two paths are consistent, the confidence of the decision will be higher; when the two paths diverge, the system also has more room for redundancy and verification. However, this integration path still faces many issues. The real-time response, physical authenticity, and lightweight deployment of world models are still being broken through; the hallucination problem of large models itself has not been completely solved. Additionally, both VLA and world models require massive computing power and data support, which in itself is a considerable threshold.
04. Concluding Remarks
So, as VLA and world model technologies mature, will it make autonomous vehicles more stable when dealing with edge cases? From the perspective of technological evolution, the answer is affirmative.
VLA makes up for the shortcoming of end-to-end models being unable to understand, and world models make up for the shortcoming of being unable to anticipate. The integration of the two provides the system with more reliable reasoning paths when facing unknown scenarios, rather than just relying on statistical patterns in the training data. However, Autonomous Driving Frontier believes that higher does not mean high enough. The essence of edge cases is that they are rare and inexhaustible. No model, regardless of how large the training data or how many parameters it has, can possibly have seen all possible extreme situations. The role of VLA and world models is to enable the system to make relatively reasonable judgments based on a deeper understanding of semantics and physics when facing unseen scenarios, rather than completely relying on pattern matching to guess a result.
In this regard, the value of these two technologies lies not in eliminating edge cases, but in giving the system more ability to think and less guesswork when facing the unknown. This transformation brings a visible improvement in safety, but there is still a long way to go before achieving true stability and reliability.
#AutonomousDriving #EndToEnd #EdgeCases