EN / 中文

Fast & Slow Thinking: VLA Models Solve Long-Tail Driving Scenarios via Multimodal Semantic Fusion

by zhijiazuiqianyan·October 8, 2026

As end-to-end technology evolves, VLA models have secured a place in autonomous driving. VLA, which stands for Vision-Language-Action models, often sparks curiosity: what exactly is the role of language in VLA? Is it similar to the cabin voice assistants we are familiar with? In fact, the role of language in VLA models extends far beyond conversing with humans. It is not a voice assistant serving passengers, but rather the cognitive hub of the autonomous driving system itself. Between perception and action, language provides crucial understanding and reasoning capabilities. The integration of this capability evolves autonomous driving from statistical conditioned reflexes to cognitive understanding followed by decision-making.

01. From Seeing Pixels to Understanding Scenes

Traditional end-to-end autonomous driving models can accurately perform object detection but struggle to understand the deeper meaning behind scenes. The language module in VLA fundamentally changes this. Through the deep integration of visual encoders and language encoders, the model can align the spatial information perceived visually with the semantic knowledge embedded in language. For instance, the system can not only identify a cone ahead but also infer that it is a construction zone requiring deceleration and preparation for detour, by combining information such as construction signs and lane line changes.

The key to achieving this understanding lies in the fusion of multimodal information. Language encoders (such as GPT-class models) can convert natural language instructions or scene descriptions into high-dimensional semantic vectors, which interact with the spatial features extracted by visual encoders in a unified representation space. It is precisely this fusion that enables VLA models to handle long-tail scenarios. Scenarios such as pedestrians suddenly appearing from gaps between vehicles in narrow residential areas, judging the speed of oncoming vehicles and right-of-way relationships at unprotected left-turn intersections, and temporary detours in construction zones, are difficult to handle relying solely on direct mapping from visual input to driving actions. The internet-scale common sense knowledge and causal reasoning capabilities embedded in the language module perfectly fill this gap.

02. How Language Drives Planning

The most profound impact of language on driving is actually reflected in the decision-making and planning stages. The language module in VLA can perform chain-of-thought (CoT) reasoning, thinking through a series of structured steps in its 'mind' just like a human driver. The AutoVLA model, accepted by NeurIPS 2025, combines chain-of-thought reasoning with the tokenization of physical actions, directly generating planning trajectories through a unified autoregressive generation process. The model designs two modes: fast thinking (outputting only the trajectory) and slow thinking (combining chain-of-thought reasoning). In complex scenarios (such as construction zones and ambiguous intersections), the model activates the slow thinking mode, generating an internal reasoning chain similar to 'The green light at the intersection ahead is about to turn yellow, and there are pedestrians waiting on the left, so I should decelerate and prepare to stop,' before generating the driving trajectory based on this.

In simple scenarios, the fast thinking mode is adopted to improve efficiency. The driving intentions inferred by language in VLA models can also directly participate in and guide trajectory generation. The MindVLA-U1 model, jointly proposed by the MMLab at The Chinese University of Hong Kong, Li Auto, and Tsinghua University, enables the driving intentions predicted by the language side to directly participate in continuous trajectory generation through a Classifier-Free Guidance (CFG) mechanism. Experimental data shows that the model achieved a Route Completion Score (RFS) of 8.20 on the validation set of the WOD-E2E autonomous driving benchmark, while the RFS for human driving reference trajectories is 8.13. This means that the trajectory quality generated by the model in open-loop evaluation surpasses the human driving reference standard for the first time.

03. A Bridge for Explainability and Human-Machine Trust

The black-box problem in autonomous driving has long constrained public trust and regulatory implementation. The introduction of the language module provides a viable technical path to solve this issue. VLA models can output their internal reasoning processes in natural language, generating the basis for decisions. NVIDIA released the open-source VLA reasoning model NVIDIA DRIVE Alpamayo-R1 (AR1) at the NeurIPS conference in December 2025, describing it as the world's first industrial-grade open-source reasoning Vision-Language-Action autonomous driving model. The model innovatively deeply integrates chain-of-thought AI reasoning with path planning technology. In areas with dense pedestrians adjacent to bike lanes, vehicles equipped with AR1 can reason through the chain of thought, complete the collection of driving path data, integrate reasoning trajectories (i.e., the system's explanations for taking specific actions), and subsequently plan the subsequent driving route.

AR1 is built on NVIDIA Cosmos Reason. This kind of explainable AI driver is of great significance for safety verification and regulatory review. The VLA large model equipped on the new upgraded version of WEY Lanshan also provides a CoT reasoning card function, which can present the reasoning process of assisted driving to users in real time, making the reasons for every braking and detour clearly visible.

04. Final Words

The addition of language essentially answers the most core question in driving: whether the vehicle truly understands its environment and the situations it needs to handle. The push of VLA towards mass production indicates that the industry has recognized this path. What needs to be discussed next is no longer whether VLA can be used, but how to use it more intelligently—that is, how to make more critical inferences with limited computing power, making decisions both explainable and trustworthy, and moving from usable to highly effective.