EN / 中文

Beyond BEV and Occupancy: Autonomous Shifts to Spatiotemporal Scene Prediction

by zhijiazuiqianyan·October 8, 2026

A car, a pedestrian, a bicycle—today's autonomous driving systems can accurately identify them all. However, for autonomous driving, the real challenge has never been what is on the road, but what will happen next.

Is the vehicle ahead decelerating to yield to a pedestrian or preparing to turn? Is the adjacent vehicle closing in to maintain its lane or intending to change lanes? Is the pedestrian standing by the crosswalk waiting statically or about to cross?

The questions autonomous driving needs to answer are shifting from what has been seen to what is happening in this scene and how it might evolve. This is also the core reason why perception technology continues to evolve toward scene understanding in 2026.

01. From Object Recognition to Understanding Space and Relationships

Traditional perception systems do not merely solve the single problem of recognizing a vehicle ahead; they simultaneously process three types of tasks: identifying what objects are present (detection), how objects are moving (tracking), and how the road constrains driving (lane lines, drivable areas, and traffic signs). These modules collectively transform sensor data into environmental information to prepare for subsequent decision-making. However, the outputs of these modules are independent of each other at the representation level.

The system can know the speed, position, and heading of a vehicle ahead, yet it lacks a unified framework to organize this information. In real-world scenarios, the spatial and temporal relationships among vehicles, pedestrians, and road structures often determine driving strategies more effectively than isolated object categories. Therefore, in recent years, autonomous driving perception has increasingly emphasized the representation of a unified space.

BEV (Bird's Eye View) is a technical route that emerged under this trend. It transforms information from different camera perspectives into a unified bird's-eye view space, enabling the vehicle to express the positional relationships of roads, vehicles, and other traffic participants within the same coordinate system. The further developed 3D Occupancy attempts to describe which areas in the 3D space around the vehicle are occupied and their semantic meanings.

Today, pure vision-based occupancy networks have been validated in mass-production solutions from companies like Tesla and XPENG, achieving complete modeling of 3D space based on the 2D perception provided by BEV. However, it must be clarified that Occupancy does not equal scene understanding. It can express the current 3D spatial state more completely than traditional bounding boxes, covering more complex and irregular spatial structures beyond vehicles and pedestrians, but it still only answers what exists where in the space, not what is happening among these elements or what will happen in the future. True scene understanding also requires the incorporation of time.

Only after continuously observing the same scene can a vehicle determine whether a target is accelerating, decelerating, stopping, or changing lanes. For example, if a vehicle approaches the edge of the road, it is difficult to judge its true intention based solely on single-frame information, but by combining continuous temporal data, it is possible to better determine its motion trend and possible next actions. The environmental representation in autonomous driving is gradually shifting from simple object recognition to dynamic scene modeling that integrates space and time. It must answer not only what is around now but also how these things are changing.

02. From Understanding the Current Scene to Predicting How the Scene Will Change

If the vehicle ahead suddenly decelerates, a traditional perception system can identify the preceding vehicle and the deceleration. But what driving decisions truly care about is why the vehicle ahead is decelerating. If there is a pedestrian ahead, the risk is completely different from normal car-following; if the preceding vehicle is merely preparing to enter the right-side road, the strategy adopted by the ego vehicle will be equally different.

Therefore, autonomous driving needs to link target states, road structures, historical trajectories, and the behaviors of other traffic participants to form predictions about future scenes. This has also propelled World Models from the laboratory to mass-production vehicles in 2026. NIO has deployed World Models to hundreds of thousands of mass-production vehicles, resulting in an 81.5% month-on-month increase in total autonomous driving mileage after deployment; Huawei ADS 4 generates millions of extreme scenarios for training through a cloud-based world engine, reducing the number of takeovers per 100 kilometers in urban areas to less than 0.5. World Models attempt to enable the system to form an internal modeling capability of the driving environment and its evolution laws, rather than merely recognizing the current frame.

Meanwhile, end-to-end autonomous driving is also promoting further integration among perception, prediction, and decision-making. Traditional autonomous driving divides perception, tracking, prediction, planning, and control into different modules, whereas the end-to-end approach in recent years has begun to enable joint learning of these tasks to reduce information loss and error propagation between modules.

XPENG's second-generation VLA (Vision-Language-Action) large model has been deployed in batches on mass-production models, and Volkswagen has been confirmed as its first customer; the cumulative delivery of DeepRoute.ai's end-to-end mapless solution has exceeded 250,000 vehicles. However, this does not mean that the traditional modular architecture has disappeared, nor does it mean that all mass-production systems adopt a completely unified model. The actual technical routes in 2026 still include different solutions such as modular, end-to-end, and a fusion of both. It is just that more and more systems no longer treat perception results as the final answer, but as part of subsequent prediction and decision-making. Therefore, the shift of autonomous driving from object recognition to scene understanding is a complete technological evolution from identifying objects to establishing unified spatial representations, then to understanding spatiotemporal relationships and predicting scene evolution, and ultimately serving driving decisions.

When the system can know what is ahead, it has already solved the first layer of the perception problem; when it can know the relationships among these things, it begins to enter scene understanding; and only when it further judges how this scene might change next does it truly begin to approach the environmental modeling capability required for advanced autonomous driving.

03. Final Words

In the autonomous driving competition of 2026, the focus is no longer just on who can identify more objects, but on who can organize scattered perception information into a more complete, continuous, and dynamic driving scene, and let this scene understanding truly improve the vehicle's prediction and decision-making. This may well be the true significance of autonomous driving evolving from seeing the world to understanding the world.