EN / 中文

Real-Time Perception Handles Dynamic Road Changes, Yet It Cannot Independently Enable L4 Autonomy

by zhijiazuiqianyan·October 4, 2026

HD maps were once considered critical infrastructure for autonomous driving. However, as perception, computing, and algorithmic capabilities steadily advance, vehicles are increasingly relying on their own ability to acquire and understand the road environment in real time.

Construction zones, temporary road closures, traffic cones, fallen debris, and suddenly appearing pedestrians and vehicles—these dynamic changes are difficult to fully cover using pre-made maps. In contrast, real-time perception can directly inform the vehicle of what is happening right in front of it.

But this raises a question: if vehicles can already understand the road environment in real time, are HD maps still that important? And can real-time perception further support Level 4 autonomous driving?

01. What Exactly Makes Real-Time Perception So Powerful?

Traditional HD maps can provide vehicles with a priori environmental information collected and processed in advance. Relatively stable information such as road centerlines, lane geometry, traffic signs, and intersection topology can be pre-loaded into the maps. During driving, the vehicle uses the positioning system to determine its location on the map, and then combines this with real-time perception to make driving decisions. However, because real-world roads change dynamically, this approach is not always reliable.

A road might be open to traffic today but have barriers added tomorrow due to construction; an intersection might maintain the same structure long-term, but temporary construction could alter lane organization; and traffic cones, fallen rocks, broken-down vehicles, and pedestrians that did not originally exist cannot all be pre-loaded into the map. Therefore, an autonomous driving system must possess the ability to promptly detect these elements through onboard sensors, even if the map does not indicate what has happened. This is the core value of real-time perception. Real-time perception is not merely about how many frames per second a camera captures.

True real-time perception requires converting data generated by sensors such as cameras, LiDAR, and millimeter-wave radar into usable environmental information for the vehicle with very low latency. This must include the position, speed, and motion state of targets, road structure, behavior of traffic participants, and environmental changes.

As autonomous driving technology evolves, modern autonomous perception is no longer satisfied with simple object detection. The system not only needs to know that there is a vehicle ahead, but also needs to continuously track its motion state, determining whether it is changing lanes or might suddenly cut into the ego lane; it not only needs to identify a traffic cone ahead, but also needs to judge whether this implies a change in the road. This is why today's autonomous perception increasingly emphasizes temporal information, scene understanding, and predictive capabilities.

What the vehicle sees is not a series of independent images, but a continuously changing 3D dynamic environment. From this perspective, real-time perception is indeed diminishing some of the tasks previously handled by HD maps. Especially regarding dynamic road information, the vehicle's own real-time perception has a natural advantage over static maps. However, an easily overlooked distinction must be noted here. The fact that real-time perception can detect changes not recorded in the map does not mean it can completely replace the map.

02. Does This Mean Maps Are No Longer Important?

HD maps and real-time perception actually address two different problems. Maps are better at describing relatively stable road priors, such as road topology, lane connectivity, and fixed traffic facilities; real-time perception is better at describing the environment the vehicle is actually facing at this moment. If the two are simply understood in terms of one replacing the other, it is easy to fall into two extremes. One view holds that HD maps will eventually disappear and all information should be perceived by the vehicle in real time; the other view argues that Level 4 must rely on maps. In reality, Level 4 does not mandate which specific technical route must be adopted.

SAE J3016 defines the ability of an automated system to perform the dynamic driving task within a specific ODD (Operational Design Domain), rather than stipulating that the system must use HD maps, LiDAR, or a specific perception algorithm. Autonomous Driving Frontier believes that a more reasonable direction might be for real-time perception to handle the dynamic world, while maps or other prior information handle the stable world, with both jointly constituting the vehicle's cognition of the environment.

For example, the vehicle knows that there is highly likely an intersection ahead, which is prior information; but upon actually entering the intersection, if it finds that one lane is under construction, it must rely on the road state perceived in real time. Similarly, the lane topology of a road will not change in a short period, but temporary traffic control, construction zones, and the motion states of traffic participants can change at any time.

The map provides what the environment should be like, while real-time perception provides what it is like right now. The real issue is actually that when the two conflict, the system must be able to determine which one is more trustworthy. This essentially pushes the technical focus of autonomous driving further from the question of whether to have a map to how to build a continuously updated environmental model.

This is also one of the important reasons why technologies such as Occupancy Networks, temporal perception, and world models have attracted attention in recent years. They attempt to enable the vehicle to no longer just output a series of isolated bounding boxes, but to continuously maintain an internal world representation capable of expressing spatial structure, dynamic targets, and environmental changes. But even if this step is achieved, real-time perception still cannot independently support Level 4. Because perception answers what the environment is, while autonomous driving must also answer what to do next.

03. What Is the True Critical Bottleneck for Level 4?

Suppose a Level 4 vehicle discovers a suddenly appearing construction zone ahead through real-time perception. The perception system can identify the location of the construction zone, as well as the positions of traffic cones, vehicles, and pedestrians; but whether the vehicle should decelerate, change lanes, detour, or stop next cannot be decided by the perception system alone.

The system also needs to predict what actions other traffic participants might take, plan its own driving trajectory, consider vehicle dynamics and safety constraints, and finally execute the decision through the control system. If the construction zone causes changes to the original road structure, the system may even need to re-establish the local road topology and determine which spaces are passable.

This means that real-time perception is only a critical link in the Level 4 closed loop, not the entirety of Level 4. Moreover, one of the biggest differences between Level 4 and Level 2 is that the system cannot ultimately hand over critical risks to the driver as a fallback. The SAE definition of Level 4 emphasizes that within a specified ODD, the autonomous driving system performs the dynamic driving task; when the system fails or exceeds its operational design domain, it must also enter the corresponding minimal risk condition according to the system design.

Therefore, what Level 4 truly tests is the reliability of the entire system. Real-time perception must be sufficiently accurate, but it must also account for sensor occlusion, rain and snow, glare, low illumination, dirt, and sensor hardware failures; positioning cannot rely on a single source; perception results need to undergo temporal tracking and multi-sensor cross-validation; planning and control need to be able to handle perception uncertainty; and the system must also possess fault detection, degradation, and minimal risk condition handling capabilities.

This is also why current Level 4 systems truly operating on public roads have not simply bet on a single perception technology. Taking Waymo as an example, its system still uses cameras, LiDAR, and radar for complementary perception, emphasizing the redundancy brought by multi-sensor fusion. Therefore, if one must answer whether real-time perception can support the future of Level 4, the answer should be that real-time perception is highly likely to become an irreplaceable core capability for Level 4, but it is not sufficient to support Level 4 on its own. What Level 4 truly needs in the future is a complete system centered on real-time environmental understanding, while integrating maps or other prior information, positioning, prediction, planning, control, and safety redundancy.

From this perspective, the real change facing HD maps is not simply being eliminated by real-time perception, but rather that its role in the autonomous driving system is changing. In the past, maps undertook a large amount of work in telling the vehicle what the road is like, but by 2026, this trend is being pushed in another direction. Some Level 4 players have already announced that their mapless solutions have achieved scaled mass production, with the system completely abandoning its prior reliance on HD maps.

This means that maps are being downgraded from a necessity to an option for autonomous driving, while real-time perception and AI decision-making are becoming the true cornerstone for the scaled commercial deployment of Level 4. Therefore, the future of Level 4 is highly likely not a choice between the map era or the real-time perception era, but a gradual shift from static prior-driven to a fusion-driven approach combining real-time perception and prior information.

And what truly determines how far Level 4 can go is not how clearly the vehicle can see, but whether it can still know what it can and cannot do, and stably bring the vehicle to a safe state when visibility is poor, the field of view is incomplete, the environment changes, or even when parts of the system experience anomalies. This is the real problem that needs to be solved after real-time perception evolves from a mere perception technology into a foundational capability for Level 4.