If the road ahead is blocked by an obstacle, a human driver will choose to reverse first, back up a bit, adjust the steering, and then continue moving forward. However, for an autonomous driving system, this is not as simple as just adding a reversing maneuver. On September 30, Horizon Robotics released HSD V2.1, introducing full-scenario reversing capabilities that cover scenarios such as passing each other on narrow roads, forward obstruction, and getting unstuck from a standstill. Horizon Robotics stated that this capability involves 3D spatial perception, complex scenario decision-making, and continuous vehicle motion control. This upgrade may seem to merely enable autonomous driving to go backward in addition to going forward, but what truly deserves attention is whether the autonomous driving system can proactively change its strategy and find a new feasible route when the current path is impassable. This is precisely where the true difficulty of full-scenario reversing lies.
01. Why Is Reversing Not Simply the Opposite of Driving Forward?
Vehicle forward and backward movements share the perception, planning, and control frameworks, but once the vehicle enters a reversing state, the relationship between the vehicle's own motion and the surrounding space changes. The system cannot merely determine whether there are obstacles behind the vehicle; it must also combine the vehicle's current posture, steering state, body dimensions, and the positions of surrounding obstacles to determine what kind of feasible trajectory can be formed next. For example, when reversing on a narrow road, turning the steering wheel does not make the entire vehicle simply move backward along a straight line. The trajectories of the front and rear ends are not the same, and the vehicle's posture continuously changes with steering, altering the space it ultimately occupies. Mathematically, this corresponds to the nonholonomic constraint characteristics of the vehicle.
Under the assumptions of no sideslip and simplifying the left and right front wheels into an equivalent front wheel, the velocity direction of the rear axle center is along the longitudinal axis of the vehicle body, and the velocity direction of the front axle center is along the orientation of the equivalent front wheel. The vehicle cannot translate laterally like a crab.
The angle between the velocity direction of the rear axle center and that of the front axle center is the equivalent front wheel steering angle δ. At low speeds, the kinematic bicycle model is commonly used to describe vehicle motion, where the rear axle center velocity v and the front wheel steering angle δ jointly determine the vehicle's position and heading rate of change at the next instant. When reversing, v < 0, and the same front wheel steering angle will cause the heading change direction to be opposite to that when driving forward. This is also one of the reasons why reversing feels opposite to driving forward.
When the steering wheel is turned to its limit, the front wheel steering angle reaches its maximum value, corresponding to a minimum turning radius. The curvature of any planned path cannot exceed this limit; otherwise, the vehicle simply cannot execute it. However, the minimum turning radius is only a kinematic lower bound. Actual planning must also consider the vehicle body swept envelope, inner wheel difference, rear overhang swing, and the safety distance to obstacles. In other words, the reversing trajectory must satisfy vehicle kinematic constraints, collision constraints, and trajectory feasibility constraints. The system must not only know what is around it but also predict what space the vehicle body swept volume will enter once the vehicle starts moving.
02. When to Reverse and How Far?
Knowing whether reversing is possible is only the first step. Determining when to reverse and how far to reverse is a key difference between the forward obstruction scenario and typical parking tasks. Automated parking generally has a relatively clear termination goal, such as entering a designated parking space, and the system needs to plan and adjust the vehicle trajectory around this goal. However, forward obstruction does not necessarily have a predetermined endpoint.
When a vehicle enters a narrow road and finds that it cannot continue forward, the system must first determine whether the current path is still feasible. If it cannot continue forward, it needs to re-evaluate whether to reverse, to what position to reverse, and how to continue moving forward after reversing.
After the vehicle reverses, its own position changes, and the relative relationship between the vehicle and surrounding vehicles and obstacles may also change. The original path may no longer be applicable, and new feasible spaces may emerge. Therefore, the system needs to continuously re-evaluate: can the current path still be taken? If not, how much should it reverse? Where should it reverse to have enough space to adjust the steering? After adjusting, should it continue forward or continue reversing? This transforms reversing from a single vehicle maneuver into a continuous replanning driving process.
03. Why Does Reversing Test End-to-End Autonomous Driving?
The above discusses the difficulties of reversing at the decision-making and planning levels. When these problems are applied to end-to-end systems, the challenges are further amplified. The reason reversing tests end-to-end autonomous driving is not because the system thereby acquires a capability completely different from forward driving. On the contrary, the reversing scenario brings together several underlying problems in autonomous driving simultaneously: spatial understanding, behavioral decision-making, trajectory generation, and continuous vehicle execution.
Horizon Robotics officially positions full-scenario reversing as a concentrated manifestation of the comprehensive capabilities of end-to-end models, pointing out that it requires the system to simultaneously possess precise perception of 3D space, decision-making judgment for complex traffic scenarios, and fine control of continuous vehicle motions. And these problems are even more severe in reversing scenarios than in forward driving.
1) The Gap Between Open-Loop Training and Closed-Loop Deployment
Current end-to-end autonomous driving models typically use imitation learning, trained on offline expert demonstration data. This paradigm often performs well in open-loop evaluation but faces the problem of objective inconsistency between open-loop training and closed-loop deployment. Imitation learning assumes that inputs are independent and identically distributed, optimizing one-step prediction accuracy under the dataset distribution. However, once deployed in a closed-loop environment, the policy affects its own observations; that is, the vehicle's actions change the environmental state, and the environment then generates new observations fed back to the model, forming a cycle of action—environment—new observation—new action.
This mismatch between training and deployment can trigger covariate shift, where small errors in the model push it into states never seen in the training data, and decisions in these states generate new errors, leading to gradual error accumulation. In closed-loop execution, this error accumulation significantly reduces system reliability. The reversing scenario is precisely where closed-loop capabilities are most needed. Because after each reversing maneuver, the relative relationship between the ego vehicle and the environment changes, the system must make new decisions based on new observations rather than simply replaying a pre-planned path. Closed-loop training is closer to real deployment conditions than open-loop training, but closed-loop training itself also faces some issues.
On the one hand, closed-loop training highly relies on high-fidelity sensor simulation, and the visual realism, sensor characteristics, and scenario diversity of sensor-level simulation remain limited compared to real driving data. On the other hand, differentiable closed-loop simulators themselves may introduce shortcut learning, where the model uses gradient flow to non-causally regret past predictions rather than truly learning recovery behaviors. These are difficulties at different levels faced by closed-loop training and should not be conflated into a single causal chain.
2) The Dilemma of Reward Design in Reinforcement Learning
Given the importance of closed-loop training, reinforcement learning becomes a natural solution approach. Horizon Robotics' HSD also introduces a dual-engine architecture of world models and end-to-end reinforcement learning, where the world model is responsible for generating scenarios, and reinforcement learning is responsible for training within those generated scenarios. However, there is a problem with using reinforcement learning for reversing: the design of the reward function is extremely difficult. Designing effective reward functions for model-free reinforcement learning under nonholonomic constraints remains an ongoing challenge, which can lead to severe local minima problems, such as policy paralysis or overly conservative risk aversion.
Specifically, the setting of collision penalties constitutes a highly fragile boundary. If the collision penalty is set too low, the model will continuously sacrifice safety for marginal progress; if set too high, the value function network will quickly succumb to a chicken strategy, where the model learns to safely avoid collisions by freezing in place or executing defensive maneuvers that completely move away from the parking area, shifting the optimization objective from task completion to minimizing the penalty flow. More challengingly, traditional dense reward design also suffers from objective mismatch.
When aggregating spatial distance error, heading deviation, and collision penalties via linear combination, a vehicle far from the goal will be prematurely penalized for heading angle deviation. This can lead to unstable local minima, such as spinning in place near the target area or far away from it. These failure modes are relatively rare in forward driving because the task space for forward driving is more open and has a larger margin for error. However, the space in reversing scenarios is extremely constrained, and minor deviations in reward function design are sharply amplified.
3) Asymmetry Between Forward and Backward Perception
Currently, the perception capabilities of almost all advanced driver-assistance systems are designed around forward perception, as forward perception has the most complete configuration and the richest training data. Once reversing to get unstuck is required, the rear becomes the front, and the system is at a disadvantage at the data distribution level. End-to-end models learn from real driving videos, and the vast majority of these videos involve driving forward, with reversing scenarios accounting for a very small proportion of the training data. This means that when the system needs to reverse to get unstuck, the rear scenario is a rare case in the model's empirical distribution.
4) Instability in Long-Horizon Behavior Switching
The reversing process involves multiple switches between forward and backward movements, and each switch is a jump in behavior mode. The instability at such behavior switching points has been observed in literature on continuous policy learning in autonomous driving. Studies have found that certain policies exhibit bifurcation characteristics between behavior modes. Minor changes in the initial state, even deviations as small as sub-centimeter level, can induce significant jumps in behavior modes, switching driving behavior from one to another. When continuous policies attempt to smoothly interpolate behavior transitions, they may instead cause constraint violations at the switching boundaries. At behavior transition points such as forward/backward switching, the model may output trajectories with correct position points but incorrect headings, or trajectories that are discontinuous across planning cycles.
04. Final Thoughts
What end-to-end autonomous driving truly needs to solve is not making the vehicle memorize a reversing maneuver. The real difficulty lies in whether the system can discover the problem on its own, change its strategy, replan, and continuously execute the new driving strategy when the original route is impassable. From this perspective, reversing is just a specific scenario. What it truly tests is the more underlying capabilities of autonomous driving: whether it can understand the spatial relationship between the vehicle and the environment, whether it can replan after a path fails, and whether it can complete new driving tasks in a continuously changing environment.