EN / 中文

L3 Admission Standard Decoded: Building Human Driver Safety Benchmark for Automated Driving Emergency Scenarios

by xiaomingshixiong·September 28, 2026

Smart Driving Society · Standard Interpretation

Author | Smart Driving Society, Xiaoming

When an OEM applies for L3 admission, the first question usually asked by the testing agency is: Is your system safer than a human?

This question seems simple but has actually been impossible to answer—when comparing to humans, which "human" are we comparing to? What actions are we comparing? How much better is enough to win? The industry has debated this for many rounds over the past few years, but has never reached a set of executable algorithms.

It was not until the draft for comments on the "Intelligent Connected Vehicles - Construction Method for Safety Boundaries in Emergency Scenarios" (Plan No. 20255713-T-339, draft completed in July 2026) was released that the "extent to which a human driver can evade in emergency scenarios" was modeled into parameters, numbers, and processes for the first time. This article breaks it down and explains it thoroughly.

01 What Problem Does This Standard Solve: Finding a Human Benchmark for "Safety"

Let's start with the background. In August 2026, the mandatory national standard "Intelligent Connected Vehicles - Safety Requirements for Automated Driving Systems" (GB 44721—2026) was officially released and will be implemented on July 1, 2027, directly applying to L3 and L4 passenger and freight vehicles. It requires system safety, but the word "safety" must be translated into something measurable within the standard.

How to test it? The traditional approach is to set up a bunch of no-go zones: no collisions allowed in this scenario, must brake to a stop in that condition. The problem is that scenarios are infinite, and the no-go zones can never be fully drawn. Moreover, these no-go zones are essentially compared against the "regulatory bottom line," not against "humans."

This draft for comments changes the coordinate system: it first measures the limit of human drivers' evasion capabilities in typical emergency scenarios, turns it into a boundary, and then requires the capabilities of the automated driving system to be no lower than this boundary. In the words of the standard, the safety boundary is "the extreme threshold of scenario parameters where the driver can control the vehicle to avoid a collision"—simply put, it is the envelope of human evasion capabilities.

Therefore, what this standard truly solves is a problem the industry has been shouting about for many years but has never had a practical tool for: How safe is safe enough for autonomous driving? For the first time, the answer becomes computable: not comparing against the regulatory bottom line, but against the human upper limit.

02 Core Concepts: What the Safety Boundary in Emergency Scenarios Is and Is Not

Let's first clarify the conceptual boundaries to avoid misunderstandings later.

What it is: A calibration and generalization framework with human drivers as the test subjects. Taking the question of "whether the driver can evade" and starting from three typical emergency scenarios (emergency braking of the preceding vehicle, emergency cut-in, and encountering a stationary preceding vehicle after emergency cut-out), it calibrates human response parameters, then performs full-scale generalization in the scenario parameter space, and draws the dividing line between "can evade" and "cannot evade."

What it is not:

    It is not a performance testing standard for systems like AEB or ESS—it tests humans, not systems; it is not a collision regulation—it does not stipulate the collision result itself, but rather the "limit of what humans can achieve"; it is not a one-time test specification—it is a complete methodology for model definition, calibration, generalization, and use case extraction.

Its most relevant aspect to the industry is: once this human boundary is calibrated, the emergency evasion capabilities of L3/L4 systems will have a comparable reference frame—the system boundary must not be lower than the human boundary. This is a dimension higher than checklist requirements like "meeting no-collision in XX scenarios."

03 Model Definition: Compressing Human Evasion Behavior into a Few Parameters

The standard defines two safety boundary models, both modeled according to the three stages of "perception-decision-execution":

Safety boundary model for the emergency braking scenario. It has three core parameters: decision response time (the delay from the appearance of risk to the driver starting to brake), brake effectiveness build-up time (the time for deceleration to linearly climb from 0 to the maximum value), and maximum braking deceleration. This model parameterizes the entire process from "the driver seeing the danger to the vehicle actually slowing down."

Safety boundary model for the emergency steering scenario. The steering wheel angle is modeled as a sine function: θ = θmax·sin(π/(Tsteer100−Tsteer0)·(t−Tsteer0)). It also has three core parameters: decision response time, steering wheel return-to-center time, and maximum steering wheel angle. Note that the maximum angle is taken as the average of the 10 frames with the largest angles within a 5-second window after the risk appears, to prevent single-frame noise.

From an engineering perspective, the value of this set of models lies in: transforming subjective descriptions like "whether the driver reacts fast enough, brakes hard enough, or steers aggressively enough" into measurable, reproducible physical quantities that can be fed into simulation models. With these parameters, human driver behavior can be "transplanted" into the simulation environment to traverse all dangerous combinations that humans have no chance to test in person.

04 Calibration Methods: The "Division of Labor" Between Real Vehicles and Simulation Is the Most Thought-Provoking Engineering Decision in This Standard

Where do the parameters come from? Using a two-pronged approach: real vehicle calibration and simulation calibration, each covering the three types of scenarios.

How is real vehicle calibration done? Taking the emergency braking scenario as an example: the VUT follows the GVT (target vehicle) at a constant speed, the GVT reaches −6 m/s² braking deceleration to a stop within 0.6 seconds, and the driver performs emergency steering to evade; then, the initial following distance is iteratively adjusted according to the formula until the critical condition of "the minimum distance to the GVT before the front of the vehicle completely leaves its own lane ≤ 0.5 m" is met. The same applies to emergency cut-in and encountering a stationary preceding vehicle after emergency cut-out, all of which force the driver to repeatedly approach the limit and "squeeze" out the boundary.

Key operating conditions for the three scenarios (Category M): emergency braking of the preceding vehicle is calibrated at three speed gears of 40/50/60 km/h; emergency cut-in covers speed combinations of VUT 40/30/50/70 km/h against GVT 5/15/30/45 km/h, with a lateral spacing of 1.6 m, a cut-in lateral speed of no less than 2 m/s, and a lateral acceleration of no less than 3 m/s²; encountering a stationary preceding vehicle after emergency cut-out is conducted at VUT speeds of 30/40/50/60 km/h.

How is simulation calibration done? The driving simulator must have steering wheel and pedal force feedback, a three-lane field of view ahead and behind, a longitudinal simulated maximum deceleration of no less than 12 m/s², a lateral deceleration of no less than 8 m/s², and a system delay of no more than 50 ms; the traffic flow on the left must have no less than 4 vehicles, and the cut-in trigger condition is strictly set at a longitudinal TTC between the VUT and GVT of less than 2 s and greater than 0.7 s.

The most critical fusion rules are here:

    Decision response time—adopts the simulation calibration results; maximum braking deceleration, brake effectiveness build-up time (as well as the return-to-center time and maximum angle in the steering parameters)—adopts the real vehicle calibration results.

Why this division of labor? Because both types of calibration have their own inaccuracies. Real vehicle measurement of decision response time will be polluted by the "learning effect"—drivers running the same scenario repeatedly will become increasingly expectant, testing faster and faster, making the data overly optimistic; simulation measurement of vehicle physical response is limited by model accuracy, and no matter how advanced the simulator is, it cannot reproduce the authenticity of whole-vehicle verification for the true limits of braking and the real feel of steering. Therefore, the standard divides the labor according to the source of error: human reactions are assigned to simulation, and vehicle responses are assigned to real vehicles.

The engineering value of this decision is much higher than listing a bunch of test conditions. It acknowledges the boundaries of each method, rather than pretending that real vehicles or simulations are omnipotent.

05 Boundary Determination: From Limited Physical Testing to Full-Space Generalization

What is obtained from calibration is a few fixed parameters, but the safety boundary must cover the entire scenario parameter space and cannot just answer "human performance under these dozen or so conditions." This step relies on parameter generalization.

Fixed parameters. Including the parameters obtained from calibration (decision response time, brake effectiveness build-up time, maximum deceleration/steering wheel angle), as well as scenario parameters: lane width of 3.5 m or 3.75 m, ground friction coefficient of no less than 0.8, and mass-produced vehicle dimensions with a wheelbase of 2300 mm to 2900 mm. The following time headway is recommended to be taken as THW = 1.06×10⁻⁴·V² − 0.02·V + 2.31.

Generalized parameters. Each of the three scenarios has a set of traversal ranges. Taking the emergency braking scenario as an example:

    VUT speed: 10 to 130 km/h, step size 10; GVT speed: 10 to 130 km/h, step size 10; GVT maximum braking deceleration: −10 to −5 m/s², step size 0.5.

The emergency cut-in scenario is a bit more complex: VUT speed 20 to 130 km/h; GVT relative speed −110 to −10 km/h; relative longitudinal distance at the start of cut-in 5 to 60 m (step size 5) plus 70/80/90/100/120/150 m; cut-in action duration 2 to 10 s (step size 1) plus 12/15 s. The cut-out scenario traverses the longitudinal distance between GVT1 and the stationary vehicle of 5 to 100 m (step size 5) and 110 to 150 m (step size 10), and cut-out duration of 2 to 15 s, under VUT/GVT1 speeds of 10 to 130 km/h.

How is the boundary calculated? All parameter combinations are traversed, and each set runs a complete simulation (simulation frequency no less than 50 Hz, control cycle no greater than 20 ms), recording whether a collision occurs. The results are presented in a scatter plot: red dots indicate a collision between the VUT and the target vehicle, blue dots indicate no collision, and yellow dots indicate collisions between non-target vehicles. The dividing line between the red and blue dots is the safety boundary.

Figure 1: Schematic diagram of collision scatter plot for generalized simulation results (schematic diagram) — the boundary between the red and blue dots is the safety boundary

The engineering significance of this section lies in: the safety boundary is not drawn arbitrarily; it is an empirical result derived from running tens of thousands of parameter combinations. Where the boundary is and how steep it is are entirely determined by the data.

06 Use Case Extraction and Evaluation: How the Boundary Model Itself Is Verified

The boundary model itself is also a model and must be verified. Appendix D provides a complete closed loop:

Where do the use cases come from? Based on multi-source data such as naturalistic driving data and accident databases, three types of emergency scenarios are extracted, and the distribution patterns and boundary values of key parameters (speed, cut-in time, longitudinal distance, braking response time, etc.) are statistically analyzed.

Parameter importance ranking. In the emergency braking scenario, the GVT maximum braking deceleration ranks first; in the emergency cut-in scenario, it is ranked as GVT relative speed > relative longitudinal distance at the start of cut-in > cut-in action duration; in the cut-out scenario, it is ranked as longitudinal distance between GVT1 and GVT2 > cut-out action duration. The ranking determines the order of stratified sampling.

Stratified sampling. First, the 10th percentile, 90th percentile, and the maximum value of the frequency distribution of the speed distribution in each calibration scenario are taken as characteristic values, and the upper and lower bounds are expanded by 5 km/h to form a speed interval; then, according to the importance ranking, for each parameter in order, 10% of its distribution characteristic value is used as the upper and lower bounds of the characteristic interval, extracting layer by layer downwards, and sequentially extracting the parameter values for each test use case.

Evaluation metrics. Four numbers are calculated using a confusion matrix: accuracy (TP+TN)/total, precision TP/(TP+FP), specificity TN/(TN+FN), and system maturity FP/(FP+TP)+TN/(FN+TN). The typical reference value in the standard is: system maturity can refer to 75%.

In other words, whether the boundary model is accurate or not is not just self-proclaimed; instead, use cases are extracted and verified back on real vehicles, scored using the four types of metrics. Only when the maturity reaches over 75% can it be said that "this human boundary is credible."

07 Typical Parameters: Those Numbers That Are About to Become Industry Benchmarks

Appendix B provides the typical calibration results for Category M and Category N vehicles. These numbers are worth remembering for everyone working in intelligent driving:

Category M (passenger vehicles) emergency braking model: decision response time 1.24 s (simulation calibration, real vehicle is only 0.42 s—the difference is exactly the learning effect), brake effectiveness build-up time 0.34 s (real vehicle), maximum braking deceleration −8.07 m/s² (real vehicle).

Category M emergency steering model: decision response time 1.24 s, steering wheel return-to-center time 0.92 s, maximum steering wheel angle −95°.

Category N (commercial vehicles) emergency braking model: decision response time 0.89 s, brake effectiveness build-up time 0.78 s, maximum braking deceleration only −4.50 m/s². The steering model is 0.87 s, 2.50 s, −48°.

Putting these numbers together can reveal a lot of information. A decision response time of 1.24 s means that at 120 km/h, from the appearance of risk to the start of braking by a human, the vehicle will slide out for 41 meters—this is the baseline reference for the emergency evasion capabilities of L3/L4. If the system cannot do better than this number, it will be hard to justify it from a regulatory perspective. Meanwhile, the maximum braking deceleration for Category N vehicles is only −4.5 m/s², reflecting the physical limitations of commercial vehicle mass, braking systems, and load conditions, indicating that the standard does not apply passenger vehicle data to commercial vehicles but calibrates them by category.

Also, pay attention to that 0.42 s vs 1.24 s: the decision response time calibrated by real vehicles is significantly smaller than that of simulation, which exactly confirms that the "learning effect" truly exists. This is also why the standard insists that the decision response time must take the simulation results—otherwise, the boundary would be inflated to an artificially high value by "drivers who get better at evading the more they are tested."

08 Relationship with GB 44721: Human Factors Baseline Enters Mandatory Regulation

Putting this draft for comments into the national standard system, its position is very clear.

GB 44721—2026 "Intelligent Connected Vehicles - Safety Requirements for Automated Driving Systems" is the mandatory top-level requirement, managing the safety bottom line of L3/L4 systems; the three recommended standards, GB/T 41798 for track testing, GB/T 44719 for road testing, and GB/T 47025 for simulation testing, provide methods for validation testing; and the Construction Method for Safety Boundaries in Emergency Scenarios fills in the missing first link—the human performance baseline. Without this baseline, all subsequent tests can only prove that "the system did not crash," but cannot prove that "the system is better than humans."

Its weight can also be seen from the drafting lineup: more than twenty OEMs, testing institutions, and universities participated jointly, including FAW, CATARC, Yinwang Intelligent Technology, Tsinghua University, Geely, Changan, SAIC Testing, Desay SV, Jilin University, Jiangsu University, AITO, Chery, NIO, Xiaomi, Li Auto, XPeng, Foton, GAC, China Merchants Testing, and Zhuoyu Technology.

Currently, this standard is in the stage of soliciting comments, and subsequent procedures for approval and release are yet to follow. Judging by the pace of GB 44721, there will usually be one or two rounds of revisions from soliciting comments to implementation, but the overall framework and core methods are highly likely to remain unchanged.

09 Implications and Boundaries for Engineers

Finally, let's talk about some practical implementation aspects. If this standard goes through the process according to the current framework, its impact on R&D and testing will be structural:

Implication 1: Calibration work must be front-loaded. In the early stages of vehicle model development, human drivers must be organized to perform dual-channel calibration of real vehicles and simulations, establishing a vehicle-model-level human boundary database. This is not about making up for a test before release, but a design input for the entire chain from perception and decision to execution.

Implication 2: Simulation capability becomes compliance capability. Simulation frequencies above 50 Hz, control cycles within 20 ms, driving simulators with force feedback, and generalized batch runs covering all parameter combinations—these used to be "R&D efficiency tools," but in the future, they will be "compliance evidence chains." If a simulation platform cannot perform generalization for emergency scenarios, it cannot even complete the calibration.

Implication 3: Data must be auditable. The original record sheets in Appendix A, the typical parameters in Appendix B, and the use cases and confusion matrix in Appendix D will constitute a complete set of traceable archives. How the calibration process was run, how the parameters were determined, and whether the model is accurate must all withstand re-examination by third-party institutions.

The boundaries must also be made clear. Currently, it is a draft for comments; the condition tables, parameter values, and generalization ranges may be adjusted in subsequent versions; the parameter differences between Category M and Category N are significant, and those working on commercial vehicle projects cannot directly apply passenger vehicle numbers; additionally, it currently covers three types of scenarios: emergency braking, emergency cut-in, and encountering a stationary preceding vehicle after cut-out. More complex interactive scenarios (multi-vehicle, special-shaped vehicles, unstructured roads) will need to be supplemented by subsequent standards.


Source |

National Public Service Platform for Standards Information: "Intelligent Connected Vehicles - Construction Method for Safety Boundaries in Emergency Scenarios" standard plan (Plan No. 20255713-T-339); full text of the draft for comments on "Intelligent Connected Vehicles - Construction Method for Safety Boundaries in Emergency Scenarios" (draft completed on 2026-07-28);

Ministry of Industry and Information Technology: Mandatory national standard "Intelligent Connected Vehicles - Safety Requirements for Automated Driving Systems" officially released (2026-08-04); Xinhua Net: Autonomous Driving Now Has a Safety Admission Baseline (2026-08-07);