New Method Enhances Autonomous Driving Safety Training

Xincong Hu (Nanjing University), Lei Ou (Nanjing University), Maosen Li (Yinwang Intelligent Technology Co., Ltd), Jingtao Zhang (Yinwang Intelligent Technology Co., Ltd), Liguo Hou (Yinwang Intelligent Technology Co., Ltd), Zongzhang Zhang (Nanjing University)· August 12, 2026 View original

Key takeaways

  • TPSP improves autonomous driving safety training with online RL.
  • It generates safety-critical scenes based on policy weaknesses.
  • Selective, policy-aware object perturbation is more effective than uniform changes.
  • Threat-guided optimization enhances learning efficiency and safety performance.

Who benefits

AutomotiveRoboticsLogisticsInsuranceDefense

Summary

Researchers propose Threat-guided Policy-aware Scene Perturbation (TPSP) to improve the safety of autonomous driving policies trained with online reinforcement learning. TPSP generates safety-critical scenes by perturbing objects based on the policy's weaknesses, leading to more efficient and robust safety learning.

Ensuring the safety of autonomous driving systems, particularly those trained with online reinforcement learning (RL), is challenging due to the rarity of dangerous real-world scenarios. A new method, Threat-guided Policy-aware Scene Perturbation (TPSP), addresses this by intelligently generating safety-critical training scenes. Unlike previous approaches that create challenging scenes without considering the evolving policy, TPSP explicitly models the interaction between the policy's behavior and its environment. TPSP uses a policy-aware scene encoder to identify and selectively perturb critical objects within a scene, rather than applying uniform modifications. It then employs a threat-guided optimization strategy, evaluating perturbed scenes based on the difference in threat levels between original and perturbed policy rollouts. This guides the generation of highly informative, safety-critical experiences, significantly improving safety learning efficiency and performance in simulated driving environments.

Why it matters

For autonomous vehicle developers, enhancing safety training efficiency and robustness is paramount for commercial viability and public acceptance. TPSP offers a promising approach to address the long-tail problem of rare dangerous scenarios.

How to implement this in your domain

  1. 1Integrate policy-aware scene generation techniques into autonomous driving simulation environments.
  2. 2Develop threat-level evaluation metrics to guide the creation of safety-critical training data.
  3. 3Focus on targeted perturbation of environmental objects based on current policy weaknesses.
  4. 4Utilize online reinforcement learning with enhanced scene perturbation for continuous safety improvement.

Original post by Xincong Hu (Nanjing University), Lei Ou (Nanjing University), Maosen Li (Yinwang Intelligent Technology Co., Ltd), Jingtao Zhang (Yinwang Intelligent Technology Co., Ltd), Liguo Hou (Yinwang Intelligent Technology Co., Ltd), Zongzhang Zhang (Nanjing University)

"arXiv:2608.10403v1 Announce Type: new Abstract: Reinforcement learning (RL) has shown promising performance in autonomous driving, yet ensuring the safety of online RL policies remains challenging due to insufficient exposure to safety-critical driving scenes. The long-tailed nat…"

View on X

Originally posted by Xincong Hu (Nanjing University), Lei Ou (Nanjing University), Maosen Li (Yinwang Intelligent Technology Co., Ltd), Jingtao Zhang (Yinwang Intelligent Technology Co., Ltd), Liguo Hou (Yinwang Intelligent Technology Co., Ltd), Zongzhang Zhang (Nanjing University) on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses