Sim2Signal Benchmarks Bridge Sim-to-Real Gap in Traffic Control

Ferdous Al Rafi, Susrik Mukherjee, Latika Liladhar Dekate, Jennifer Yawa Lavoe, Huaiyuan Yao, Shlok Mohanty, Longchao Da, Xuesong Zhou, Hua Wei· September 3, 2026 View original

Key takeaways

  • The Sim-to-Real gap is a major challenge for RL in traffic signal control.
  • Sim2Signal is a new benchmark to systematically analyze and mitigate this gap.
  • The gap is decomposed into observation, action, transition, and reward components.
  • Effective mitigation often involves estimating gap changes, not just domain randomization.

Who benefits

Smart CitiesTransportationLogisticsAutomotive

Summary

Sim2Signal is a new benchmark designed to systematically measure and mitigate the "Sim-to-Real" gap in reinforcement learning for traffic signal control. It decomposes the gap into observation, action, transition, and reward components, evaluating 18 mitigation methods across diverse real-world network settings.

This paper introduces Sim2Signal, a novel benchmark aimed at addressing the significant challenge of the "Sim-to-Real" gap in reinforcement learning (RL) applications for traffic signal control. While RL policies often perform well in simulated environments, their effectiveness frequently diminishes when deployed in real-world traffic systems. This discrepancy stems from various sources, including mismatches in sensing, action execution, traffic dynamics, and control objectives. Sim2Signal systematically breaks down this gap into four distinct components: observation, action, transition, and reward gaps, each corresponding to a mismatch in the underlying Markov Decision Process (MDP). The benchmark allows for inducing each gap in isolation under a standardized protocol. Researchers evaluated 18 different mitigation methods on two base controllers across 33 gap settings and 10 calibrated networks derived from five real-world locations. The findings indicate that direct transfer of policies consistently degrades performance across all gap sources. Interestingly, the severity of this degradation does not reliably predict the effectiveness of mitigation strategies. Instead, mitigation success is highly dependent on the specific network and gap setting. The study concludes that the most effective methods generally involve estimating what the gap changes, rather than relying on domain randomization or invariant representations to make policies insensitive.

Why it matters

Urban planners, transportation engineers, and AI developers can use this benchmark to develop and validate more robust RL-based traffic control systems that perform reliably in real-world conditions.

How to implement this in your domain

  1. 1Utilize the Sim2Signal benchmark to evaluate new RL algorithms for traffic signal control.
  2. 2Identify specific Sim-to-Real gap sources (observation, action, transition, reward) relevant to your deployment.
  3. 3Experiment with mitigation methods that estimate gap changes rather than just randomizing domains.
  4. 4Calibrate simulation environments using real-world traffic data to reduce initial discrepancies.
  5. 5Collaborate with researchers to contribute to the benchmark and share findings on effective mitigation.

Original post by Ferdous Al Rafi, Susrik Mukherjee, Latika Liladhar Dekate, Jennifer Yawa Lavoe, Huaiyuan Yao, Shlok Mohanty, Longchao Da, Xuesong Zhou, Hua Wei

"arXiv:2609.01676v1 Announce Type: new Abstract: Reinforcement learning achieves strong traffic signal control performance in simulation, yet policies trained in simulators often fail once deployed in the real world, a failure known as the Sim-to-Real gap. When RL is applied to tr…"

View on X

Originally posted by Ferdous Al Rafi, Susrik Mukherjee, Latika Liladhar Dekate, Jennifer Yawa Lavoe, Huaiyuan Yao, Shlok Mohanty, Longchao Da, Xuesong Zhou, Hua Wei on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses