Runtime-Tunable AI Optimizes Transit Signal Priority

Philip-Roman Adam, Stefanie Schmidtner· July 22, 2026 View original

Summary

This paper introduces a preference-conditioned multi-objective reinforcement learning controller for Transit Signal Priority (TSP) that can be tuned at runtime. The controller balances bus delay reduction with minimizing impact on non-bus traffic, offering operational flexibility without retraining.

Transit signal priority (TSP) systems aim to reduce bus delays but must also manage the impact on general traffic and prevent excessive waits for other vehicles. Current reinforcement learning (RL) approaches for TSP often use fixed reward functions, limiting their adaptability when operational priorities, such as time-of-day or disruption conditions, change. This research presents a novel preference-conditioned TSP controller, denoted as $\pi(a \mid s,w)$. This controller can select the optimal signal phase while adhering to minimum/maximum green times and transition feasibility constraints. Crucially, it can be dynamically tuned at runtime using a preference parameter 'w' to adjust the trade-off between prioritizing buses and minimizing overall traffic delay, all without requiring a complete retraining of the model. Implemented within the IntersectionZoo framework, the system was extended with bus-prevalence augmentation and timetable-based bus insertion for robust training. Experiments comparing it against fixed-time control, rule-based TSP, and fixed-weight PPO specialists showed that the single learned conditioned policy effectively spans a smooth empirical trade-off frontier. It consistently outperformed baselines, maintained constraint feasibility, and demonstrated that non-bus externalities remained limited under moderate bus-priority settings.

Why it matters

Urban planners and traffic management professionals can deploy more flexible and adaptive transit signal systems, optimizing traffic flow in real-time based on changing priorities, leading to improved public transit efficiency and reduced congestion.

How to implement this in your domain

  1. 1Evaluate current traffic signal priority systems for their adaptability to dynamic urban conditions and changing operational goals.
  2. 2Investigate integrating multi-objective reinforcement learning with runtime-tunable preference parameters into your traffic management software.
  3. 3Pilot a preference-conditioned TSP controller in a simulated environment or a controlled intersection to assess its performance and flexibility.
  4. 4Develop interfaces that allow traffic operators to easily adjust preference parameters (e.g., bus priority vs. general traffic flow) in real-time.
  5. 5Collaborate with public transit agencies to define optimal trade-off strategies for different times of day or special events.

Who benefits

Smart CitiesTransportationUrban PlanningPublic Transit

Key takeaways

  • A new RL controller allows runtime tuning of transit signal priority.
  • It balances bus delay reduction with minimizing impact on general traffic.
  • The system adapts to changing priorities without requiring retraining.
  • This offers greater operational flexibility for urban traffic management.

Original post by Philip-Roman Adam, Stefanie Schmidtner

"arXiv:2607.18286v1 Announce Type: new Abstract: Transit signal priority (TSP) requires balancing competing objectives: reducing bus delay while limiting adverse impacts on non-bus traffic and avoiding extreme waits for a subset of vehicles. Existing reinforcement-learning (RL) ap…"

View on X

Originally posted by Philip-Roman Adam, Stefanie Schmidtner on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses