The Big Picture
Learned traffic controllers trained near capacity handle a wider range of traffic demands than a classical rule-based method; decentralized controllers perform nearly as well as a fully informed central controller, and a shared-policy approach can be reused on larger road networks to create self-timed “green waves.”
ON THIS PAGE
Core Insights
Reinforcement-learning-based signal controllers remain stable under a much larger set of traffic loads than the MaxPressure baseline, even when trained at a single busy demand. A centralized controller that sees the whole network gives the lowest travel times, but a fully decentralized controller that only uses local data performs almost as well. A parameter-sharing decentralized controller trains more efficiently and can be deployed on longer corridors than it was trained on, producing coordinated ‘green-wave’ behavior (green-wave behavior) (vehicles passing without stopping) despite no explicit junction-to-junction coordination.
Test your agentsValidate against real scenarios
By the Numbers
1Controllers were trained at 700 vehicles/hour per direction and tested across demand grids up to total flows of 1400 vehicles/hour.
2A parameter-sharing controller deployed on a longer corridor produced a peak zero-stop (no-stop) ratio of ~57% at 700 m spacing between junctions (about a 50 s travel time at free-flow speed).
3A single-junction baseline yields ~93% zero-stop; chaining that performance across 13 independent junctions predicts ~39–43% (0.93^13), which matches observed multi-junction behavior and suggests near-optimal emergent coordination.
Why It Matters
Traffic systems engineers and city mobility teams can use these findings to justify trials of learned signal controllers that generalize beyond their training demands and that can be deployed without full network retraining. AI teams and researchers building multi-agent control systems will care because decentralized and shared-policy approaches offer a practical trade-off between performance, training cost, and deployability on different network sizes.
Key Figures

Fig 1: Figure 1 : Multi-junction corridor simulation set-up

Fig 2: (a) maxPressure
Ready to evaluate your AI agents?
Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.
Learn MoreLimitations
Results come from simulation of a simple corridor: one lane per direction, straight movements only, and Poisson vehicle arrivals, so real-world road geometry and arrival patterns may change outcomes. Parameter-sharing sacrifices some optimality when junctions play different roles, so per-junction tuning might still be needed in asymmetric networks. Robustness to sensor noise, incidents, and non-stationary traffic patterns was not tested and will be critical for real deployments. (Note: non-stationary traffic patterns are relevant to Dynamic Task Routing considerations.)
Full Analysis
An urban corridor with equally spaced signalized intersections was simulated using a traffic micro-simulator. Vehicles arrived on entry links as stochastic flows, and experiments compared a classical MaxPressure controller to three learned controllers: a centralized controller that sees the whole network, a fully decentralized controller where each intersection learns independently from its local observations, and a parameter-sharing decentralized controller that trains a single policy applied to every intersection. Learning used a common policy-gradient method and was performed at a near-capacity demand (700 vehicles/hour per direction) across many long simulations until performance converged. Learned controllers showed substantially larger capacity regions — the range of demand combinations where queues remain stable — than MaxPressure. When demand combinations were within all controllers’ capacity, the centralized learned controller gave the lowest average travel times, but the fully decentralized controller was only slightly worse despite having no global view. The centralized learned controller gave the lowest travel times, while the parameter-sharing approach traded off some performance for much greater training efficiency and, importantly, transferred to longer corridors: when deployed on a 13-intersection stretch it produced emergent green-wave behavior with a peak ~57% of vehicles traversing without stopping at certain link lengths. The work highlights using capacity-region analysis to evaluate robustness, and suggests shared-policy decentralization as a promising path toward scalable, reusable traffic controllers — with the caveat that simulation simplifications and unseen real-world variability still need to be addressed in follow-up work.
Not sure where to start?Get personalized recommendations
Credibility Assessment:
ArXiv preprint with no listed affiliations and very low author h-indices (1–2). No top venue or recognizable institutional signals, suggesting limited credibility per rubric.