Papers for
urban traffic management teams
Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.
Coordinated lane speed limits and ramp metering cut traffic risks
Coordinated Lane-Level Variable Speed Limits and Ramp Metering for Successive Weaving Segments Considering Merging/Diverging Risks: A Hybrid Model Predictive Control and Multi-Agent Reinforcement Learning Approach
Abstract: Successive weaving segments (SWSs) on urban expressways are bottlenecks prone to recurrent congestion and collisions, requiring fine-grained active traffic management (ATM). Existing approaches struggle to balance the adaptive performance of data-driven optimization with the resilience and transferability of model-based control. We propose a hybrid framework to coordinate lane-level variable speed limits (VSLs) and ramp metering across SWSs. First, we reconstruct L-METANET, a lane-level macroscopic traffic flow model that captures free and forced lane changes. Second, we combine XGBoost-SHAP with a random-parameters binary logit (RPBL) model to derive analytical equations for merging and diverging collision risks and formulate system cost and reward functions. Third, we develop MPC-STMAPPO, a hierarchical controller integrating model predictive control (MPC) and multi-agent reinforcement learning (MARL). Its upper MPC layer uses L-METANET for long-horizon rolling optimization and generates baseline commands; its lower spatiotemporal MAPPO (ST-MAPPO) layer, enhanced with Mamba cells and graph attention, produces residual actions for short-horizon adjustment. Real-world experiments on the 18-km Eastern Expressway in Changchun, China, show that L-METANET accurately reproduces lane-changing-induced flow redistribution and capacity drops, with state evolution aligned with ground truth. XGBoost-SHAP-RPBL achieves AUCs above 0.80 in most tasks, outperforming conventional logit models. MPC-STMAPPO converges faster and performs better across multiple metrics than MPC- and MARL-based baselines. Under randomly fluctuating demand, it also significantly outperforms pure MARL in generalization, demonstrating strong potential for industrial deployment.
Spatiotemporal prediction improves by masking ambiguous sensor data
Addressing Spatial Indistinguishability in Spatiotemporal Prediction via Optimal Transport-Guided Masking
Abstract: Spatiotemporal prediction aims to learn discriminative representations from correlated temporal signals over spatial structures for accurate future inference. A central challenge is \emph{spatial indistinguishability}: different nodes may share similar historical patterns yet evolve toward divergent futures, severely degrading forecasting performance in real-world sensor networks. Existing embedding-based and graph neural network (GNN)-based approaches can partially detect such ambiguous nodes but rely on historical similarity, struggling to capture \emph{future behavioral divergence}. We propose \textbf{STOT} (\textbf{S}patio\textbf{T}emporal \textbf{O}ptimal \textbf{T}ransport), a self-supervised framework that resolves spatiotemporal ambiguity via structured masking guided by optimal transport. Our key idea treats indistinguishability as a \emph{disambiguation} problem: future states are inferred by exploiting concurrent spatial correlations and their time-varying similarity. We design a similarity-aware metric for dynamic inter-node relationships and an optimal transport-based masking strategy to emphasize ambiguous positions during pre-training. A batch consistency constraint preserves semantic coherence, while a random-walk masking mechanism promotes structured context exploration. Experiments on six real-world datasets show that STOT performs competitively with state-of-the-art baselines on the evaluated benchmarks and improved interpretability through transport-plan visualizations.
Directional traffic light recognition improves city driving decisions
One Perception, All Maneuvers: Directional Traffic Signal Understanding for Maneuver-Level Signal Intent Prediction
Abstract: Traffic lights are a key regulatory signal for autonomous driving at urban intersections, yet existing traffic signal perception is still predominantly formulated as instance-level detection or color recognition. Such formulations identify where traffic lights are and what colors they display, but leave a critical semantic gap before downstream planning: which ego maneuver is controlled by each visible signal and what dynamic permission the signal expresses for that maneuver. In this paper, we formulate Directional Traffic Signal Understanding, a decision-oriented task that predicts structured signal states for straight, left-turn, right-turn, and U-turn maneuvers from a front-view image. Each state contains the associated signal color and signal-implied passability. Based on OpenLane-V2, we provide a direction-level benchmark with maneuver-level supervision and metrics for color recognition, passability, full-frame consistency, and safety-critical errors. A direction-aware baseline combines global context, localized traffic-light evidence, and maneuver-specific representations. Experiments show that direction-level modeling improves passability prediction over image-level classifiers and detection-oriented pipelines, particularly at complex multi-signal intersections. The resulting representation provides a direct and interpretable traffic-signal interface for downstream planning together with topology, route, and surrounding-agent information.
ITS Fairy improves vehicle safety by supplying missing object info
ITS Fairy: Occlusion Assistance Selected Against a Recipient's Own Perception Reports
Abstract: Cooperative perception can expose object state beyond a vehicle's onboard sensors, but sensing occlusion can still leave a local safety application without the objects its collision computation needs. To tackle this challenge, we present the ITS Fairy, an infrastructure-side Server Local Dynamic Map (S-LDM) service whose decision unit is the pair (recipient, missing conflict-relevant object): among objects absent from a recipient's CPM-derived reported awareness, it sends only those relevant to a Time of Closest Approach (TCA) conflict test. Comparable services predict what a vehicle can perceive; the ITS Fairy instead reads what it has already reported. The recipient inserts the selected state into its local LDM and uses its unchanged collision-avoidance controller. We evaluate this application-level mechanism in SUMO--ms-van3t--S-LDM emulation, since extended as VaN3Twin, using a sensing-occluded lane merge and four-way intersection scenario. At every main-sweep speed, the smallest assisted per-encounter minimum TCA exceeds the largest local-only value in the archived data. Additionally, assisted medians remain in the multi-second range where local-only operation repeatedly approaches zero. In the lane-merge robustness data, the median benefit persists at 80% configured assistance omission with 10 and 5 Hz analysis, but largely disappears at 100-120 km/h when 80% omission is combined with 1 Hz analysis. These results demonstrate the application-level value of supplying object state selected against what a recipient has itself reported. They are not a vehicular wireless-channel evaluation, and they do not quantify what selectivity saves relative to forwarding every nearby object.
Deep reinforcement learning optimizes 5G resources for vehicle communication
A DRL-Driven Optimization of RAN Slice Resource Partitioning for V2X SLA Compliance in 5G Networks
Abstract: Vehicle-to-Everything (V2X) communications impose very demanding requirements in terms of latency and reliability, which must be met in scenarios where multiple services with diverse performance targets coexist. In such scenarios, traffic-intensive services compete for limited radio resources, complicating the fulfillment of V2X service demands. Within this context, Network Slicing (NS) emerges as a key factor that enables the creation of multiple slices and the allocation of resources among them to satisfy heterogeneous service requirements. In particular, this work addresses the Radio Access Network (RAN) slicing problem from the perspective of Physical Resource Block (PRB) partitioning under high traffic demand conditions. To this end, a reinforcement learning approach based on Proximal Policy Optimization (PPO) is proposed to determine PRB allocations that satisfy the strict latency and reliability requirements of V2X services, while improving resource utilization efficiency and minimizing performance degradation of enhanced Mobile BroadBand (eMBB) services. The proposed solution is evaluated through simulation-based experiments under various traffic loads and different V2X service requirements, demonstrating its ability to adapt resource partitioning to network conditions and service demands.
Multimodal system improves emergency vehicle recognition with audio and video
Multimodal Emergency Vehicle Classification via Audio-Visual Transformers and Knowledge Distillation
Abstract: Emergency vehicle detection in autonomous driving is a safety-critical perception task that demands robustness under diverse and adverse real-world conditions. Existing approaches rely on a single modality, either audio or video, which leads to systematic failure when that modality is degraded: microphone-based systems fail in noisy urban environments, and camera-based systems fail at night or under occlusion. This report presents AVNet, a multimodal audio-visual transformer that classifies emergency vehicles (ambulance, fire engine, police car) and road background using both audio and video, while gracefully handling the absence of either modality at inference time. AVNet introduces three key contributions: (1) a temporally aligned cross-modal fusion module that performs second-level cross-attention between audio spectrogram tokens and video frame tokens, exploiting their exact temporal correspondence without any learned alignment mechanism; (2) learned null embeddings that substitute for missing modality tokens, enabling a single unified model to operate in audio-only, video-only, or joint audio-visual mode without retraining; and (3) a knowledge distillation training strategy in which specialist unimodal teacher models transfer inter-class dark knowledge into the multimodal student fusion branch via soft probability targets. Evaluated on 281 clips from the Google AudioSet dataset, AVNet achieves 66.6% overall accuracy in audio-visual mode, outperforming the audio-only branch by +10.4% and the video-only branch by +15.0%. The largest per-class gain is observed for the hardest class, Ambulance, where fusion achieves +29.5% over either unimodal branch alone, demonstrating that the two modalities provide complementary information that the aligned cross attention mechanism successfully exploits.
Vectorized maps forecast beyond vehicle view for safer driving
Generation of Vectorized Maps Beyond Vehicle View
Abstract: Autonomous driving relies on High Definition (HD) maps for safe navigation. Traditional HD maps construction is costly in hardware, data and human resources, which together with its update limitations hinders scalability. Recent works have proposed online alternatives for HD vectorized mapping from onboard sensors. However, sensor field of view is limited, and the range of the reconstructed maps ahead of the vehicle is insufficient for safe planning. This paper aims to address this limitation by proposing the novel beyond-view vectorized map generation problem: given vectorized maps of the area sensed by the vehicle (in-view), to generate plausible map continuations. To experimentally assess its feasibility, we propose BeyondFormer, which, to the best of out knowledge, is the first work designed towards beyond-view map generation. Given the novelty of the problem, we generate the first dataset specifically designed for it and evaluate the proposed approach. The results demonstrate consistent performance across diverse scenarios, establishing learning-based methods as a promising direction for map forecasting in autonomous driving. Beyond demonstrating the feasibility of the task, we provide an extensive discussion of the method's limitations and identify key future research directions for scaling it to more complex driving conditions. Code is available at https://git-autopia.car.upm-csic.es/beyondformer.