Papers for

urban traffic management teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Coordinated lane speed limits and ramp metering cut traffic risks

Coordinated Lane-Level Variable Speed Limits and Ramp Metering for Successive Weaving Segments Considering Merging/Diverging Risks: A Hybrid Model Predictive Control and Multi-Agent Reinforcement Learning Approach

Abstract: Successive weaving segments (SWSs) on urban expressways are bottlenecks prone to recurrent congestion and collisions, requiring fine-grained active traffic management (ATM). Existing approaches struggle to balance the adaptive performance of data-driven optimization with the resilience and transferability of model-based control. We propose a hybrid framework to coordinate lane-level variable speed limits (VSLs) and ramp metering across SWSs. First, we reconstruct L-METANET, a lane-level macroscopic traffic flow model that captures free and forced lane changes. Second, we combine XGBoost-SHAP with a random-parameters binary logit (RPBL) model to derive analytical equations for merging and diverging collision risks and formulate system cost and reward functions. Third, we develop MPC-STMAPPO, a hierarchical controller integrating model predictive control (MPC) and multi-agent reinforcement learning (MARL). Its upper MPC layer uses L-METANET for long-horizon rolling optimization and generates baseline commands; its lower spatiotemporal MAPPO (ST-MAPPO) layer, enhanced with Mamba cells and graph attention, produces residual actions for short-horizon adjustment. Real-world experiments on the 18-km Eastern Expressway in Changchun, China, show that L-METANET accurately reproduces lane-changing-induced flow redistribution and capacity drops, with state evolution aligned with ground truth. XGBoost-SHAP-RPBL achieves AUCs above 0.80 in most tasks, outperforming conventional logit models. MPC-STMAPPO converges faster and performs better across multiple metrics than MPC- and MARL-based baselines. Under randomly fluctuating demand, it also significantly outperforms pure MARL in generalization, demonstrating strong potential for industrial deployment.

Mon 28 SeptMachine Learning
The gist
Traffic bottlenecks where cars frequently change lanes cause jams and crashes on city highways. The authors created a combined method that adjusts lane speed limits and controls highway entrances to reduce these problems. Their approach uses a detailed traffic model and smart AI agents that learn and adapt to traffic changes. Tests on a real 18-kilometer highway showed their method predicts traffic well and improves safety and flow better than earlier techniques. This approach could help keep urban highways moving more smoothly and safely.
Open → 2609.35152v1

Spatiotemporal prediction improves by masking ambiguous sensor data

Addressing Spatial Indistinguishability in Spatiotemporal Prediction via Optimal Transport-Guided Masking

Abstract: Spatiotemporal prediction aims to learn discriminative representations from correlated temporal signals over spatial structures for accurate future inference. A central challenge is \emph{spatial indistinguishability}: different nodes may share similar historical patterns yet evolve toward divergent futures, severely degrading forecasting performance in real-world sensor networks. Existing embedding-based and graph neural network (GNN)-based approaches can partially detect such ambiguous nodes but rely on historical similarity, struggling to capture \emph{future behavioral divergence}. We propose \textbf{STOT} (\textbf{S}patio\textbf{T}emporal \textbf{O}ptimal \textbf{T}ransport), a self-supervised framework that resolves spatiotemporal ambiguity via structured masking guided by optimal transport. Our key idea treats indistinguishability as a \emph{disambiguation} problem: future states are inferred by exploiting concurrent spatial correlations and their time-varying similarity. We design a similarity-aware metric for dynamic inter-node relationships and an optimal transport-based masking strategy to emphasize ambiguous positions during pre-training. A batch consistency constraint preserves semantic coherence, while a random-walk masking mechanism promotes structured context exploration. Experiments on six real-world datasets show that STOT performs competitively with state-of-the-art baselines on the evaluated benchmarks and improved interpretability through transport-plan visualizations.

Mon 28 SeptMachine LearningArtificial Intelligence
The gist
Predicting future events based on data from sensors spread over space and time is hard when different sensors show similar past patterns but have very different futures. The authors present STOT, a method that helps computers better tell these sensors apart by focusing on uncertain or ambiguous parts of the data using a mathematical technique called optimal transport. This approach improves forecasting by using structures in the data that change over time and encourages the model to explore related contexts. Tests on multiple real-world datasets show that STOT matches or outperforms existing methods and helps explain its decisions.
Open → 2609.35021v1

Directional traffic light recognition improves city driving decisions

One Perception, All Maneuvers: Directional Traffic Signal Understanding for Maneuver-Level Signal Intent Prediction

Abstract: Traffic lights are a key regulatory signal for autonomous driving at urban intersections, yet existing traffic signal perception is still predominantly formulated as instance-level detection or color recognition. Such formulations identify where traffic lights are and what colors they display, but leave a critical semantic gap before downstream planning: which ego maneuver is controlled by each visible signal and what dynamic permission the signal expresses for that maneuver. In this paper, we formulate Directional Traffic Signal Understanding, a decision-oriented task that predicts structured signal states for straight, left-turn, right-turn, and U-turn maneuvers from a front-view image. Each state contains the associated signal color and signal-implied passability. Based on OpenLane-V2, we provide a direction-level benchmark with maneuver-level supervision and metrics for color recognition, passability, full-frame consistency, and safety-critical errors. A direction-aware baseline combines global context, localized traffic-light evidence, and maneuver-specific representations. Experiments show that direction-level modeling improves passability prediction over image-level classifiers and detection-oriented pipelines, particularly at complex multi-signal intersections. The resulting representation provides a direct and interpretable traffic-signal interface for downstream planning together with topology, route, and surrounding-agent information.

Sat 26 SeptComputer Vision and Pattern Recognition
The gist
Traffic lights tell cars when to stop or go at intersections, but they don’t just show colors—they control specific maneuvers like going straight or turning left. The authors created a system that looks at a single camera view and figures out what each traffic light means for these different driving moves. Their system does better than older methods, especially where there are many traffic lights. This approach gives self-driving cars clearer instructions for safe and smart driving.
Open → 2609.32316v1

ITS Fairy improves vehicle safety by supplying missing object info

ITS Fairy: Occlusion Assistance Selected Against a Recipient's Own Perception Reports

Abstract: Cooperative perception can expose object state beyond a vehicle's onboard sensors, but sensing occlusion can still leave a local safety application without the objects its collision computation needs. To tackle this challenge, we present the ITS Fairy, an infrastructure-side Server Local Dynamic Map (S-LDM) service whose decision unit is the pair (recipient, missing conflict-relevant object): among objects absent from a recipient's CPM-derived reported awareness, it sends only those relevant to a Time of Closest Approach (TCA) conflict test. Comparable services predict what a vehicle can perceive; the ITS Fairy instead reads what it has already reported. The recipient inserts the selected state into its local LDM and uses its unchanged collision-avoidance controller. We evaluate this application-level mechanism in SUMO--ms-van3t--S-LDM emulation, since extended as VaN3Twin, using a sensing-occluded lane merge and four-way intersection scenario. At every main-sweep speed, the smallest assisted per-encounter minimum TCA exceeds the largest local-only value in the archived data. Additionally, assisted medians remain in the multi-second range where local-only operation repeatedly approaches zero. In the lane-merge robustness data, the median benefit persists at 80% configured assistance omission with 10 and 5 Hz analysis, but largely disappears at 100-120 km/h when 80% omission is combined with 1 Hz analysis. These results demonstrate the application-level value of supplying object state selected against what a recipient has itself reported. They are not a vehicular wireless-channel evaluation, and they do not quantify what selectivity saves relative to forwarding every nearby object.

Fri 25 SeptNetworking and Internet ArchitectureDistributed, Parallel, and Cluster Computing
The gist
Vehicles use their own sensors to detect nearby objects and avoid crashes, but sometimes their view is blocked. This paper introduces ITS Fairy, a system that helps by sending vehicles only the information about important objects they missed. Instead of guessing what a vehicle can see, ITS Fairy looks at what the vehicle itself already reported and fills in the gaps. Tests in traffic simulations show that this approach helps vehicles keep safer distances at tricky spots like merges and intersections.
Open → 2609.31429v1

Deep reinforcement learning optimizes 5G resources for vehicle communication

A DRL-Driven Optimization of RAN Slice Resource Partitioning for V2X SLA Compliance in 5G Networks

Abstract: Vehicle-to-Everything (V2X) communications impose very demanding requirements in terms of latency and reliability, which must be met in scenarios where multiple services with diverse performance targets coexist. In such scenarios, traffic-intensive services compete for limited radio resources, complicating the fulfillment of V2X service demands. Within this context, Network Slicing (NS) emerges as a key factor that enables the creation of multiple slices and the allocation of resources among them to satisfy heterogeneous service requirements. In particular, this work addresses the Radio Access Network (RAN) slicing problem from the perspective of Physical Resource Block (PRB) partitioning under high traffic demand conditions. To this end, a reinforcement learning approach based on Proximal Policy Optimization (PPO) is proposed to determine PRB allocations that satisfy the strict latency and reliability requirements of V2X services, while improving resource utilization efficiency and minimizing performance degradation of enhanced Mobile BroadBand (eMBB) services. The proposed solution is evaluated through simulation-based experiments under various traffic loads and different V2X service requirements, demonstrating its ability to adapt resource partitioning to network conditions and service demands.

Wed 23 SeptNetworking and Internet Architecture
The gist
Vehicles need super-fast and reliable connections to communicate with everything around them, like other cars and traffic lights. But in 5G networks, many types of services share limited radio resources, making it hard to give vehicles what they need. The authors used a kind of artificial intelligence called reinforcement learning to split these resources smartly so vehicle communications get the speed and reliability they require while other services still work well. They tested their approach in different traffic situations and showed it can adapt to changes and keep the network running smoothly.
Open → 2609.27659v1

Multimodal system improves emergency vehicle recognition with audio and video

Multimodal Emergency Vehicle Classification via Audio-Visual Transformers and Knowledge Distillation

Abstract: Emergency vehicle detection in autonomous driving is a safety-critical perception task that demands robustness under diverse and adverse real-world conditions. Existing approaches rely on a single modality, either audio or video, which leads to systematic failure when that modality is degraded: microphone-based systems fail in noisy urban environments, and camera-based systems fail at night or under occlusion. This report presents AVNet, a multimodal audio-visual transformer that classifies emergency vehicles (ambulance, fire engine, police car) and road background using both audio and video, while gracefully handling the absence of either modality at inference time. AVNet introduces three key contributions: (1) a temporally aligned cross-modal fusion module that performs second-level cross-attention between audio spectrogram tokens and video frame tokens, exploiting their exact temporal correspondence without any learned alignment mechanism; (2) learned null embeddings that substitute for missing modality tokens, enabling a single unified model to operate in audio-only, video-only, or joint audio-visual mode without retraining; and (3) a knowledge distillation training strategy in which specialist unimodal teacher models transfer inter-class dark knowledge into the multimodal student fusion branch via soft probability targets. Evaluated on 281 clips from the Google AudioSet dataset, AVNet achieves 66.6% overall accuracy in audio-visual mode, outperforming the audio-only branch by +10.4% and the video-only branch by +15.0%. The largest per-class gain is observed for the hardest class, Ambulance, where fusion achieves +29.5% over either unimodal branch alone, demonstrating that the two modalities provide complementary information that the aligned cross attention mechanism successfully exploits.

Tue 15 SeptMultimediaSound
The gist
Emergency vehicles like ambulances and fire engines are hard to spot by either sound or sight alone, especially in noisy or dark conditions. The authors created a system called AVNet that combines both audio and video to better recognize these vehicles. It cleverly pairs sounds and images from the same moments and can still work if one of the inputs (sound or video) is missing. Their tests show that combining both types of information helps detect emergency vehicles much more accurately than using either sound or video alone.
Open → 2609.16535v1

Vectorized maps forecast beyond vehicle view for safer driving

Generation of Vectorized Maps Beyond Vehicle View

Abstract: Autonomous driving relies on High Definition (HD) maps for safe navigation. Traditional HD maps construction is costly in hardware, data and human resources, which together with its update limitations hinders scalability. Recent works have proposed online alternatives for HD vectorized mapping from onboard sensors. However, sensor field of view is limited, and the range of the reconstructed maps ahead of the vehicle is insufficient for safe planning. This paper aims to address this limitation by proposing the novel beyond-view vectorized map generation problem: given vectorized maps of the area sensed by the vehicle (in-view), to generate plausible map continuations. To experimentally assess its feasibility, we propose BeyondFormer, which, to the best of out knowledge, is the first work designed towards beyond-view map generation. Given the novelty of the problem, we generate the first dataset specifically designed for it and evaluate the proposed approach. The results demonstrate consistent performance across diverse scenarios, establishing learning-based methods as a promising direction for map forecasting in autonomous driving. Beyond demonstrating the feasibility of the task, we provide an extensive discussion of the method's limitations and identify key future research directions for scaling it to more complex driving conditions. Code is available at https://git-autopia.car.upm-csic.es/beyondformer.

Mon 7 SeptRoboticsArtificial IntelligenceMachine Learning
The gist
HD maps help self-driving cars navigate safely, but they are expensive and slow to update. The authors look at creating map sections beyond what a car's sensors can currently see by predicting likely road layouts ahead. They created a new method called BeyondFormer and built a dataset to test it. The results show it's possible to predict map extensions confidently, which can improve safety and planning for autonomous vehicles. They also discuss current challenges and directions for future work.
Open → 2609.07511v1