Papers for

smart city infrastructure teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Closure-guided communication cuts data and boosts vehicle perception

SemRD-V2X: Closure-Guided Communication with Bounded Inference for Cooperative Perception

Abstract: Vehicle-to-Everything (V2X) cooperative perception improves 3-D detection by sharing intermediate features, but dense remote features may repeat context that the ego agent can infer locally. Most communication-efficient designs optimize masks or codes empirically, leaving a more basic question open: which remote evidence is indispensable given the receiver's own observation? We introduce a closure-fidelity perspective on ego conditioned remote perception. Under a finite deductive abstraction and explicit conditions, its rate--distortion function decomposes over an irredundant core, and the exact zero-distortion rate becomes $P_A H(π_A)$. This analysis suggests a concrete design principle: transmit compact evidence and recover derivable context with bounded receiver-side inference. Guided by this principle, SemRD-V2X is an operational neural proxy that combines exact-budget BEV support selection, pointwise channel compression, and masked shared-weight reconstruction before standard fusion. Experiments on simulated V2XSet and real-world DAIR-V2X validate the resulting design. In a controlled five-run V2XSet comparison against a locally reproduced V2X-ViT-v1 baseline on one Tesla V100, SemRD-V2X reduces the analytical feature payload by $26.6\times$ while improving AP@0.5/AP@0.7 by 4.13/8.57 points, with 3.81\% additional mean compute latency. These results position closure fidelity as both an analytical lens and an actionable design principle for communication-efficient cooperative perception.

Mon 28 SeptArtificial IntelligenceInformation Theory
The gist
Vehicles working together to detect objects around them can send lots of data, often repeating what each car already knows. The authors study how to pick only the most crucial shared information that can't be guessed from a vehicle's own sensors. They design a system called SemRD-V2X that smartly compresses and shares only this essential data, letting each car fill in the rest by itself. Tests show this method cuts data sent by over 26 times while improving detection accuracy and only slightly slowing down computation.
Open → 2609.34353v1

TrafficFab autonomously manages city traffic using edge and cloud AI

TrafficFab: An Autonomic Edge-Cloud Testbed Fabric forAI-Driven Traffic Management

Abstract: Traffic management in emerging megacities requires real-time analytics over thousands of CCTV video streams under latency, bandwidth, compute and energy constraints. We present TrafficFab, an autonomic edge--cloud testbed for AI-driven traffic management, designed to validate a representative slice of a megacity deployment. TrafficFab combines RTSP stream emulation, heterogeneous edge inference using DNNs, cloud-based nowcasting and forecasting using Spatio-Temporal Graph Neural Network (ST-GNN), and continual model adaptation through foundation-model (FM)-assisted Federated Learning (FL). Its autonomic control enables fine-grained scale-out/in of edge inference through energy- and migration-aware scheduling, elastic scale-up/down of GNN forecasting on public clouds, and periodic adaptation of the DNN on edge accelerators and private cloud, without centralized video collection. We evaluate TrafficFab on a Bangalore-city inspired deployment, spanning Raspberry Pis, Jetson accelerators, GPU fogs, private cloud servers, and cloud VMs, sustaining real-time analytics for $\approx 400$ live camera streams (10% of Bangalore) and analytically characterize larger setups. The results demonstrate that TrafficFab offers a practical validation-scale platform for closed-loop traffic analytics, short-term operational decision support, and longer-horizon planning analyses in megacity scales.

Thu 24 SeptDistributed, Parallel, and Cluster Computing
The gist
Managing traffic in huge cities is challenging because it requires very fast analysis of many live video streams from traffic cameras. The authors created TrafficFab, a system that uses smart cameras and cloud computers to watch and predict traffic patterns without sending all video to one place. It automatically controls where and how AI models run on small devices near cameras and big cloud servers to save energy and computing power. They tested TrafficFab using a setup inspired by Bangalore city with hundreds of cameras to show it can work at city scale.
Open → 2609.29223v1

Domain adaptation method improves segmentation in bad weather

ICM: Intra-class Mixing for Domain Adaptation in Adverse Weather

Abstract: Unsupervised domain adaptation (UDA) for semantic segmentation remains challenging under adverse weather conditions because severe appearance changes enlarge the domain gap and degrade the reliability of pseudo labels in the target domain. To address this problem, we propose an Intra-Class Mixing Consistency (ICM) framework that enforces prediction consistency between an intra-class mixed image and its original counterpart. Unlike previous mixing-based consistency methods that combine regions across different images or domains and may introduce unrealistic semantic inconsistencies, ICM performs mixing within the same image and semantic class, preserving realistic semantic layout for consistency regularization. With ICM, we establish a new state-of-the-art performance for clear-to-adverse-weather unsupervised domain adaptation (UDA) in semantic segmentation. On the Cityscapes $\rightarrow$ ACDC benchmark, our method achieves 75.7\% mIoU, outperforming the previous state of the art by +1.9 pp, demonstrating its effectiveness in mitigating class confusion under challenging environmental conditions. The code is provided in the supplementary material.

Wed 23 SeptComputer Vision and Pattern Recognition
The gist
Semantic segmentation is a way for computers to understand images by identifying each part of a scene. This task gets harder in bad weather because the images look very different, making the computer confused. The authors created a new method called intra-class mixing consistency (ICM), which mixes parts of the same object within an image to help the computer learn better without confusing different things. Their method improved accuracy on a well-known test for understanding images taken in bad weather, showing it helps computers see more clearly when conditions are tough.
Open → 2609.27533v1

Trust management challenges and features in edge-enabled IoT security

Trust in Edge-Enabled IoT Security: Features, Challenges and Research Directions

Abstract: Providing autonomous intelligence, pervasive connectivity and usability to human life and industry has led to the emergence of the Internet of Things (IoT). To support time-sensitive and resource-constrained applications, IoT systems nowadays increasingly rely on edge computing. This brings computation and decision-making closer to end devices. In edge-enabled IoT architecture, latency and communication overhead are reduced, but interactions among a larger and more diverse set of devices, edge nodes, services, and data sources are introduced as well. In such environments, security and privacy mechanisms provide the foundation for protection, while trust management can assess the reliability of interacting entities and adapting secure decisions. In this paper, we systematically review the current state of trust management in edge-enabled IoT. To this end, we propose a comprehensive taxonomy that maps physical, network, and application architectural IoT layers against the consumer, commercial, industrial, and infrastructure IoT domains. We further investigate state-of-art research based on their trust design, how trust integrated into secure IoT operations, the attacks that effect trust management process. Based on these findings, we identify key gaps in current research and outline future directions for context-aware and adaptive trust management in edge-enabled IoT.

Mon 21 SeptCryptography and SecurityArtificial IntelligenceComputers and Society
The gist
As many everyday devices connect to the internet through the Internet of Things (IoT), keeping these devices secure becomes more complicated when edge computing is used to process data closer to the devices. This paper by the authors explains how trust management helps decide which devices and services can be trusted in these complex networks. They review existing research on how trust is designed and attacked in edge-enabled IoT and highlight gaps that need more work. Their findings guide future research toward making trust systems more adaptive and aware of different contexts in IoT setups.
Open → 2609.24669v1

Vast toolchain boosts testing of cooperative autonomous driving systems

VAST: V2X/Dynamic Map-Aware Autonomous Driving Systems Validation Toolchain

Abstract: Cooperative autonomous driving in the IoT-to-Edge-to-Cloud continuum requires system-level validation across vehicles, infrastructure sensors, edge-side Dynamic Map services, and in-vehicle autonomous-driving stacks. This paper presents VAST, a V2X/Dynamic Map-aware validation toolchain that connects Scenic, Scenario Simulator v2, AWSIM, Autoware, and SIM-LDM. VAST does not introduce a new search algorithm; instead, it addresses interoperability challenges, including Lanelet2-to-Scenic mapping, ROS 2-based co-simulation through SS2, Dynamic Map object injection into Autoware, and collection of TTC, PET, collision, timeout, and performance measurements. In occluded-intersection scenarios, Lanelet2-compatible constrained sampling increases the edge-case discovery rate from 40.0% to 80.0% and reduces the average time per discovered edge case from 259.7 s to 110.4 s. Under the same generated scenario distribution, Dynamic Map availability reduces the collision rate from 78.0% to 40.0% and increases non-collision outcomes from 22.0% to 60.0%, with statistically significant TTC/PET shifts. A throughput study with 1-16 NPCs shows that sampling remains below 0.1 s, whereas AWSIM/Autoware execution and restart overhead dominate runtime. These results position VAST as a practical validation infrastructure for cooperative autonomous-driving CPSs.

Thu 17 SeptRobotics
The gist
Testing self-driving cars that communicate with other vehicles and traffic systems is complicated because many different software parts and sensors have to work together. The authors made VAST, a toolchain that connects several simulators and autonomous-driving software, to check how these systems behave together in different scenarios. Their tool helps find tricky situations where accidents could happen faster and more often, and shows that having detailed dynamic maps can make driving safer. VAST also runs efficiently even with multiple simulated cars around, making it a useful platform for testing real-world cooperative driving technology.
Open → 2609.19681v1

Multi-view pedestrian tracking improves with fewer cameras

GRACE: Geometry- and Ray-Aware Camera-Efficient Multi-View Pedestrian Tracking

Abstract: Reducing the number of cameras reduces the deployment cost but removes views that correct BEV responses stretched away from true pedestrian positions by projection and short score drops that can split tracks} in Bird's-Eye View (BEV) tracking. We introduce GRACE, a camera-efficient multi-view tracker with three components. Volumetric-Guided Fusion combines homography-based BEV features with features lifted through 3D space. Ray Conditioning exposes each camera's viewing direction to the fusion network. Its tracking component, BEV Track Recovery (BTR), uses low-confidence detections only to continue existing tracks. The same detections cannot start new tracks. With two WildTrack cameras, GRACE improves MOTA from 83.54 for TrackTacular, our baseline, to 91.07.

Tue 15 SeptComputer Vision and Pattern Recognition
The gist
Tracking people using cameras is easier with many cameras, but using fewer cameras cuts costs. The authors introduce GRACE, a method that helps track pedestrians accurately even with just two cameras. It combines different types of 3D and bird’s-eye view features and conditions on each camera’s direction to improve tracking. Their method reduces errors and keeps track of people better than earlier approaches.
Open → 2609.16872v1

Multi sensor fusion improves vehicle network beam prediction accuracy

Robust Beam Prediction for V2X Networks with Multi-Modal Sensing

Abstract: Integrated sensing and communication (ISAC) provides a promising foundation for beam prediction in future vehicle-to-everything (V2X) networks. However, existing sensing-assisted beamforming methods still rely heavily on radio-frequency sensing, which may become unreliable in complex vehicular environments. Meanwhile, the growing availability of heterogeneous sensors, such as cameras and LiDAR, offers new opportunities to improve beam prediction through richer environmental perception. Motivated by this, this paper proposes a multi-modal beam prediction framework for V2X networks. Specifically, we develop BeamTransFuser, a hierarchical Transformer-based architecture that progressively fuses camera, LiDAR, radar, and GPS observations for robust beam prediction. In addition, to handle possible missing modalities in practical deployment, we introduce a generative module that reconstructs missing modality features from the available observations. Experimental results on a real-world multi-modal V2X dataset show that the proposed framework consistently outperforms representative baselines, while the generative module further improves robustness under incomplete sensing conditions.

Wed 9 SeptMachine Learning
The gist
In vehicle communication networks, it's important to focus signals accurately to maintain good connections. The authors noticed current methods rely mostly on radio signals, which can fail in tricky environments. They designed a system called BeamTransFuser that combines data from cameras, LiDAR, radar, and GPS to better guess where to direct signals. Their system also copes when some sensors aren’t available by filling in missing data from the others. Tests on real vehicle data showed this approach predicts signal directions more reliably than before.
Open → 2609.10200v1

LiDAR detection improves urban vehicle and pedestrian tracking

Solution for UCF UrbanTwin V2X-Real Track: Sim-to-Real Urban LiDAR 3D Object Detection

Abstract: Bridging the simulation-to-reality gap in roadside LiDAR requires addressing several coupled discrepancies, including scene geometry, sampling density, return patterns, and pedestrian scale. This report presents a multi-source collaborative training and class-aware fusion framework for Sim2Real 3D detection. The method organizes digital-twin scans, diffusion-redrawn scans, density-stabilized scans, and pedestrian morphology-aligned samples into a unified training pool with complementary roles. Within a common DSVT detection formulation, source-specialized expert branches preserve those roles while optimizing for the same detection objective. At inference, a predefined class-aware fusion pathway integrates geometry-stable and calibration-aware branches for vehicles, sampling-complementary branches for trucks, and morphology-consistent evidence for pedestrians. A label-free point-cloud center blend then refines geometric localization. On the UrbanTwin V2X-Real hidden test set, the unified system achieves a combined score of 0.7421, with 3D mAP@0.5 of 0.4518 and a realism score of 0.8871. The results indicate that a stable, interpretable collaboration among data sources is more valuable than unconstrained aggregation of model outputs.

Mon 7 SeptComputer Vision and Pattern Recognition
The gist
Detecting objects like vehicles and pedestrians using 3D LiDAR scanners in cities is tricky because real-world data differs from simulated data. The authors created a system that combines different types of simulated scans to better match real-world conditions. They use specialized parts of the system for different object types and then fuse the results in a smart way to improve accuracy. Their approach works well on a challenging urban dataset, showing that carefully mixing data sources is better than just combining many model outputs randomly.
Open → 2609.07608v1