Papers for
urban planners
Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.
Robots as social elements shaping human group experiences
A Robot Among People:From Social Imitation to the Social Becoming of Human Groups
Abstract: Robots designed to mediate human groups often fall into the solutionist trap: they are framed as sociable agents that fix problems such as conflict, disengagement, or lack of coordination. We suggest a different way of thinking. Rather than discrete agents, robots can be understood as situated elements of shared environments; catalysts and carriers of group experience whose meaning emerges through how people position, interpret, and interact with them. From this perspective, robots are not there to repair some ostensible dysfunctionality, but to enable group-level sense-making around care, norms, and identity. Our prior work on robotic street furniture suggests that this does not happen by imitating human sociality but by taking the shape of deliberately constrained, group-facing entities that happen and act for \textit{us} without being socially entangled as one of us. We thus understand robots in public spaces not in terms of autonomy or intelligence, but as a relational capacity. This implies designing robots not in our image or for our utility, but grounded in our needs in being and becoming together.
Strategyproof facility location with learned predictions improves accuracy
Learning-Augmented Strategyproof Facility Location in $\mathbb{R}^d$ with $\ell_p$ Distances
Abstract: We study learning-augmented mechanism design for locating a single facility in $\mathbb{R}^d$ to minimize the sum of the agents' $\ell_p$ distances to the facility. We analyze the coordinate-wise median with predictions (CMP) mechanism, which adds $cn$ virtual agents at a predicted optimal facility location to the $n$ reported locations and returns their coordinate-wise median. The parameter $c\in[0,1)$ represents confidence in the prediction; CMP is known to be strategyproof when both are fixed independently of the reports. We determine its approximation guarantees under correct predictions (consistency) and arbitrary predictions (robustness) in two settings. First, for $d=2$, we establish exact guarantees for every $p\in[1,+\infty]$. For $1<p<+\infty$, the consistency and robustness are respectively $\Bigl(1+\bigl(\frac{1-c}{1+c}\bigr)^{\frac{p}{p-1}}\Bigr)^{\frac{p-1}{p}}$ and $\Bigl(1+\bigl(\frac{1+c}{1-c}\bigr)^{\frac{p}{p-1}}\Bigr)^{\frac{p-1}{p}}$. We use the median condition in each coordinate to compare the mechanism's cost with the optimal cost, and establish tightness using instances with agents at only three distinct locations. Second, for $1<p<+\infty$, we obtain dimension-independent consistency and robustness upper bounds valid for every $d\ge1$. For each fixed $p$ and $c$, we construct families of instances whose ratios approach the respective upper bounds as $d\to\infty$, proving asymptotic tightness. We also establish exact guarantees for $p=1$ in every dimension and asymptotically tight bounds for $p=+\infty$. Our high-dimensional results recover the prediction-free bounds of Gravin and Jia (STOC 2025) when $c=0$ and their learning-augmented bounds in arbitrary-dimensional Euclidean spaces when $p=2$.
Strategyproof facility location on a circle with improved approximation ratio
A Randomized $\frac32$-Approximation for Strategic Facility Location on a Circle
Abstract: We study strategyproof mechanisms for locating a single facility on a circle so as to serve a set of strategic agents under the utilitarian social cost objective. We analyze a simple parity-dependent mechanism that randomizes between the Random Dictator (RD) mechanism and the Proportional Circle Distance (PCD) mechanism. For an odd number \(n\) of agents, this mechanism mixes RD and PCD with equal probability, as originally proposed by Rogowski and Dziubi{ń}ski (IJCAI 2025). For even \(n\), it mixes RD with a random-deletion extension of PCD, in which one agent is removed uniformly at random before PCD is applied to the remaining agents. Our main result shows that this mechanism is strategyproof and achieves an approximation ratio of \(\frac32\) for every \(n\ge 3\). This improves the \(\frac74\) upper bound of Rogowski and Dziubi{ń}ski, which applied only to odd \(n\), and extends the guarantee to even numbers of agents. Finally, we establish a lower bound of \(\frac{11}{10}\) on the approximation ratio of any randomized strategyproof mechanism, improving on the previous lower bound of \(1.0456\) due to Meir (SAGT 2019).
Machine learning predicts pedestrian counts from city map features
Estimating Pedestrian Volumes from GIS-Derived Built-Environment Features: A Machine Learning Framework
Abstract: Transportation agencies need pedestrian volume estimates across entire road networks to prioritize safety investments, yet manual counts are expensive and cover only a small share of intersections. We present a machine learning pipeline that predicts 2-hour PM peak pedestrian volume at 101 urban intersections in Portland, Oregon, from built-environment, land-use, and street-network features drawn from open GIS data. Starting from the Negative Binomial GLM used in practice, we add feature selection, count-aware gradient boosting, and repeated cross-validation, selecting one configuration by a combined rank over RMSE, MAPE, and SMAPE across four cross-validation strategies. The winner, a histogram-based gradient boosting model with Poisson loss and L1 Lasso feature selection, reduces cross-validated RMSE by 12% over the GLM baseline (89.8 to 78.7) and holdout RMSE by 19% (108.0 to 87.9). Code is released on GitHub.
Improved road scene segmentation with RGB and near infrared images
HSI-Road Relabeled: Surface-Aware Road-Scene Segmentation
Abstract: The HSI-Road dataset provides paired RGB and 25-channel NIR (600--960~nm) images with binary masks but no surface-level labels.~This paper introduces a manually labeled six-class taxonomy: Background, Asphalt, Concrete, Dirt, Water, and Grass, and an RGB-to-NIR registration pipeline with corresponding annotations. Six semantic-segmentation models (SSMs) are evaluated under four input configurations: original-resolution RGB (RGB$_{\text{ori}}$), registered low-resolution RGB (RGB$_{\text{reg}}$), NIR, and channel-stacked RGB$_{\text{reg}}$--NIR (RGBN$_{\text{stk}}$). The comparison quantifies the effect of spatial-resolution reduction on RGB, along with evaluation of NIR and RGBN$_{\text{stk}}$, with results reported using per-class and mean IoU and F1 scores. RGB$_{\text{ori}}$ achieves the highest overall performance but contains 12$\times$ more pixels than the matched-resolution inputs. At the matched 192$\times$384 resolution, RGBN$_{\text{stk}}$ outperforms NIR for all six SSMs and RGB$_{\text{reg}}$ for five of six, with the most consistent gains for the Water class. These results highlight the importance of spatial resolution while showing that NIR provides complementary information to RGB.
Perceived food access reveals barriers beyond store proximity
Population-level measures of perceived food access reveal barriers beyond geographic proximity
Abstract: Food access is multidimensional, but population-level measurement still relies heavily on geography because perceived dimensions of access are difficult to measure at scale. Here, we use 25,125 Google Maps reviews from 49 grocery stores in Raleigh, North Carolina, to measure five dimensions of food access: availability, accessibility, affordability, accommodation, and acceptability. We identify review topics with unsupervised topic modeling and assign them to access dimensions using zero-shot classification, with 85.4% agreement against manual coding. The resulting store-level measures capture distinct aspects of food access and reveal barriers that geographic proximity alone does not capture. Comparisons between nearby stores in the same chain further show that identical store policies can be perceived very differently across locations, consistent with food access reflecting the fit between residents and their food environment. Perceived food access also follows systematic socioeconomic and demographic patterns that broadly parallel, but do not replicate, those observed for geographic access. These results show that online grocery reviews can provide a scalable complement to geographic measures of food access.
Digital twins must include human and social system factors
Missing Dimensions: Integrating Human and Social Systems into Digital Twin Engineering
Abstract: Digital twins (DTs) have emerged as a key technology at the core of digital transformation, yet their engineering practice remains too narrowly focused on engineered and natural systems. This paper argues that four system dimensions must be explicitly recognized in DT engineering: Engineered, Natural/Biological, Human, and Social. Each dimension brings distinct properties, modeling requirements, and ethical obligations that fundamentally shape what a DT must represent and how it must be built. We further argue that as DTs extend into the Human and Social dimensions, the emphasis should shift from automated control toward decision support and occupant empowerment. We illustrate our arguments using a smart building as a running example and identify four open research challenges for multi-dimensional DT engineering.
Super-resolution improves digital elevation models using satellite images
Guided Super-Resolution of Digital Elevation Models with Diffusion-Based Image Generators
Abstract: High-resolution digital surface models (DSMs) play an important role in urban analysis, 3D building reconstruction, and infrastructure monitoring, yet their availability remains limited due to the high cost and complexity of data acquisition. In contrast, coarse DSMs from commercial satellite missions are widely accessible, and high-resolution optical imagery is increasingly available from aerial and satellite platforms. We address the resulting mismatch in spatial resolution and propose a DSM superresolution approach that enhances 5 m DSMs to 0.5 m resolution, using guidance from high-resolution spectral images. Our method employs denoising diffusion to transfer information that is visible only in the image, like crisp outlines and detailed roof structures, into the elevation maps. In this way, surface details are reconstructed more accurately than with conventional interpolation or filtering techniques. Experiments on several cities in Central Europe demonstrate that the proposed approach produces high-quality DSMs with improved structural detail and accurate surface geometry. Our results highlight the potential of guided super-resolution with foundational image priors as a means of reconstructing high-resolution surface models.
Geospatial models reveal health insights missing in social risk data
Geospatial Foundation Models Capture Health-Relevant Dimensions of Place Beyond Conventional Social Risk Indices
Abstract: Area-based social risk indices summarize residents' socioeconomic conditions but incompletely capture physical features of place that may affect health. We evaluated whether numerical representations of physical place produced by four geospatial foundation model families from 2022 satellite data explained residual variance in tract-level associations between the Area Deprivation Index, Social Deprivation Index, and Social Vulnerability Index with health outcomes. We used LightGBM to predict variables from the American Community Survey and 40 chronic disease and health-behavior outcomes from CDC PLACES across 82,646 census tracts in the contiguous United States, evaluating performance across 10 held-out states. Among survey variables, models were moderately predictive of some variables including housing type (R-squared up to 0.54) but weak for disability, unemployment, and income disparity. For health outcomes, models explained up to 54% of variance left unexplained by social risk indices, with the largest gains for annual checkups, arthritis, and high blood pressure. Mean total variance explained by geospatial foundation models across the 40 health-related outcomes increased from 0.31 in the smallest tract-size decile to 0.39 in the largest. Geospatial foundation models capture health-relevant features of place not represented by conventional social risk indices and may usefully augment them in epidemiological analyses.
Visual motion on screens changes pedestrian walking paths in public
Visual-Motion-Induced Modulation of Pedestrian Trajectories Using Spatially Distributed Multi-Display Signage in Public Spaces
Abstract: Multi-display signage (MDS), now ubiquitous in urban environments, has the potential to influence human behavior and experience in public spaces. However, despite its unique capability to present spatially distributed dynamic visual stimuli, its current use is mainly limited to advertising. In this study, we propose a perception-based approach for laterally modulating pedestrian trajectories as a nonverbal means of guiding pedestrians in public spaces. The approach is motivated by vection, the illusion of self-motion, and uses laterally moving monochrome stripes, a standard stimulus in vection research, presented across spatially distributed displays to elicit postural responses that may bias pedestrian trajectories. We evaluated the approach through a controlled laboratory experiment and a real-world field deployment involving actual pedestrian flows in a national museum. The laboratory experiment examined whether the MDS setup induced trajectory shifts in the direction predicted by prior research on the behavioral effects of vection. The field deployment investigated whether comparable effects would emerge in aggregate pedestrian behavior during unconstrained movement under conditions closer to those of urban public spaces. In the laboratory, full-screen motion significantly biased walking trajectories in the direction of visual motion, whereas partial-stripe motion produced no significant directional effect. In the field deployment, opposing full-screen motion conditions produced direction-consistent differences in aggregate pedestrian positions. The field results, observed despite the substantial variability in real-world pedestrian flows, extend the controlled laboratory findings and provide ecologically valid evidence supporting practical MDS-based pedestrian modulation in public settings. The results further suggest that sufficient visual-motion coverage may be important.
Conceptual framework guides digital tech for social wellbeing in cities
Designing Technology for Social Wellbeing in Built Environments: A Conceptual Framework
Abstract: Digital technologies are deployed in urban built environments with the aim of supporting social dimensions. However, the research offers no unified guidance for such digital technology design. Additionally, evidence shows that week social wellbeing contributes to mental and physical health outcomes which suggests that the technology designed to strengthen social wellbeing could also function as a form of health promoting and preventive intervention. This chapter addresses this research gap by developing a conceptual framework that supports the design of digital technologies for social wellbeing in built environments. We propose a conceptual framework composed of three core components which are drawn on the synthesis of selected empirical studies on technologies embedded in built environments for social wellbeing. First, a social wellbeing dimensions model that identifies what digital technology could address. Second, a digital technology contribution matrix that distinguishes the types of contributions a digital technology could make. Third, levels that maps the scope at which technology could support social wellbeing. This conceptual framework could help researchers, practitioners and policymakers to design and guide digital technology interventions that target social wellbeing in the built environment.
StreetDiff improves urban street scene generation with consistent views
StreetDiff: Multi-view Street Scenes Generation via Cross-view Consistent Multi-view Stable Diffusion with Structure Prompts
Abstract: Multi-view diffusion models have shown strong performance in scenes with strong geometric priors and sparse semantics, such as indoor rooms or simple outdoor environments (e.g., fields, courtyards). However, they often fail to maintain cross-view consistency under camera rotation, especially in structurally complex urban environments. Without explicit modeling of spherical correspondence across views, existing approaches tend to produce object duplication, structural distortion, and layout inconsistency. To address this limitation, we propose StreetDiff, a multi-view diffusion framework that explicitly enforces cross-view alignment during denoising. StreetDiff introduces a Panorama--Perspective Synergy design to decouple global layout reasoning from local detail synthesis, and incorporates a Panorama Alignment Module (PAM) that establishes spherical-projection-based attention constraints across views. By injecting structured alignment constraints without modifying the diffusion backbone, our framework achieves robust cross-view coherence in challenging urban street scene generation tasks. In addition, we construct Street360, a large-scale HDR multi-view urban panorama dataset. Extensive experiments demonstrate that StreetDiff significantly improves structural consistency and visual fidelity compared to prior multi-view diffusion generation methods.
Sociodemographic data adds no clear value to predicting next locations
Who You Are Adds Nothing Detectable to Where You Go Next: Sociodemographic Conditioning in LLM Next-Location Prediction
Abstract: Large language models (LLMs) are increasingly used for individual next-location prediction, while sociodemographic conditioning is common in LLM-based travel simulation. Yet the incremental predictive value of sociodemographic attributes remains unclear. To directly test this contribution, sociodemographic records were linked with passively sensed mobility data from 5,000 Shenzhen residents to construct a closed-set benchmark in which models rank 100 candidate destinations. Each prediction instance is evaluated with and without age, gender, occupation and income, while holding mobility history, candidates and all other prompt content fixed. Results show that across four history lengths, the paired change in top-1 accuracy ranges from -0.8 to +0.5 percentage points, with no detectable gain from attributes. This result remains consistent when stay history is withheld, across alternative prediction times, in two additional LLMs and in a supervised reranker trained on the same benchmark. The null does not reflect a lack of model responsiveness to demographic information, as permuted attributes reduce LLM accuracy whereas correctly matched attributes do not improve it. A further asymmetry emerges in the reverse predictive direction, as pre-cut mobility trajectories recover income with an AUC of 0.708, while sociodemographic attributes contribute little to next-location prediction. Beyond demographic conditioning, candidate construction exerts a much larger influence on reported performance. Removing distance raises top-1 accuracy by 7.7 percentage points under proximity sampling but lowers it by 22.3 points under popularity sampling, with the reversal reproduced across all three LLMs. These results distinguish demographic association from incremental predictive usefulness and show that sampled next-location accuracy depends strongly on how candidate alternatives are constructed.
CityPlanner improves urban planning with interactive sandbox agent
CityPlanner: A Sandbox Agent for Executable Urban Planning
Abstract: Urban planning is a real-world spatial optimization problem that requires selecting feasible actions from large candidate spaces under practical objectives such as cost and service quality. Existing optimization and reinforcement learning methods are effective for fixed formulations, but often depend on task-specific representations and constraint handling. We propose \emph{CityPlanner}, a sandbox-agent framework for executable urban planning. CityPlanner introduces \emph{UrbanSandbox}, a unified file-based environment where agents inspect task files, generate plans, run evaluators, and revise decisions based on executable feedback. To make learning tractable, we further propose atomic-task reinforcement learning, which decomposes long sandbox trajectories into \emph{BuildPlan} for initial construction and \emph{ImprovePlan} for feedback-based refinement. Experiments on a real-world benchmark show that CityPlanner consistently outperforms heuristic, task-specific RL, and general LLM-agent baselines. Ablations verify the contributions of UrbanSandbox, atomic-task RL, and iterative deployment. We release the code and dataset at https://anonymous.4open.science/r/co-agent-C1C8
Urban region embeddings improve by combining mobility flow and connectivity patterns
Synergistic Fusion of Topological Structure and Temporal Semantics of Mobility for Urban Region Embedding
Abstract: Urban region embeddings have shown promising results in diverse urban sensing tasks such as crime, income, and service-call prediction. Recent methods improve representation quality by integrating mobility data with auxiliary modalities, using cross-view attention or contrastive objectives to align heterogeneous features into a unified region representation. However, leveraging the temporal dynamics of human mobility remains under-explored. Regional inflow and outflow fluctuate throughout the day, and inter-region connections emerge, persist, and dissolve over time. Moreover, prevailing fusion strategies combine views additively and miss the joint signal that emerges only when views co-occur. To address these gaps, we propose Mobility Stream-Structure Synergy (MoSS), which derives complementary views from mobility data: a Sequence view that preserves each region's hourly inflow/outflow profile, and a Structure view based on zigzag persistence diagrams that capture how regional connectivity emerges, persists, and dissolves over time. A synergy module then extracts emergent representations from the co-occurrence of these views through multi-degree interactions, explicitly capturing higher-order signal across views. Extensive experiments on New York City and Chicago show that MoSS achieves state-of-the-art performance across three downstream tasks using mobility data alone, outperforming baselines that rely on auxiliary modalities.
CrowdTraj benchmark reveals challenges in dense crowd trajectory prediction
CrowdTraj: A Benchmark for Dense Crowd Trajectory Prediction in Realistic Crowded Environments
Abstract: In real-world applications, pedestrian trajectory prediction models rely on inputs from detection and tracking systems. Prior trajectory prediction benchmarks either contain relatively sparse pedestrian interactions, assume perfect tracking inputs, or rely on overhead viewpoints that minimize occlusion and perspective distortion, limiting evaluation in realistic dense-crowd scenarios. We present CrowdTraj, a benchmark for pedestrian trajectory prediction in natural dense crowd scenes. Unlike previous datasets, CrowdTraj supports end-to-end evaluation from detection through tracking to trajectory prediction under severe occlusion in CCTV views. It also captures diverse, natural pedestrian behaviours, including abrupt directional changes rarely observed in existing benchmarks. CrowdTraj includes five diverse scenes, with an average of 1,146 unique pedestrians per scene, maximum frame-level densities ranging from 114 to 372 pedestrians, and over 3.2 million annotated head bounding boxes. CrowdTraj provides pixel and real-world coordinates via per-scene homography matrices for physically meaningful analysis. Our experimental results show that tracking accuracy (IDF1) drops to 0.68 to 0.70 in the densest scenes, compared with approximately 0.90 in less crowded scenes. Trajectory prediction training also becomes substantially more computationally expensive in dense scenes, with training times increasing by up to 8 times. These findings show that CrowdTraj exposes limitations in current trajectory prediction pipelines that remain hidden on existing sparse-crowd benchmarks, particularly in robustness to tracking noise and computational scalability.
Cities map where people travel using only crowd counts
Inferring Urban Mobility Interactions from Aggregated Dynamics
Abstract: Real-time urban governance depends not only on knowing where people are, but on how they move between places, directional flows that could be conventionally resolved by tracking individuals through space, i.e., expensive to sustain and built on traces that are highly unique and readily re-identifiable. Here we show that this directional structure need not be observed to be known: aggregated counts which cities already collect retain enough information to reconstruct the temporal evolution of origin-destination (OD) matrix. Using an uncertainty-aware physics-informed framework, we infer future OD flows from area-level counts alone across twelve mobility datasets from cities in the United States and China, reaching accuracy comparable to models that take historical OD matrices as input. Probabilistic modeling corrects the systematic underestimation of sparse, high-value corridors and yields calibrated predictions consistent with observed flows. Architectures that respect the generation-before-assignment logic of transport planning recover interactions more faithfully, indicating that location-level spatial heterogeneity should be preserved before pairwise interactions are reconstructed. Because inference requires only aggregated observations after training, recovering interactions this way reduces reliance on continuous individual-level tracking, pointing toward a more deployable and less exposure-heavy basis for real-time urban intelligence.
Improved strategies for fair and random obnoxious facility placement
Improved Randomized Approximations for Strategic Obnoxious Facility Location
Abstract: We study randomized strategyproof mechanisms for strategic obnoxious facility location on a line segment, where agents wish the facility to be located as far away from them as possible and their utility is their distance from the facility, under the social utility and minimum utility objectives. For social utility, we propose a novel randomized mechanism that breaks the previously best known \(\frac32\)-approximation of [Cheng, Yu, and Zhang, TCS 2013], achieving an approximation ratio of at most \(1.47359\). We also raise the lower bound on the approximation ratio of randomized strategyproof mechanisms from \(\frac{2}{\sqrt{3}}\approx1.15470\) [Feigenbaum et al., JAAMAS 2020] to \(\frac{105}{88}\approx1.19318\). For minimum utility, following the profile-independent approach of [Chan, Lin and Wang, AAMAS 2026], we design a simple randomized mechanism that reduces the approximation guarantee from \(\sqrt{2n}+O(1)\) to \(\sqrt n+O(1)\), where \(n\) is the number of agents. Finally, we prove that no randomized strategyproof mechanism can achieve an asymptotic approximation ratio strictly smaller than \(2\), strengthening the previous asymptotic lower bound of \(\frac32\) [Feigenbaum et al., JAAMAS 2020]. Thus, all four bounds considered in this paper strictly improve upon the corresponding previously known results.