Papers for

urban planners

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Robots as social elements shaping human group experiences

A Robot Among People:From Social Imitation to the Social Becoming of Human Groups

Abstract: Robots designed to mediate human groups often fall into the solutionist trap: they are framed as sociable agents that fix problems such as conflict, disengagement, or lack of coordination. We suggest a different way of thinking. Rather than discrete agents, robots can be understood as situated elements of shared environments; catalysts and carriers of group experience whose meaning emerges through how people position, interpret, and interact with them. From this perspective, robots are not there to repair some ostensible dysfunctionality, but to enable group-level sense-making around care, norms, and identity. Our prior work on robotic street furniture suggests that this does not happen by imitating human sociality but by taking the shape of deliberately constrained, group-facing entities that happen and act for \textit{us} without being socially entangled as one of us. We thus understand robots in public spaces not in terms of autonomy or intelligence, but as a relational capacity. This implies designing robots not in our image or for our utility, but grounded in our needs in being and becoming together.

Fri 11 SeptRoboticsComputers and SocietyHuman-Computer Interaction
The gist
Sometimes robots are designed to fix problems between people, like stopping fights or helping people work together. This paper suggests a different idea: robots don’t need to copy humans exactly, but instead can act as parts of the environment that help groups understand each other better. The authors studied robots in public places and found that robots work best when they act for the group as a whole, not as another human. This helps groups create shared meanings about care, rules, and identity without the robot being a person.
Open 2609.12937v1

Strategyproof facility location with learned predictions improves accuracy

Learning-Augmented Strategyproof Facility Location in $\mathbb{R}^d$ with $\ell_p$ Distances

Abstract: We study learning-augmented mechanism design for locating a single facility in $\mathbb{R}^d$ to minimize the sum of the agents' $\ell_p$ distances to the facility. We analyze the coordinate-wise median with predictions (CMP) mechanism, which adds $cn$ virtual agents at a predicted optimal facility location to the $n$ reported locations and returns their coordinate-wise median. The parameter $c\in[0,1)$ represents confidence in the prediction; CMP is known to be strategyproof when both are fixed independently of the reports. We determine its approximation guarantees under correct predictions (consistency) and arbitrary predictions (robustness) in two settings. First, for $d=2$, we establish exact guarantees for every $p\in[1,+\infty]$. For $1<p<+\infty$, the consistency and robustness are respectively $\Bigl(1+\bigl(\frac{1-c}{1+c}\bigr)^{\frac{p}{p-1}}\Bigr)^{\frac{p-1}{p}}$ and $\Bigl(1+\bigl(\frac{1+c}{1-c}\bigr)^{\frac{p}{p-1}}\Bigr)^{\frac{p-1}{p}}$. We use the median condition in each coordinate to compare the mechanism's cost with the optimal cost, and establish tightness using instances with agents at only three distinct locations. Second, for $1<p<+\infty$, we obtain dimension-independent consistency and robustness upper bounds valid for every $d\ge1$. For each fixed $p$ and $c$, we construct families of instances whose ratios approach the respective upper bounds as $d\to\infty$, proving asymptotic tightness. We also establish exact guarantees for $p=1$ in every dimension and asymptotically tight bounds for $p=+\infty$. Our high-dimensional results recover the prediction-free bounds of Gravin and Jia (STOC 2025) when $c=0$ and their learning-augmented bounds in arbitrary-dimensional Euclidean spaces when $p=2$.

Fri 11 SeptComputer Science and Game Theory
The gist
Finding the best location for something like a public facility to minimize travel distances can be tricky, especially when people might lie about their locations. This paper studies a method that uses machine learning predictions combined with a strategy that prevents people from benefiting by lying. The authors analyze how well this method works when predictions are good and when they are wrong, in spaces of different dimensions and for various distance measures. They provide exact and asymptotic guarantees on how close their approach gets to the best possible location.
Open 2609.12829v1

Strategyproof facility location on a circle with improved approximation ratio

A Randomized $\frac32$-Approximation for Strategic Facility Location on a Circle

Abstract: We study strategyproof mechanisms for locating a single facility on a circle so as to serve a set of strategic agents under the utilitarian social cost objective. We analyze a simple parity-dependent mechanism that randomizes between the Random Dictator (RD) mechanism and the Proportional Circle Distance (PCD) mechanism. For an odd number \(n\) of agents, this mechanism mixes RD and PCD with equal probability, as originally proposed by Rogowski and Dziubi{ń}ski (IJCAI 2025). For even \(n\), it mixes RD with a random-deletion extension of PCD, in which one agent is removed uniformly at random before PCD is applied to the remaining agents. Our main result shows that this mechanism is strategyproof and achieves an approximation ratio of \(\frac32\) for every \(n\ge 3\). This improves the \(\frac74\) upper bound of Rogowski and Dziubi{ń}ski, which applied only to odd \(n\), and extends the guarantee to even numbers of agents. Finally, we establish a lower bound of \(\frac{11}{10}\) on the approximation ratio of any randomized strategyproof mechanism, improving on the previous lower bound of \(1.0456\) due to Meir (SAGT 2019).

Fri 11 SeptComputer Science and Game Theory
The gist
This paper looks at how to fairly choose a spot on a circular path to build one facility so that a group of people, each with their own honest preferences, are served well. The authors study a method that mixes two known ways to pick this spot, ensuring that no one can benefit by lying about their location. They prove this mixed method works better than previous approaches, giving a closer-to-optimal solution for any number of people. They also show that no random method can do much better than their guaranteed performance.
Open 2609.12792v1

Machine learning predicts pedestrian counts from city map features

Estimating Pedestrian Volumes from GIS-Derived Built-Environment Features: A Machine Learning Framework

Abstract: Transportation agencies need pedestrian volume estimates across entire road networks to prioritize safety investments, yet manual counts are expensive and cover only a small share of intersections. We present a machine learning pipeline that predicts 2-hour PM peak pedestrian volume at 101 urban intersections in Portland, Oregon, from built-environment, land-use, and street-network features drawn from open GIS data. Starting from the Negative Binomial GLM used in practice, we add feature selection, count-aware gradient boosting, and repeated cross-validation, selecting one configuration by a combined rank over RMSE, MAPE, and SMAPE across four cross-validation strategies. The winner, a histogram-based gradient boosting model with Poisson loss and L1 Lasso feature selection, reduces cross-validated RMSE by 12% over the GLM baseline (89.8 to 78.7) and holdout RMSE by 19% (108.0 to 87.9). Code is released on GitHub.

Thu 10 SeptMachine Learning
The gist
Transportation planners need to know how many people walk at different intersections to make roads safer, but counting pedestrians manually is hard and expensive. The authors created a computer model that uses information from maps about streets and buildings to predict how many people walk at certain times in Portland. Their model does better than the standard method, cutting down prediction errors by around 12 to 19 percent. This helps make better decisions about where to focus on pedestrian safety without needing lots of manual counting.
Open 2609.12173v1

Improved road scene segmentation with RGB and near infrared images

HSI-Road Relabeled: Surface-Aware Road-Scene Segmentation

Abstract: The HSI-Road dataset provides paired RGB and 25-channel NIR (600--960~nm) images with binary masks but no surface-level labels.~This paper introduces a manually labeled six-class taxonomy: Background, Asphalt, Concrete, Dirt, Water, and Grass, and an RGB-to-NIR registration pipeline with corresponding annotations. Six semantic-segmentation models (SSMs) are evaluated under four input configurations: original-resolution RGB (RGB$_{\text{ori}}$), registered low-resolution RGB (RGB$_{\text{reg}}$), NIR, and channel-stacked RGB$_{\text{reg}}$--NIR (RGBN$_{\text{stk}}$). The comparison quantifies the effect of spatial-resolution reduction on RGB, along with evaluation of NIR and RGBN$_{\text{stk}}$, with results reported using per-class and mean IoU and F1 scores. RGB$_{\text{ori}}$ achieves the highest overall performance but contains 12$\times$ more pixels than the matched-resolution inputs. At the matched 192$\times$384 resolution, RGBN$_{\text{stk}}$ outperforms NIR for all six SSMs and RGB$_{\text{reg}}$ for five of six, with the most consistent gains for the Water class. These results highlight the importance of spatial resolution while showing that NIR provides complementary information to RGB.

Thu 10 SeptComputer Vision and Pattern Recognition
The gist
It can be hard for computer programs to tell different road surfaces apart, like asphalt, concrete, or water. This paper adds new detailed labels to a dataset that has regular camera images and special near-infrared images. The authors tested models using different types of input images and found that using both registered RGB and near-infrared images together helps recognize surfaces better at similar image resolutions. They also show that image resolution is important for accurate results.
Open 2609.12151v1

Perceived food access reveals barriers beyond store proximity

Population-level measures of perceived food access reveal barriers beyond geographic proximity

Abstract: Food access is multidimensional, but population-level measurement still relies heavily on geography because perceived dimensions of access are difficult to measure at scale. Here, we use 25,125 Google Maps reviews from 49 grocery stores in Raleigh, North Carolina, to measure five dimensions of food access: availability, accessibility, affordability, accommodation, and acceptability. We identify review topics with unsupervised topic modeling and assign them to access dimensions using zero-shot classification, with 85.4% agreement against manual coding. The resulting store-level measures capture distinct aspects of food access and reveal barriers that geographic proximity alone does not capture. Comparisons between nearby stores in the same chain further show that identical store policies can be perceived very differently across locations, consistent with food access reflecting the fit between residents and their food environment. Perceived food access also follows systematic socioeconomic and demographic patterns that broadly parallel, but do not replicate, those observed for geographic access. These results show that online grocery reviews can provide a scalable complement to geographic measures of food access.

Thu 10 SeptComputation and Language
The gist
Measuring how easy it is for people to get food usually looks at how close stores are, but that doesn’t tell the whole story. The authors used thousands of online reviews to understand five different ways people think about food access, like whether food is affordable or if the store feels welcoming. They found that even stores close to each other can feel very different to shoppers depending on the neighborhood and store experience. This means that looking at just the location of stores misses important parts of how people experience getting food.
Open 2609.12132v1

Digital twins must include human and social system factors

Missing Dimensions: Integrating Human and Social Systems into Digital Twin Engineering

Abstract: Digital twins (DTs) have emerged as a key technology at the core of digital transformation, yet their engineering practice remains too narrowly focused on engineered and natural systems. This paper argues that four system dimensions must be explicitly recognized in DT engineering: Engineered, Natural/Biological, Human, and Social. Each dimension brings distinct properties, modeling requirements, and ethical obligations that fundamentally shape what a DT must represent and how it must be built. We further argue that as DTs extend into the Human and Social dimensions, the emphasis should shift from automated control toward decision support and occupant empowerment. We illustrate our arguments using a smart building as a running example and identify four open research challenges for multi-dimensional DT engineering.

Thu 10 SeptSoftware Engineering
The gist
Digital twins are computer models that replicate real-world systems for monitoring and control. But most digital twins only focus on physical or natural elements, leaving out humans and their social interactions. The authors say we should build digital twins that also model human behavior and social dynamics, which means shifting from just automatic control to helping people make decisions. They use smart buildings as an example and highlight challenges in creating these multi-dimensional digital twins.
Open 2609.12131v1

Super-resolution improves digital elevation models using satellite images

Guided Super-Resolution of Digital Elevation Models with Diffusion-Based Image Generators

Abstract: High-resolution digital surface models (DSMs) play an important role in urban analysis, 3D building reconstruction, and infrastructure monitoring, yet their availability remains limited due to the high cost and complexity of data acquisition. In contrast, coarse DSMs from commercial satellite missions are widely accessible, and high-resolution optical imagery is increasingly available from aerial and satellite platforms. We address the resulting mismatch in spatial resolution and propose a DSM superresolution approach that enhances 5 m DSMs to 0.5 m resolution, using guidance from high-resolution spectral images. Our method employs denoising diffusion to transfer information that is visible only in the image, like crisp outlines and detailed roof structures, into the elevation maps. In this way, surface details are reconstructed more accurately than with conventional interpolation or filtering techniques. Experiments on several cities in Central Europe demonstrate that the proposed approach produces high-quality DSMs with improved structural detail and accurate surface geometry. Our results highlight the potential of guided super-resolution with foundational image priors as a means of reconstructing high-resolution surface models.

Thu 10 SeptComputer Vision and Pattern Recognition
The gist
High-resolution 3D maps of surfaces, like cities, are very useful but hard and expensive to make. The authors found a way to improve low-resolution elevation maps by using clear satellite photos to add details. They use a special AI method called denoising diffusion to bring sharp shapes and features from the images into the elevation data. This method works better than older techniques and can create better 3D models of cities.
Open 2609.11886v1

Geospatial models reveal health insights missing in social risk data

Geospatial Foundation Models Capture Health-Relevant Dimensions of Place Beyond Conventional Social Risk Indices

Abstract: Area-based social risk indices summarize residents' socioeconomic conditions but incompletely capture physical features of place that may affect health. We evaluated whether numerical representations of physical place produced by four geospatial foundation model families from 2022 satellite data explained residual variance in tract-level associations between the Area Deprivation Index, Social Deprivation Index, and Social Vulnerability Index with health outcomes. We used LightGBM to predict variables from the American Community Survey and 40 chronic disease and health-behavior outcomes from CDC PLACES across 82,646 census tracts in the contiguous United States, evaluating performance across 10 held-out states. Among survey variables, models were moderately predictive of some variables including housing type (R-squared up to 0.54) but weak for disability, unemployment, and income disparity. For health outcomes, models explained up to 54% of variance left unexplained by social risk indices, with the largest gains for annual checkups, arthritis, and high blood pressure. Mean total variance explained by geospatial foundation models across the 40 health-related outcomes increased from 0.31 in the smallest tract-size decile to 0.39 in the largest. Geospatial foundation models capture health-relevant features of place not represented by conventional social risk indices and may usefully augment them in epidemiological analyses.

Thu 10 SeptMachine Learning
The gist
Social risk indexes help show how neighborhood factors might affect health, but they miss physical details about the place itself. The authors found that models analyzing satellite images can capture these physical features and explain health differences better. These models provided new information beyond usual social risk measures, particularly about health issues like arthritis and high blood pressure. This approach can improve health studies by including more detailed data on the environment where people live.
Open 2609.11689v1

Visual motion on screens changes pedestrian walking paths in public

Visual-Motion-Induced Modulation of Pedestrian Trajectories Using Spatially Distributed Multi-Display Signage in Public Spaces

Abstract: Multi-display signage (MDS), now ubiquitous in urban environments, has the potential to influence human behavior and experience in public spaces. However, despite its unique capability to present spatially distributed dynamic visual stimuli, its current use is mainly limited to advertising. In this study, we propose a perception-based approach for laterally modulating pedestrian trajectories as a nonverbal means of guiding pedestrians in public spaces. The approach is motivated by vection, the illusion of self-motion, and uses laterally moving monochrome stripes, a standard stimulus in vection research, presented across spatially distributed displays to elicit postural responses that may bias pedestrian trajectories. We evaluated the approach through a controlled laboratory experiment and a real-world field deployment involving actual pedestrian flows in a national museum. The laboratory experiment examined whether the MDS setup induced trajectory shifts in the direction predicted by prior research on the behavioral effects of vection. The field deployment investigated whether comparable effects would emerge in aggregate pedestrian behavior during unconstrained movement under conditions closer to those of urban public spaces. In the laboratory, full-screen motion significantly biased walking trajectories in the direction of visual motion, whereas partial-stripe motion produced no significant directional effect. In the field deployment, opposing full-screen motion conditions produced direction-consistent differences in aggregate pedestrian positions. The field results, observed despite the substantial variability in real-world pedestrian flows, extend the controlled laboratory findings and provide ecologically valid evidence supporting practical MDS-based pedestrian modulation in public settings. The results further suggest that sufficient visual-motion coverage may be important.

Thu 10 SeptHuman-Computer Interaction
The gist
Big screens showing moving stripes can trick people’s brains into feeling like they are moving sideways, even when they are not. The authors tested this idea in a lab and at a national museum and found that when these patterns cover a large part of people’s field of view, they tend to walk slightly in the direction the stripes move. This shows that such visual illusions on multiple displays can subtly guide where crowds walk without saying a word. It could help designers manage foot traffic in busy places.
Open 2609.11088v1

Conceptual framework guides digital tech for social wellbeing in cities

Designing Technology for Social Wellbeing in Built Environments: A Conceptual Framework

Abstract: Digital technologies are deployed in urban built environments with the aim of supporting social dimensions. However, the research offers no unified guidance for such digital technology design. Additionally, evidence shows that week social wellbeing contributes to mental and physical health outcomes which suggests that the technology designed to strengthen social wellbeing could also function as a form of health promoting and preventive intervention. This chapter addresses this research gap by developing a conceptual framework that supports the design of digital technologies for social wellbeing in built environments. We propose a conceptual framework composed of three core components which are drawn on the synthesis of selected empirical studies on technologies embedded in built environments for social wellbeing. First, a social wellbeing dimensions model that identifies what digital technology could address. Second, a digital technology contribution matrix that distinguishes the types of contributions a digital technology could make. Third, levels that maps the scope at which technology could support social wellbeing. This conceptual framework could help researchers, practitioners and policymakers to design and guide digital technology interventions that target social wellbeing in the built environment.

Wed 9 SeptComputers and SocietyHuman-Computer Interaction
The gist
Many digital technologies are used in cities to help people connect and support each other, but there isn’t clear guidance on how to design these tools effectively. The authors created a framework to help design digital technologies that improve social wellbeing in buildings and urban spaces. Their framework includes understanding what social wellbeing means, how technology can help, and the different levels at which technology can support communities. This work can help planners, designers, and policymakers create better technology for healthier social environments.
Open 2609.10779v1

StreetDiff improves urban street scene generation with consistent views

StreetDiff: Multi-view Street Scenes Generation via Cross-view Consistent Multi-view Stable Diffusion with Structure Prompts

Abstract: Multi-view diffusion models have shown strong performance in scenes with strong geometric priors and sparse semantics, such as indoor rooms or simple outdoor environments (e.g., fields, courtyards). However, they often fail to maintain cross-view consistency under camera rotation, especially in structurally complex urban environments. Without explicit modeling of spherical correspondence across views, existing approaches tend to produce object duplication, structural distortion, and layout inconsistency. To address this limitation, we propose StreetDiff, a multi-view diffusion framework that explicitly enforces cross-view alignment during denoising. StreetDiff introduces a Panorama--Perspective Synergy design to decouple global layout reasoning from local detail synthesis, and incorporates a Panorama Alignment Module (PAM) that establishes spherical-projection-based attention constraints across views. By injecting structured alignment constraints without modifying the diffusion backbone, our framework achieves robust cross-view coherence in challenging urban street scene generation tasks. In addition, we construct Street360, a large-scale HDR multi-view urban panorama dataset. Extensive experiments demonstrate that StreetDiff significantly improves structural consistency and visual fidelity compared to prior multi-view diffusion generation methods.

Wed 9 SeptComputer Vision and Pattern Recognition
The gist
Generating realistic images of city streets from different camera angles is hard because objects can appear duplicated or distorted when views change. The authors created StreetDiff, a new method that keeps these different views aligned by using a special way to connect panoramic and perspective images. They also built a large dataset of urban panoramic images, called Street360, to help train and test their method. This approach produces more coherent and detailed street scene images than previous techniques.
Open 2609.09890v1

Sociodemographic data adds no clear value to predicting next locations

Who You Are Adds Nothing Detectable to Where You Go Next: Sociodemographic Conditioning in LLM Next-Location Prediction

Abstract: Large language models (LLMs) are increasingly used for individual next-location prediction, while sociodemographic conditioning is common in LLM-based travel simulation. Yet the incremental predictive value of sociodemographic attributes remains unclear. To directly test this contribution, sociodemographic records were linked with passively sensed mobility data from 5,000 Shenzhen residents to construct a closed-set benchmark in which models rank 100 candidate destinations. Each prediction instance is evaluated with and without age, gender, occupation and income, while holding mobility history, candidates and all other prompt content fixed. Results show that across four history lengths, the paired change in top-1 accuracy ranges from -0.8 to +0.5 percentage points, with no detectable gain from attributes. This result remains consistent when stay history is withheld, across alternative prediction times, in two additional LLMs and in a supervised reranker trained on the same benchmark. The null does not reflect a lack of model responsiveness to demographic information, as permuted attributes reduce LLM accuracy whereas correctly matched attributes do not improve it. A further asymmetry emerges in the reverse predictive direction, as pre-cut mobility trajectories recover income with an AUC of 0.708, while sociodemographic attributes contribute little to next-location prediction. Beyond demographic conditioning, candidate construction exerts a much larger influence on reported performance. Removing distance raises top-1 accuracy by 7.7 percentage points under proximity sampling but lowers it by 22.3 points under popularity sampling, with the reversal reproduced across all three LLMs. These results distinguish demographic association from incremental predictive usefulness and show that sampled next-location accuracy depends strongly on how candidate alternatives are constructed.

Wed 9 SeptComputers and Society
The gist
Predicting where people will go next using large language models is common, but it's unclear if personal details like age or income help. The authors tested this by linking movement data with demographics for 5,000 residents and found that adding these personal details made almost no difference in prediction accuracy. Interestingly, while past locations can reveal income, knowing income doesn’t improve next-move predictions. They also found that how possible next places are chosen to compare affects prediction results much more than demographics do.
Open 2609.09609v1

CityPlanner improves urban planning with interactive sandbox agent

CityPlanner: A Sandbox Agent for Executable Urban Planning

Abstract: Urban planning is a real-world spatial optimization problem that requires selecting feasible actions from large candidate spaces under practical objectives such as cost and service quality. Existing optimization and reinforcement learning methods are effective for fixed formulations, but often depend on task-specific representations and constraint handling. We propose \emph{CityPlanner}, a sandbox-agent framework for executable urban planning. CityPlanner introduces \emph{UrbanSandbox}, a unified file-based environment where agents inspect task files, generate plans, run evaluators, and revise decisions based on executable feedback. To make learning tractable, we further propose atomic-task reinforcement learning, which decomposes long sandbox trajectories into \emph{BuildPlan} for initial construction and \emph{ImprovePlan} for feedback-based refinement. Experiments on a real-world benchmark show that CityPlanner consistently outperforms heuristic, task-specific RL, and general LLM-agent baselines. Ablations verify the contributions of UrbanSandbox, atomic-task RL, and iterative deployment. We release the code and dataset at https://anonymous.4open.science/r/co-agent-C1C8

Wed 9 SeptArtificial IntelligenceComputation and Language
The gist
Urban planning involves choosing the best ways to design cities by balancing costs and services, which is very complex. The authors created CityPlanner, a tool where an agent uses a sandbox environment to try out city plans, check how good they are, and then improve them step-by-step. To make this easier to learn, the system breaks the planning into two parts: building an initial plan and then refining it based on feedback. Tested on real city data, CityPlanner performs better than other existing methods for urban planning.
Open 2609.09578v1

Urban region embeddings improve by combining mobility flow and connectivity patterns

Synergistic Fusion of Topological Structure and Temporal Semantics of Mobility for Urban Region Embedding

Abstract: Urban region embeddings have shown promising results in diverse urban sensing tasks such as crime, income, and service-call prediction. Recent methods improve representation quality by integrating mobility data with auxiliary modalities, using cross-view attention or contrastive objectives to align heterogeneous features into a unified region representation. However, leveraging the temporal dynamics of human mobility remains under-explored. Regional inflow and outflow fluctuate throughout the day, and inter-region connections emerge, persist, and dissolve over time. Moreover, prevailing fusion strategies combine views additively and miss the joint signal that emerges only when views co-occur. To address these gaps, we propose Mobility Stream-Structure Synergy (MoSS), which derives complementary views from mobility data: a Sequence view that preserves each region's hourly inflow/outflow profile, and a Structure view based on zigzag persistence diagrams that capture how regional connectivity emerges, persists, and dissolves over time. A synergy module then extracts emergent representations from the co-occurrence of these views through multi-degree interactions, explicitly capturing higher-order signal across views. Extensive experiments on New York City and Chicago show that MoSS achieves state-of-the-art performance across three downstream tasks using mobility data alone, outperforming baselines that rely on auxiliary modalities.

Tue 8 SeptMachine LearningArtificial Intelligence
The gist
Understanding how people move around a city can help predict things like crime rates or income levels in different neighborhoods. The authors found that current methods often miss the time patterns of these movements and how areas connect differently throughout the day. They created a new way to capture both the flow of people in and out of areas and the changing connections between regions over time. Their approach combines these two views in a way that finds new insights and performs better than older methods, even without extra data types.
Open 2609.08268v1

CrowdTraj benchmark reveals challenges in dense crowd trajectory prediction

CrowdTraj: A Benchmark for Dense Crowd Trajectory Prediction in Realistic Crowded Environments

Abstract: In real-world applications, pedestrian trajectory prediction models rely on inputs from detection and tracking systems. Prior trajectory prediction benchmarks either contain relatively sparse pedestrian interactions, assume perfect tracking inputs, or rely on overhead viewpoints that minimize occlusion and perspective distortion, limiting evaluation in realistic dense-crowd scenarios. We present CrowdTraj, a benchmark for pedestrian trajectory prediction in natural dense crowd scenes. Unlike previous datasets, CrowdTraj supports end-to-end evaluation from detection through tracking to trajectory prediction under severe occlusion in CCTV views. It also captures diverse, natural pedestrian behaviours, including abrupt directional changes rarely observed in existing benchmarks. CrowdTraj includes five diverse scenes, with an average of 1,146 unique pedestrians per scene, maximum frame-level densities ranging from 114 to 372 pedestrians, and over 3.2 million annotated head bounding boxes. CrowdTraj provides pixel and real-world coordinates via per-scene homography matrices for physically meaningful analysis. Our experimental results show that tracking accuracy (IDF1) drops to 0.68 to 0.70 in the densest scenes, compared with approximately 0.90 in less crowded scenes. Trajectory prediction training also becomes substantially more computationally expensive in dense scenes, with training times increasing by up to 8 times. These findings show that CrowdTraj exposes limitations in current trajectory prediction pipelines that remain hidden on existing sparse-crowd benchmarks, particularly in robustness to tracking noise and computational scalability.

Mon 7 SeptComputer Vision and Pattern Recognition
The gist
CrowdTraj is a new dataset that helps test how well computers can predict where people will move in very crowded spaces, using real security camera footage. Earlier tests mostly looked at simpler cases with fewer people or perfect tracking, which doesn’t happen in real life. This dataset includes very busy scenes with lots of occlusions and sudden movements, making predictions harder. The authors found that current methods struggle more with this complexity and need more computing power to learn from such crowded examples.
Open 2609.07685v1

Cities map where people travel using only crowd counts

Inferring Urban Mobility Interactions from Aggregated Dynamics

Abstract: Real-time urban governance depends not only on knowing where people are, but on how they move between places, directional flows that could be conventionally resolved by tracking individuals through space, i.e., expensive to sustain and built on traces that are highly unique and readily re-identifiable. Here we show that this directional structure need not be observed to be known: aggregated counts which cities already collect retain enough information to reconstruct the temporal evolution of origin-destination (OD) matrix. Using an uncertainty-aware physics-informed framework, we infer future OD flows from area-level counts alone across twelve mobility datasets from cities in the United States and China, reaching accuracy comparable to models that take historical OD matrices as input. Probabilistic modeling corrects the systematic underestimation of sparse, high-value corridors and yields calibrated predictions consistent with observed flows. Architectures that respect the generation-before-assignment logic of transport planning recover interactions more faithfully, indicating that location-level spatial heterogeneity should be preserved before pairwise interactions are reconstructed. Because inference requires only aggregated observations after training, recovering interactions this way reduces reliance on continuous individual-level tracking, pointing toward a more deployable and less exposure-heavy basis for real-time urban intelligence.

Mon 7 SeptMachine Learning
The gist
Knowing how people move between different parts of a city helps manage urban life better, but tracking individuals is expensive and raises privacy concerns. The authors show that just counting how many people are in various areas over time is enough to guess these movements accurately. They developed a smart method that predicts where people will go next without needing detailed individual tracking. This method works well on real data from cities in the US and China and could help cities monitor movement in real-time while protecting privacy.
Open 2609.07349v1

Improved strategies for fair and random obnoxious facility placement

Improved Randomized Approximations for Strategic Obnoxious Facility Location

Abstract: We study randomized strategyproof mechanisms for strategic obnoxious facility location on a line segment, where agents wish the facility to be located as far away from them as possible and their utility is their distance from the facility, under the social utility and minimum utility objectives. For social utility, we propose a novel randomized mechanism that breaks the previously best known \(\frac32\)-approximation of [Cheng, Yu, and Zhang, TCS 2013], achieving an approximation ratio of at most \(1.47359\). We also raise the lower bound on the approximation ratio of randomized strategyproof mechanisms from \(\frac{2}{\sqrt{3}}\approx1.15470\) [Feigenbaum et al., JAAMAS 2020] to \(\frac{105}{88}\approx1.19318\). For minimum utility, following the profile-independent approach of [Chan, Lin and Wang, AAMAS 2026], we design a simple randomized mechanism that reduces the approximation guarantee from \(\sqrt{2n}+O(1)\) to \(\sqrt n+O(1)\), where \(n\) is the number of agents. Finally, we prove that no randomized strategyproof mechanism can achieve an asymptotic approximation ratio strictly smaller than \(2\), strengthening the previous asymptotic lower bound of \(\frac32\) [Feigenbaum et al., JAAMAS 2020]. Thus, all four bounds considered in this paper strictly improve upon the corresponding previously known results.

Mon 7 SeptComputer Science and Game Theory
The gist
This paper looks at how to fairly decide where to put a facility along a line when people want it as far away from them as possible. The authors propose better random methods that improve fairness and efficiency compared to older ones. They show their methods work closer to the best possible fairness and prove limits on how good any method can be. These results help understand the best ways to balance people’s competing interests in facility placement.
Open 2609.07261v1