Cross-View Urban Sensing: Mapping Subjective Streetscape Perception via AlphaEarth Embeddings and Urban Context

2026-08-17Computer Vision and Pattern Recognition

Computer Vision and Pattern Recognition
AI summary

The authors developed CVLNet, a model that predicts how people feel about city streets without needing street-level photos each time. Instead, it uses satellite data and other city info to estimate street qualities in four Southeast Asian cities. Their model predicts five perception aspects better than previous methods and covers entire cities, not just where photos exist. They also studied how these street qualities vary across different populations and areas, highlighting inequalities. This approach shows a way to map urban environments more fully using remote sensing data.

Street View ImageryPerceptual qualitiesAlphaEarth embeddingsCross-View LearningUrban contextual dataAdjusted R-squaredRemote sensingPopulation exposure inequalityDeficit Palma RatioSoutheast Asian cities
Authors
Peilin Li, Pengfei Chen, Jingyu Wang, Zhifeng Yang, Tiansheng Chen, Mengjie Gong, Xiao Cheng
Abstract
Residents' perception of the urban streetscape is an important factor in public health, active mobility, and social wellbeing. Street view imagery (SVI) has emerged as a widely used data source for assessing these perceptual qualities, yet its uneven coverage and irregular updating limit large-scale measurement. Here, we present CVLNet, a Cross-View Learning Network that predicts street-level perception from AlphaEarth embeddings and multi-source urban contextual data without requiring SVI at inference. CVLNet applies per-task adaptive gating to jointly model five perceptual dimensions, using labels from the pretrained SVI-Percept model as ground truth. The proposed method is evaluated across four Southeast Asian cities: Singapore, Kuala Lumpur, Jakarta, and Manila. CVLNet achieves a median road-segment-level Adjusted $R^{2}$ of 0.76 and consistently outperforms the baseline models, with gains ranging from 5.9--11.3% across the five perceptual dimensions. Ablation experiments show that AlphaEarth features and urban contextual features contribute complementary information. We further produce citywide road-level streetscape perception maps for five subjective perceptual dimensions across all four cities, extending perception estimation from the 13--31% of the road network directly covered by available SVI to the complete road network of each city. Integrating these maps with WorldPop gridded population data, we quantify exposure inequality across population-density, demographic, and land-use groups using the Deficit Palma Ratio. These results demonstrate that remote sensing can serve as a scalable alternative to SVI for citywide streetscape perception mapping, enabling a more comprehensive assessment of urban environmental inequality.