Papers for

surveillance technology developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

ReVA dataset advances remote sensing video question answering worldwide

ReVA: A Scene-Centric Dataset Beyond Repetition for Remote Sensing Video Question Answering

Abstract: Multimodal Large Language Models (MLLMs) have demonstrated remarkable advances in remote sensing. However, existing remote sensing multimodal reasoning benchmarks exhibit two critical limitations: they rely on (i) template-driven questions, which causes repetitive questions; and (ii) static images that fail to capture the inherent temporal nature of drone/UAV videos. This leaves systematic evaluation of remote sensing video reasoning largely unexplored. To address this gap, we introduce ReVA, a new dataset for remote sensing video question answering, designed to assess spatiotemporal, scene-centric, and reasoning-oriented capabilities of MLLMs. ReVA comprises 2,438 drone videos spanning 18 cities worldwide (580K frames) and 22K high-quality question-answer pairs across 11 challenging QA tasks. We develop a semi-automatic annotation pipeline that leverages Text LLMs and MLLMs for question-answer generation with human verification. We evaluate 23 proprietary and open-source Video LLMs on ReVA, exposing fundamental limitations of current models. These findings position ReVA as a critical benchmark toward better remote sensing video understanding and temporal reasoning capabilities for real-world deployments. Our code and dataset are available at: https://github.com/zyaocoder/ReVA

Mon 28 SeptComputer Vision and Pattern Recognition
The gist
Understanding videos taken from drones over cities is hard because questions often repeat and images don’t show change over time. The authors created ReVA, a large collection of drone videos with lots of unique questions and answers about what’s happening in the scenes. They tested many video AI models on ReVA and found that current technology still struggles with understanding these complex videos. ReVA helps improve AI’s ability to reason about changes and scenes in remote sensing videos.
Open → 2609.35507v1

Origins method improves gait recognition with unified template learning

Learning A Unified Template for Gait Recognition

Abstract: "What I cannot create, I do not understand."Human wisdom reveals that creation is one of the highest forms of learning. For example, Diffusion Models have demonstrated remarkable semantic structure and memory in image generation, understanding, and restoration, which intuitively benefits representation learning. However, current gait networks rarely embrace this perspective, relying primarily on learning by contrasting gait samples under varying complex conditions, leading to semantic inconsistency and uniformity issues. To address these issues, we propose Origins with generative capabilities whose underlying philosophy is that different entities are generated from a unified template, inherently regularizing gait representations within a consistent and diverse semantic space to capture accurate gait differences. Admittedly, learning this unified template is exceedingly challenging, as it requires the comprehensiveness of the template to encompass gait representations with various conditions. Inspired by Diffusion Models, Origins diffuses the unified template into timestep templates for gait generative learning, and meanwhile transfers the unified template for gait representation learning. Especially, gait generative and representation learning serve as a unified framework for end-to-end joint training. Extensive experiments on CASIA-B, CCPG,SUSTech1K, Gait3D, GREW and CCGR-MINI demonstrate that Origins performs unified generative and representation learning, achieving superior performance.

Wed 16 SeptComputer Vision and Pattern Recognition
The gist
Gait recognition means identifying people by the way they walk, which is tricky because walking changes with different conditions. The authors created a new method called Origins that learns a single, unified template to represent all walking styles and conditions. By using ideas from image generation called Diffusion Models, Origins can both generate and recognize walking patterns more consistently. This approach helps computers tell people apart more accurately by focusing on key walking differences within a stable framework.
Open → 2609.18490v1