Papers for

security system developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Generative method finds people from text descriptions without labels

Generative Retrieval for Unsupervised Text-Based Person Search

Abstract: Text-based person search (TBPS) aims to retrieve images of a target person from a large image gallery based on a given natural language description. Most existing methods rely on supervised learning with manually annotated image-text pairs. In this paper, we explore unsupervised TBPS, with only unlabeled images. We propose GTR+, a two-stage generation-then-retrieval framework. In the generation stage, we introduce a tiered description generation framework designed to produce fine-grained and stylistically diverse textual descriptions through a three-tier sequential process. The base tier leverages an automated question-and-answer mechanism to generate basic visual attribute descriptions; the intermediate tier enhances fine-grained detail using an inter-sample contrastive mechanism; the advanced tier further enriches textual diversity via a stylized expansion mechanism. In the retrieval stage, to mitigate the impact of noisy pseudo texts, we develop an adaptive confidence-weighted retrieval learning framework. We model image-text pairs as clean or noisy using a Gaussian Mixture Model, calibrated by real-time image-text similarity and static text generation probability from the prior stage, yielding adaptive sample weights during training. Beyond that, we also contribute LargeFine-Person, a large-scale TBPS dataset with high-quality, fine-grained, and diverse textual annotations, enabling a practical and generalizable TBPS pre-training benchmark under unsupervised setting. Experiments on multiple TBPS benchmarks demonstrate the effectiveness and generalization of both GTR+ and LargeFine-Person. Code is available at: https://github.com/Flame-Chasers/GTR.

Fri 11 SeptComputer Vision and Pattern RecognitionArtificial Intelligence
The gist
Finding pictures of a person from a description usually needs many labeled examples, which take a lot of work. The authors propose a new two-step approach that first creates detailed and varied text descriptions from unlabeled images, then uses these to train a system to match text to images more accurately. They also introduce a large new dataset with detailed text annotations to help improve this kind of search. Their experiments show this method works well even without any labeled training data.
Open 2609.12965v1

Rgb thermal detection adapts to unreliable sensor data for better results

RA-SOD: Reliability-Aware RGB-T Salient Object Detection under Modality Degradation

Abstract: RGB-Thermal (RGB-T) salient object detection leverages complementary cues from visible and thermal modalities to improve robustness in challenging environments. However, in real-world scenarios, the reliability of each modality is inherently unstable: RGB images degrade under low illumination, motion blur, and noise, while thermal imagery often suffers from contrast compression and sensor artifacts. Such degradation introduces unreliable perceptual evidence that can mislead cross-modal fusion and significantly deteriorate detection performance. To address this challenge, we propose RA-SOD, a reliability-aware RGB-T salient object detection framework that explicitly models modality reliability and integrates it into feature learning and cross-modal fusion. First, we introduce a reliability-conditioned representation that adaptively compensates degraded modality features while preserving structural cues. Second, an uncertainty-guided dual-stream refinement strategy progressively corrects cross-modal representations while suppressing unreliable evidence. Finally, we propose a pixel-wise modality competition mechanism that dynamically selects modality cues according to spatial reliability for fine-grained fusion. Extensive experiments on four benchmarks (VT821, VT1000, VT5000, and VT-IMAG) demonstrate that RA-SOD achieves state-of-the-art performance and exhibits strong robustness under severe modality degradation. Code and models are available at https://github.com/zaoxienian/RA-SOD.

Fri 11 SeptComputer Vision and Pattern Recognition
The gist
Detecting important objects using regular and thermal cameras can be tricky when the camera images are blurry, noisy, or unclear. The authors created a system called RA-SOD that learns how reliable each camera is at every moment and uses this to combine their images better. This approach helps the system focus on trustworthy information, improving object detection even when one camera's image is poor. They tested this technique on several datasets and found it performs better than other methods.
Open 2609.12622v1

Person re-identification improves with clothing change awareness

SCORE: SubDistribution-aware Collaborative Knowledge Reinforcing for Cloth-Hybrid Lifelong Person Re-Identification

Abstract: Lifelong Person Re-Identification (LReID) aims to train a unified person retrieval model from a non-stationary data stream. Existing LReID methods mainly focus on scenarios where the clothing of each person is consistent. Recently, the Cloth-Hybrid LReID (CH-LReID) where cloth-consistent and cloth-changing data alternately occur, has emerged as a more practical and challenging scenario. Due to the conflict between clothing-relevant and clothing-irrelevant knowledge, the well-known catastrophic forgetting problem is significantly exacerbated in this task. To address this issue, we propose a SubDistribution-aware COllaborative Knowledge REinforcing (SCORE) framework, where our key idea is explicitly modeling the intra-identity diversity to continually consolidate distinct cloth-consistent and cloth-changing knowledge. Specifically, an Adaptive SubDistribution Modeling mechanism is developed, where a set of distributional subprototypes is assigned to each identity to capture the intra-identity diversity, improving the compatibility between cloth-consistent and cloth-changing knowledge. Then, a Distributional Knowledge Reinforcement scheme is introduced, where the knowledge of old distributional subprototypes is retained in the new ones by a collaborative aligning mechanism. Extensive experiments show that our SCORE achieves the state-of-the-art performance. Our code is available at https://github.com/zhoujiahuan1991/ECCV2026-SCORE

Fri 11 SeptComputer Vision and Pattern Recognition
The gist
Person re-identification means recognizing the same person across different images or video, even when they change clothes. This is a hard problem because the model can forget old information when learning new data, especially when clothing changes. The authors propose a method called SCORE that models different clothing styles for the same person separately and keeps old knowledge from being lost. This method helps the system better identify people despite these clothing changes and improves performance.
Open 2609.12577v1

Gait recognition improved with multimodal data and unified identity encoding

MMGait: Benchmarking and Unifying Gait Recognition across Heterogeneous Modalities

Abstract: Gait recognition is commonly studied using RGB videos or their derived silhouettes and poses. Yet human walking produces heterogeneous photometric, geometric, and motion cues that cannot be systematically examined with RGB-centered benchmarks. We present MMGait, a large-scale multi-sensor benchmark that brings visible, infrared, depth, LiDAR, and radar observations into sequence-level correspondence. It provides diverse modalities spanning appearance, contours, geometry, motion, and body structure. Under a shared impostor-augmented protocol, we evaluate single-modal recognition, cross-modal recognition via directed retrieval, and multi-modal recognition using task-specific experts. Across settings, modality rankings vary with probe conditions, cross-modal alignment remains difficult, and fusion often provides complementary gains. This analysis exposes a scalability problem: individual modalities, modality pairs, and fusion configurations are typically handled by separately trained experts. We formulate Omni-Modal Gait Recognition, which unifies single-modal, cross-modal, and multi-modal recognition within a shared identity space. OmniGait++ uses modality-specific front ends followed by a shared identity encoder to preserve modality-dependent cues while learning comparable identity descriptors. An anchor-guided fusion module aggregates modality subsets of varying size without frame-level synchronization. A jointly trained checkpoint covers all three recognition settings and accommodates modality subsets of different compositions and cardinalities. Experiments show OmniGait++ remains competitive with task-specific experts in many shared settings and extends to higher-cardinality fusion unavailable to fixed-pair models. The results establish MMGait as a common testbed for heterogeneous gait sensing and demonstrate the feasibility of unified recognition under varying modality availability.

Thu 10 SeptComputer Vision and Pattern Recognition
The gist
Recognizing people by the way they walk, called gait recognition, usually uses video images or simplified outlines. The authors point out that walking generates many different types of data like infrared, depth, radar, and more, which are rarely studied together. They created a big dataset called MMGait that combines all these types of data to compare and test gait recognition across different sensors. They also developed a system called OmniGait++ that can learn from and combine multiple sensor types at once, working well whether it uses one type of data or many. Their work shows it’s possible to have one system that adapts to different sensor combinations and still identify people accurately.
Open 2609.11601v1

Prototype-based method improves visible-infrared person matching

Prototype Matters: Modality-unified Prototype Self-distillation for Unsupervised Visible-infrared Person Re-identification

Abstract: Estimating reliable cross-modality association is crucial to unsupervised visible-infrared person re-ID. While optimal transport is shown to be a practical solution for cross-modality association, it suffers from the rigidness of hard label assignment without considering the impact of cluster noise. Moreover, enforcing only cross-modality contrast is also suboptimal, as it fails to jointly optimize the similarity relation within and across modality. In this paper, we propose a novel framework for cross-modality learning by well exploitation of prototypes: First, instead of contrasting with cross-modality prototypes, we show that modality-unified prototypical contrast facilitates better modality invariance by jointly and simultaneously optimizing similarity relation within and across-modality. Taking self-prototype as a steady teacher, we further refine the instance-prototype online relation through prototype-guided self-distillation. The two components are optimized in a unified framework, leading to a simple yet effective model. On standard VI-ReID benchmarks, we perform extensive comparison and analysis, validating the effectiveness of our proposed method. Code is available at: https://github.com/Terminator8758/PoSeD.

Thu 10 SeptComputer Vision and Pattern Recognition
The gist
Matching people captured in visible light with those in infrared is hard because the images look very different. The authors show that using shared group examples called prototypes helps computers learn better links between the two kinds of images. They also refine this process by letting the system teach itself, leading to better understanding inside and across both image types. Their combined method improves matching accuracy without needing labeled training data.
Open 2609.11514v1

Vision language models speed up error detection in image classifiers

Beyond Benchmarks: Using VLMs to Reveal Systematic Classification Failures Under Real World Conditions

Abstract: Verification and validation (V&V) of classification models is crucial to enable a wide range of sensor processing applications. Currently, the V&V process relies on time-consuming manual inspection of erroneous samples to find meaningful patterns. This work explores the use of Vision Language Models (VLMs) to speed up this laborious process. VLMs are trained to embed images into a semantically meaningful vector representation, from which human-interpretable systematic errors can be distilled. Deploying such VLM-based methods in a defence context introduces two major challenges: (1) the defence domain is underrepresented in the training data of VLMs, and (2) surroundings and context are less diverse than for other domains. This study provides an initial assessment of the suitability of VLM-based methods for V&V of defence applications. We propose a VLM-based error slice detection (ESD) method that independently groups and labels systematic errors made by a classification model. We demonstrate that this method is able to identify operationally-relevant artificially added perturbations in a non-military dataset. In a military context, our method clusters and describes images based on their surroundings, but also exhibits overlap between cluster descriptions. We further investigate the difference in embedding variation between our military and non-military dataset, which remains a topic of interest. Although the results do not yet warrant fully automated V&V through VLM-based ESD, they show that VLMs could be used to accelerate V&V processes in the future.

Thu 10 SeptComputer Vision and Pattern RecognitionArtificial IntelligenceEmerging Technologies
The gist
Checking where image recognition systems make mistakes usually involves a lot of manual work. The authors looked at using vision language models, which understand both images and words, to group and describe these errors automatically. They found these models can spot added errors in regular images and help cluster images in defense-related ones, though with some overlap. Their work shows this approach could make error checking faster but isn’t reliable enough yet to replace humans.
Open 2609.11126v1

Adaptive question answering improves police body camera video captioning

BodyCam-VQA: Enhanced Body-Worn Camera Video Captioning via Multimodal Reasoning and Probe Question Generation

Abstract: Police body-worn camera (BWC) footage has emerged as a critical aspect of law enforcement that ensures legal transparency, officer accountability, and the protection of civil rights. However, effectively processing this data remains a significant challenge due to its multimodal video format. BWC videos, in many cases, comprise chaotic scenes with low visual quality, rapid movement/interactions, and high-noise audio that make visual understanding a challenge for even SOTA multimodal models. Current Vision-Language Models (VLMs) frequently overlook critical forensic details, such as the presence of valuable evidence or the latent nuances of suspect-officer interactions, which are vital for fair legal outcomes and civilian/officer safety. To address these limitations, we propose an Adaptive Visual Question Answering (VQA) framework engineered for high-stakes law enforcement. Our framework employs a structured reasoning approach to extract fine-grained visual evidence that traditional captioning systems fail to capture. We experiment with multiple question generation models, including foundation models and fine-tuned open-weight models, to observe performance variation among question generation model implementations. Our results demonstrate that this VQA-driven architecture provides a more reliable, objective, and detailed record of enforcement events, ultimately serving as a powerful tool to protect both law enforcement officers and the public through AI-assisted forensic clarity.

Wed 9 SeptComputer Vision and Pattern RecognitionComputation and Language
The gist
Police body cameras record important interactions, but the videos can be hard to understand because they are often blurry, noisy, and full of fast actions. The authors created a system that asks detailed questions about what's happening in the video to find important details that usual methods miss. This helps make clearer and more accurate descriptions of the events, which can protect both police officers and civilians. They tested different ways to generate these questions to see which work best.
Open 2609.10815v1

Synthetic thermal data boosts visible to thermal face recognition

SynThermFace: Amplifying Limited Paired Data for Visible-Thermal Face Recognition via Synthetic Data Generation

Abstract: Face recognition (FR) is a widely used modality for biometric authentication, but conventional models rely on visible-spectrum imagery and degrade when high-quality RGB images cannot be captured. Cross-spectral face recognition addresses this limitation by matching visible images with other modalities such as thermal imagery, enabling more reliable performance in low-light, nighttime, and unconstrained conditions. However, progress is limited by the scarcity of paired visible-thermal data, which is difficult and costly to collect at scale. We propose SynThermFace, a framework that amplifies limited real visible-thermal supervision into larger paired adaptation datasets for cross-spectral face recognition. A diffusion model is first adapted using a limited set of paired visible--thermal images and then used to generate large-scale paired visible--synthetic thermal data from existing real or synthetic visible face datasets. The generated pairs are used to adapt a pretrained visible-spectrum face recognition model into a CFR model. Unlike synthesis-based approaches that require image translation at test time, the proposed method shifts generation to the training stage and performs inference with a single forward pass through the adapted recognition model. Under the same MCXFace real-pair protocol, PACT improves over the evaluated CFR adaptation baselines, isolating the effect of the proposed adaptation objective. Training PACT on larger generated paired datasets provides additional improvements over both the unadapted model and the real-pair PACT configuration. Cross-database evaluation on the Tufts dataset provides evidence that the learned representation transfers to an unseen database. The source code and trained models will be made publicly available.

Wed 9 SeptComputer Vision and Pattern Recognition
The gist
Face recognition usually works best with normal visible-light photos but struggles in the dark or low-light conditions. This paper shows how using computer-generated thermal images paired with visible images can improve face recognition when real thermal data is scarce. The researchers adapted a model to create synthetic thermal images from visible faces, which helps train recognition systems better. This method speeds up recognition since it doesn't require image conversion during use, and it works across different face datasets.
Open 2609.10303v1

Afid framework improves automated fingermark recognition and analysis accuracy

AFID: A Unified Open Framework for Automated Fingermark Identification, Quality Assessment and Feature Extraction

Abstract: Automated fingermark identification is the foundation of forensic investigation, yet progress in the field is held back by fragmented, closed-source solutions trained on private or discontinued data. We present AFID, a unified open-source framework for friction ridge image processing that performs recognition, quality assessment, and feature extraction based on a single shared encoder, trained exclusively on publicly available data. At its core is a fixed-length representation learned for identity discrimination, trained under heavy augmentation. Despite applying essentially no preprocessing beyond resizing and padding at inference, AFID sets a new state of the art in fixed-length fingermark recognition, leading identification across NIST SD 27 (67.6% rank-1) , SD 302 (54.9% rank-1), and SD 303 (67.6% rank-1), surpassing a commercial matcher on fingermarks. From the same frozen backbone, a quality assessment module predicts recognition utility more accurately than any compared baseline and generalizes across independent matchers, while lightweight decoders recover minutiae, ridge orientation, and segmentation competitive with dedicated methods. The framework proves that a single, efficiently trained encoder can support the full fingermark processing pipeline, from recognition through quality assessment all the way to feature extraction. To accelerate research on fingermark analysis even further, we release the code, models, and annotations to the community.

Mon 7 SeptComputer Vision and Pattern Recognition
The gist
Identifying fingerprints from crime scenes is important, but many current computer tools are closed and use private data. The authors created AFID, an open and unified software system that can recognize fingerprints, judge their quality, and find features using the same underlying model. AFID works well even without complex image adjustments and matches or beats commercial tools on standard forensic tests. It also predicts which fingerprints are good enough to use and extracts key details accurately, all with one shared model.
Open 2609.07439v1

Event camera system improves privacy aware emotion recognition

Emo-DVS: A Multimodal Benchmark for Privacy-Aware Emotion Recognition with Event Cameras

Abstract: Emotion analysis is a fundamental task in computer vision, but its practical deployment remains constrained by the privacy risks inherent to conventional RGB cameras. Bio-inspired event cameras present a promising hardware-level solution because they capture asynchronous brightness changes, thereby reducing exposure of facial identity details while leveraging high dynamic range for robust perception under challenging illumination conditions. Despite these advantages, existing event-based methods struggle in complex real-world settings due to limited dataset scales, simple acquisition conditions, and reliance on single-modality visual cues. To address these, we establish a challenging tri-modal benchmark with event, audio, and text modalities and propose the Information-Guided Gated Fusion (IGF) framework, which first pre-trains an event encoder on the FAU subset of Emo-DVS to capture fine-grained facial dynamics, then employs adaptive modality gating to suppress modality-specific noise, and finally leverages mutual information maximization to align robust cross-modal representations. To alleviate data scarcity, we introduce Emo-DVS, the first large-scale event-based emotion analysis dataset, which couples dynamic illumination with the Facial Action Unit (FAU) subset and emotion subset. Extensive experiments demonstrate that IGF achieves state-of-the-art performance.

Mon 7 SeptComputer Vision and Pattern RecognitionArtificial Intelligence
The gist
Emotion recognition using regular cameras risks exposing people's private facial details. The authors use special event cameras that capture only changes in light, which helps keep identities private while still detecting emotions. They created a new large dataset with event camera data plus audio and text to make emotion recognition better. They also developed a method to combine these different types of information effectively, improving emotion detection in real-life situations.
Open 2609.06928v1