Papers for

video surveillance teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Hyperspectral tracker improves video target tracking with spectral memory

HyperDAM: Hyperspectral Distractor-Aware Memory with Amodal Expansion for SAM 3 Tracking

Abstract: Hyperspectral video provides material cues that can disambiguate targets with similar false-color appearance, yet foundation-model trackers update memory primarily from spatial and appearance evidence. We present HyperDAM, a DAM4SAM3-based hyperspectral tracker with three principal contributions. First, HOTC2026-Modal adds human-verified frame-wise modal masks and mask-tight boxes to all 481 organizer-provided HOTC 2026 videos. Second, a frame-zero-calibrated HSI gate rejects spectrally inconsistent updates to the distractor-resolving memory (DRM) without altering the current prediction. Third, a causal spatiotemporal expander adds outward-only amodal corrections from frozen SAM features. Static-scene recovery and empty-mask RTS smoothing address target switches and full occlusion. Model selection prioritizes cross-domain robustness over leaderboard-specific optimization. The final system ranked second in HOTC 2026, achieving 68.0093% AUC and 87.7703% DP@20 in the organizer's private evaluation.

Mon 28 SeptComputer Vision and Pattern Recognition
The gist
Tracking objects in videos can be tricky when different things look very similar. The authors developed a system called HyperDAM that uses special hyperspectral video data to tell apart targets based on their material properties, not just their looks. They also improved how the system decides when to update its memory to avoid confusion from things that look like the target but are different. Their method ranked second in a major competition, showing better accuracy in following targets over time.
Open → 2609.34396v1

Multi-camera pedestrian location without calibration or training needed

CAT-Free: Multi-View Pedestrian Localization without Calibration, Annotations, or Target-Scene Training via Adaptive Geometric Filtering

Abstract: Multi-camera pedestrian localization is useful for wide-area monitoring in public and commercial spaces. However, deploying these systems often requires considerable setup for each new environment. Existing methods typically require camera calibration, position annotations, or target-scene training. CAT-Free removes all three requirements. It uses synchronized RGB video as its only scene-specific input. Camera configuration is estimated directly from the video. Pedestrian locations are then estimated by combining observations from multiple cameras. Automatic camera estimation is not always accurate. This can produce unreliable pedestrian locations. CAT-Free therefore introduces two adaptive geometric filters. They remove unreliable position estimates. Their thresholds are estimated from each input sequence. CAT-Free achieves 82.5, 84.5, and 65.7 MODA on WildTrack, MultiviewX, and GMVD. It uses no supplied calibration, position annotations, or target-scene training. Published methods using such scene-specific information report 88.2--95.0 MODA on WildTrack and 83.9--96.5 on MultiviewX under their respective protocols. CAT-Free also transfers without retuning. It reaches 74.9 MODA on four additional sequences and 78.6 on an unseen 8-camera installation. Finally, localization uncertainty predicts MODA with $r=-0.98$. This provides a label-free estimate of localization reliability.

Mon 28 SeptComputer Vision and Pattern Recognition
The gist
Tracking where people are in big spaces using many cameras usually requires setting up each camera carefully, marking exact positions, or training the system beforehand. The authors developed a method called CAT-Free that only needs synchronized video from cameras and no extra setup or training. It figures out camera positions automatically from the video and combines information to locate people. Because the automatic setup can be inaccurate, it uses special filters that remove unreliable location guesses, making the system work well without prior calibration.
Open → 2609.34302v1

Pipeline filters miss workers in rare low poses on construction sites

Auditing Quality Filters for Long-Tail Human Data Curation

Abstract: Robots on construction sites must detect workers who are kneeling or bending, which we call low poses. These workers can be lost from training datasets during automatic labeling. We study a pipeline that detects people, estimates their body joints using NLF, and groups similar poses. Low poses account for only about 2 percent of the retained examples. This low share may partly reflect the pipeline's quality filter, which rejects examples with low detection confidence or uncertain joint estimates. We examine this filtering using four alternative pose clues: bounding-box shape, vertical body span, pose grouping aligned to the scene's vertical direction, and image appearance. All four suggest that low poses are rejected by the filter more often. Separately, controlled simulated scenes show that a person detector fine-tuned on a public construction dataset misses more workers in these poses even when we correct their bounding box height is matched to that of standing workers. These findings suggest that low poses are scarce and hard to find, and we cannot rely on bounding boxes or poses for long-tail human data curation.

Fri 25 SeptComputer Vision and Pattern RecognitionArtificial Intelligence
The gist
When robots look for workers on construction sites, it’s important to spot people who are bending or kneeling, since these low poses are less common and harder to detect. The authors found that the automatic system’s quality checks tend to reject these rare low poses more often than normal standing ones. They tested this using different clues about the body’s shape and image features and confirmed the effect. This means that these low poses are not only rare but also difficult to find with current automatic detection methods.
Open → 2609.31896v1

Lifecycle-aware memory improves multi-object tracking with SAM2 models

LiAM-SAM: Lifecycle-Aware Memory for Robust SAM2-Based MOT

Abstract: Segmentation-based multi-object tracking (MOT) with foundation video models such as SAM2 offers strong localization quality, yet remains fragile in crowded, real-world scenes. In detector-prompted SAM2 pipelines, failures typically arise at three stages of the object lifecycle: (i) erroneous or duplicate track initiation, (ii) memory drift during close interactions, and (iii) unreliable re-identification after long occlusions or re-entry. These errors corrupt object memory and accumulate over time, making long-horizon tracking unstable. In this paper, we reframe MOT as a lifecycle memory integrity problem. We present LiAM-SAM, a Lifecycle-Aware Memory (LiAM) framework with targeted mechanisms for each of the three failure modes. At track birth, to prevent faulty or duplicate initiations, we apply contrastive track initiation, which conditions each prompt on existing nearby tracked instances. To preserve memory integrity during strong interactions, we introduce motion- and geometry-grounded memory correction that resolves interaction confusions and suppresses drift. For reliable re-identification after disappearance, we maintain an adaptive context memory that promotes diverse and trustworthy references as long-term identity anchors. Finally, similarity aware spatial pruning optionally selects the memory tokens to retain at cross-attention time, improving efficiency with minimal accuracy loss. LiAM-SAM represents a modular, detector-agnostic, SAM2-based MOT system that achieves state-of-the-art HOTA and IDF1 on the evaluated benchmarks. In association-challenging environments, our ablations show that LiAM improves a detector+SAM2 baseline by +10.5 HOTA, +17.4 AssA, and reduces identity switches by 96%.

Wed 23 SeptComputer Vision and Pattern Recognition
The gist
Tracking multiple objects in videos can get confused when objects are close together, disappear, or reappear after hiding. The authors rethink this problem as keeping good memory of each object’s history. They build a system called LiAM-SAM that carefully manages object information through their whole lifecycle, fixing mistakes like duplicate tracking or mixing up identities. Their approach helps machines track objects more accurately and reliably, especially in crowded or challenging scenes.
Open → 2609.28078v1

Multi-object tracking improved by direction-aware occlusion handling

DOA-SORT: Directional Occlusion-Aware Multi-Object Tracking with Distributional Observations

Abstract: Identity association in multi-object tracking (MOT) is vulnerable to partial occlusion, truncated detections, and fluctuating confidence scores. Existing motion-dominant trackers commonly represent occlusion as a scalar penalty. This treatment misses the directional observation bias caused by occlusion: left, right, top, and bottom occlusions distort the location and shape of a detection in different ways. We propose \ours{} (Directional Occlusion-Aware SORT), an online and training-free tracker that models these biases explicitly. First, it infers a soft front--back ordering from box overlap and relative bottom positions, and estimates directional occlusion coverage and depth. It then constructs a mixture of one clean and four directional occlusion observation components. The model uses a five-dimensional observation comprising box center, area, confidence, and aspect ratio, and adapts observation noise to predicted occlusion and detection confidence. The directional mixture likelihood is used in high-confidence association, low-confidence association, and track recovery; ambiguity penalties and local order-consistency swaps further reduce identity errors among nearby objects. On the DanceTrack validation split, \ours{} improves HOTA from 63.00 to 66.34, AssA from 45.10 to 49.57, and IDF1 from 62.19 to 65.28 over OA-SORT with the same detector and evaluation protocol. The gains are concentrated in association quality while detection accuracy remains stable. Additional local evaluations on MOT17 and MOT20 train splits characterize cross-dataset behavior under the same no-ReID tracking protocol.

Sat 19 SeptComputer Vision and Pattern RecognitionInformation Retrieval
The gist
Tracking multiple moving objects on camera gets tricky when some objects block part of others. The authors introduce a new way to understand how objects get hidden from different sides and how this affects their detected position and shape. Their method models these directional hiding effects to better guess where objects really are, helping keep their identities correct over time. This leads to more accurate tracking without needing extra training data.
Open → 2609.22706v1

Simulated video features help find people from text descriptions

SCOUT: Sim-to-Real Text-Based Person Retrieval by Embedding-Space Prediction over Frozen Video Features

Abstract: Text-based person retrieval under a sim-to-real gap (synthetic training data, a real-image gallery) is usually tackled with costly fine-tuned cross-encoders. We ask whether a frozen-encoder system can compete. We present SCOUT, which casts cross-modal retrieval as prediction in embedding space. A trainable predictor maps the patch tokens of a frozen video encoder into the embedding space of a frozen text encoder under a bidirectional InfoNCE objective, and no encoder is fine-tuned in the base model. The video encoder is V-JEPA, the text encoder is EmbeddingGemma, and the predictor is initialized from a Qwen3.5-0.8B decoder. We make three findings. First, the best frozen text encoder is simply the one whose geometry best matches the video features. A training-free alignment score ranks three candidate text encoders in the same order as their retrieval accuracy on our held-out split (Spearman $ρ= 1.0$); a fourth, LLM-based encoder shows the rule is metric-dependent, holding for a neighborhood-overlap score ($ρ= 0.8$) but not for a linear probe ($ρ= -0.2$). Second, two precision-targeted levers, parameter-efficient ExPLoRA adaptation of the video encoder and a training-free attribute-decomposed reranker built on a vision-language model, improve the top-rank precision that otherwise limits the frozen system, adding 2.2 points of leaderboard R@1. Third, a local-versus-public calibration study explains which interventions transfer to the real domain. On AI City Challenge 2026 Track 4 the full retrieve-fuse-rerank system reaches 84.25 mAP@10 on the final leaderboard, while a single frozen model submitted alone reaches 60.63. Our trained components cost about 95 GPU-hours. CMP, the dataset authors' fine-tuned cross-encoder that trains for sixteen GPU-days, is one fusion member of the full system, not an alternative. Code and annotations: https://github.com/abtraore/SCOUT-ECCV

Wed 16 SeptComputer Vision and Pattern RecognitionInformation Retrieval
The gist
Finding people in real videos using text descriptions is hard when training only uses simulated videos. The authors show that using fixed video and text processing models plus a trainable middle part can work well without retraining big models. They find that matching how video and text models represent data is important, and adding smart reranking improves results a lot. Their full system performs competitively on a real video challenge and uses far less training time than prior methods.
Open → 2609.19483v1

Video anomaly detection improved with training-free severity probing

Probe-VAD: Ordinal Likelihood Probing for Training-Free Video Anomaly Detection

Abstract: Video anomaly detection (VAD) aims to localize anomalous events in untrimmed videos. Vision-language models (VLMs) provide rich visual understanding for training-free VAD, but existing approaches impose restrictive interfaces between visual understanding and anomaly scoring. Caption-based pipelines compress visual evidence into text, potentially discarding subtle cues, while direct numerical generation forces the model to express its judgment through a small set of predefined scores. Such interfaces can obscure subtle differences in anomaly severity, causing visually distinct clips to receive similar representations or scores and thereby limiting the resolution of anomaly ranking. We propose \textbf{Probe-VAD}, an ordinal binary-probing framework that directly probes severity preferences from a frozen VLM. Given raw video clips, Probe-VAD queries ten ordered severity thresholds and extracts constrained \textit{YES}/\textit{NO} continuation likelihoods. Their normalized preferences form a cumulative severity profile, from which tail evidence is aggregated into a continuous anomaly score, with isotonic projection enforcing ordinal consistency. Experiments on public VAD benchmarks demonstrate superior performance with low computational cost. Probe-VAD provides a simple interface for translating frozen VLM visual understanding into continuous, rank-sensitive anomaly scores without task-specific training or caption-based compression. Code is available at: https://github.com/yvestine/COVAS-VAD.

Tue 15 SeptComputer Vision and Pattern Recognition
The gist
Video anomaly detection tries to find unusual events in long videos without prep work. The authors found that current methods lose important details when turning video into text or limited scores. They created Probe-VAD, which asks a fixed video-language model about levels of anomaly severity using simple yes/no questions. This method keeps subtle differences and ranks anomalies better without extra training, using a clever math step to keep the rankings consistent. Tests show it works well and uses less computing power.
Open → 2609.17211v1

Video segmentation predicts object movement without extra inputs

DynEoMT: Learning Object Dynamicity from Online Segmentation Queries

Abstract: Video segmentation models recognize and track objects over time, but they do not indicate whether each segmented region moves independently of the observing camera. This dynamicity attribute cannot be inferred from semantics alone and is confounded by camera ego-motion. We introduce \method, an online framework that augments query-based video segmentation with region-level dynamicity prediction. It jointly produces the original segmentation outputs and a dynamic or static state for each predicted region. At inference, DynEoMT uses only the current frame and propagated queries, without optical flow, depth, camera pose, previous RGB frames, or feature maps. Because established video segmentation benchmarks do not annotate this attribute, we also introduce a class-agnostic offline supervision pipeline using camera-compensated optical flow and confidence-aware temporal filtering. Across VIPSeg, OVIS, YouTube-VIS 2022, and VSPW, DynEoMT achieves balanced accuracies of 84.3, 68.0, 68.6, and 87.6, respectively, while largely preserving segmentation performance. These results show that segmentation-region dynamicity can be learned from propagated queries, enabling its online prediction without a dedicated motion-processing pipeline at inference. The complete code will be released as open source to enable full reproduction of the method and experiments.

Sun 13 SeptComputer Vision and Pattern RecognitionRobotics
The gist
Video segmentation helps find and follow objects in videos, but it can't tell if an object is moving on its own or if the whole camera is moving. The authors created DynEoMT, a method that can label each object as either moving independently or static, using only the current video frame and some previous information from earlier frames. They also developed a way to train the system without needing special labels showing which objects move. Their method works well on several video datasets while keeping good object tracking accuracy. This means it can figure out object movement online without needing complex extra data like optical flow or depth.
Open → 2609.14466v1

Black-box attack lowers accuracy of human pose and action models

A Black-Box Adversarial Attack on Human Pose Estimation and Keypoint-Based Action Recognition Models

Abstract: Human pose estimation and keypoint-based action recognition models are increasingly deployed as components of video understanding pipelines, yet their vulnerability to adversarial attacks remains insufficiently studied. Temporally coherent black-box attacks have been previously studied in visual object tracking, where the attack feedback can be defined using bounding-box overlap measures such as Intersection over Union (IoU). However, human pose estimation produces keypoint configurations rather than enclosing boxes, making box-level similarity poorly suited for measuring pose degradation. We propose OKS Attack, a decision-based black-box attack that uses Object Keypoint Similarity (OKS) as the attack feedback signal, directly targeting the spatial structure of human poses rather than their enclosing boxes. Experiments on the Penn Action dataset show that OKS Attack consistently reduces pose quality across evaluated pose estimators, with mean OKS decreases ranging from 0.0802 to 0.1494. In a downstream cross-dataset action-recognition evaluation, the attack reduces accuracy by 6.18 to 13.86 percentage points and outperforms query-matched random-noise perturbations. The attack is effective across both top-down and single-stage pose estimation models. The source code will be made publicly available at https://github.com/KacperM33/OKS_attack

Mon 7 SeptComputer Vision and Pattern Recognition
The gist
Human pose estimation and action recognition models try to understand body movements in videos, but this paper shows they can be fooled by special subtle changes. Instead of using simple box comparisons, the authors created a new way to attack these models by confusing how key body points are detected. Their method lowers the accuracy of pose detection and action recognition across many models. This shows these systems are vulnerable to attacks that work without knowing model details.
Open → 2609.08013v1