Papers for

security monitoring operators

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

MetaSampling reduces frames for better long video question answering

MetaSampling: Making Frame Samplers Efficient for Long-Video Question Answering

Abstract: Frame selection is an important component of long-video question answering (VQA) with Multimodal Large Language Models (MLLMs). Existing frame-selection methods improve over simple top-$k$ embedding retrieval and uniform sampling, but are typically applied under a fixed global selection budget. We introduce \textbf{MetaSampling}, a training-free, plug-and-play sampling strategy that can be applied on top of existing frame selectors. MetaSampling improves downstream VQA efficiency by dynamically reducing the number of frames passed to the MLLM while preserving, and in some cases improving, answer accuracy. We evaluate MetaSampling across 36 paired frame-selector--MLLM-backbone--VQA-benchmark configurations. MetaSampling reduces the number of selected frames in all 36 configurations and improves accuracy in 25 of them, yielding an average frame reduction of $8.9\%$ while slightly improving accuracy overall.

Sun 27 SeptComputer Vision and Pattern Recognition
The gist
Answering questions about very long videos can be hard because a computer has to look at many video frames to understand what's happening. The authors created MetaSampling, a method that helps choose fewer important frames to show the computer without losing accuracy in the answers. MetaSampling works on top of other frame selection methods and can sometimes even improve how well the computer answers questions. It was tested with many different video question answering systems and consistently reduced the number of frames needed.
Open → 2609.33998v1

Accurate streaming detection of action changes in video sequences

Groupwise Selective State-Space Filtering for Accurate and Streaming Action Boundary Detection

Abstract: Action boundary detection partitions untrimmed video into intervals without assigning action classes. We present a boundary-detection adapter operating on pre-extracted video features, learning temporal representations via groupwise selective scans. Learned group fusion and temporal modeling convert these into transition scores, which are decoded into boundary timestamps. Trained with boundary-time supervision, the class-agnostic model is evaluated on Breakfast, GTEA, and 50Salads using temporal tolerances and bipartite matching, achieving boundary $F_1$ scores of 0.457, 0.622, and 0.611. A stateful variant enables feature-streaming inference with zero neural look-ahead, one-sample peak confirmation, and bounded memory. Downstream systems can subsequently assign s

Sun 27 SeptComputer Vision and Pattern Recognition
The gist
Detecting where one action ends and another begins in long videos is important for analyzing activities without labeling what the actions are. The authors describe a method that looks at video features over time in groups to find these boundaries accurately. Their approach works without knowing the action types and can handle live video streams with little delay and limited memory. It was tested on different datasets and improved the detection of action transitions.
Open → 2609.33400v1

Omnimodal agent improves referring video segmentation with precise reasoning

OPERA: A Unified Omnimodal Progressive Spatio-Temporal Reasoning Agent for Referring Video Segmentation

Abstract: Referring video segmentation with heterogeneous multimodal queries---spanning text, audio, and reference images---demands both robust cross-modal understanding and precise spatio-temporal reasoning. We propose OPERA (Omnimodal Progressive spatio-tEmporal Reasoning Agent), a unified reasoning agent built on a single MLLM that performs dual-axis progressive reasoning via three specialized stages. Along the temporal axis, a Temporal Reasoning Agent narrows the frame search space through coarse-to-fine filtering to identify the most informative key frame. Along the spatial axis, a Distillation Agent establishes what to locate via cross-modal semantic distillation, and a Grounding Agent enhanced with GRPO determines where the target appears, with dense mask propagation completing the pixel-level output. OPERA sets a new state of the art on OmniAVS and Ref-AVS and transfers zero-shot to standard referring video segmentation benchmarks.

Sun 27 SeptComputer Vision and Pattern Recognition
The gist
Referring video segmentation means identifying specific things or moments in a video based on clues like words, sounds, or pictures. The authors created OPERA, a system that carefully looks through video frames over time and uses different types of clues to find exactly what to highlight in the video. OPERA filters information step-by-step to focus on important frames and spots, producing detailed masks that outline the target. This new system works better than previous ones on several video segmentation tests.
Open → 2609.33338v1

Semi supervised visible infrared object detection with aligned consensus teacher

Aligned Consensus Teaching for Label-Efficient Oriented Object Detection in Weakly-Aligned Visible-Infrared Imagery

Abstract: Visible-infrared object detection (VIOD) detects objects with oriented bounding boxes from paired visible and infrared images. Existing methods depend on costly dual-modality annotations. Semi-supervised learning can reduce this burden, but extending it from single-modal detection to VIOD is challenging. In the practical image-pair-level setting considered here, only a few pairs are labeled in both modalities, while the rest are completely unlabeled. This limited supervision creates three challenges: (i) too few labeled boxes for robust cross-modal alignment; (ii) pseudo-label errors caused by branch-wise misses accumulate during self-training; and (iii) tail-class annotations become critically scarce as the labeling budget decreases. We propose Aligned Consensus Teacher (ACT) for label-efficient VIOD in this setting. Its Cycle-Consistent Region Alignment (CRA) combines cycle consistency and sparse anchors with reliability-weighted regional matching. Cross-Modal Consensus Mean-Teacher (CMC-MT) forms consensus pseudo labels under pair-preserving views to recover branch-wise misses and supervise unlabeled pairs. Text-Guided Cross-Modal Instance Augmentation (TG-CMIA) uses a vision-language scene prior to compose tail-class instance pairs while preserving RGB--IR offsets. To the best of our knowledge, ACT is the first framework to study semi-supervised VIOD under this image-pair-level setting. Experiments on DroneVehicle and VEDAI show consistent gains across annotation ratios. With 10\% labeled pairs on DroneVehicle, ACT reaches 94.3\% of the mAP obtained by the same detector under full supervision. Code and models will be available on GitHub to facilitate future work.

Wed 16 SeptComputer Vision and Pattern Recognition
The gist
Detecting objects in images taken by both regular cameras and infrared cameras usually needs lots of detailed labels, which are expensive to get. The authors developed a method called Aligned Consensus Teacher that uses only a small amount of labeled image pairs and many unlabeled pairs to learn accurately. Their approach smartly aligns objects between visible and infrared images and improves predictions by combining knowledge from both types of images. This method works well even when only a small fraction of the image pairs are labeled, achieving performance close to fully labeled training.
Open → 2609.18124v1