Papers for

video analytics teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Label guided method improves 3D CNN video action recognition

Label-Guided Knowledge Distillation for 3D-CNNs in Action Recognition

Abstract: As a key model compression technique, knowledge distillation aims to transfer knowledge from a high-capacity teacher model to a lightweight student model for enhancing the latter's performance. In this work, we reviewed the feature knowledge distillation for 3D-CNNs and observed that most feature distillation methods in video analysis are simple adaptations of those used in image analysis, often neglecting the differences of video features in the temporal dimension. To address this issue, we proposed Label-Guided Knowledge Distillation (LGKD) to guide the distillation of student model features using ground truth labels. Our method entails two components: sample-wise distillation and class-wise distillation, enabling the student model to learn feature representation of the teacher model at two levels. Sample-wise distillation utilizes label information and the teacher's probability distribution to guide the learning of features that significantly impact temporal accuracy while mitigating noise. Meanwhile, class-wise feature distillation employs a prototype network to further capture the relational knowledge among samples within the same category, enhancing the student's ability to learn higher-dimensional semantic information and improving model generalization. To demonstrate the effectiveness and superiority of our method, we conducted comprehensive experiments on two benchmark action recognition datasets, UCF101 and HMDB51, achieving competitive results.

Fri 11 SeptComputer Vision and Pattern RecognitionArtificial Intelligence
The gist
Video action recognition models often use simpler versions of image analysis techniques that ignore the timing in videos, which can reduce accuracy. The authors introduced a way to help smaller, faster video models learn better from bigger ones by using labels to guide the learning process at both individual video and category group levels. This approach helps the small models understand timing and category relationships better, which improves their accuracy. They tested this on standard video action datasets and found their method competitive with existing techniques.
Open 2609.13024v1

AutoSkill improves long video question answering by adapting frame selection

One Skill Does Not Fit All: Automatic Discovery and Taxonomy-Guided Routing of Frame-Selection Skills for Long-Video Question Answering

Abstract: Long-Video Question Answering (LVQA) requires locating decisive evidence in hour-scale videos under a limited frame budget. Most training-free methods apply the same frame-selection strategy to all questions, despite substantial variation in the evidence required by different question types. Our analysis shows that the relative effectiveness of frame-selection strategies varies across semantic categories and benchmarks, motivating adaptive evidence acquisition. In this paper, we introduce AutoSkill, a source-supervised framework for automatically discovering and routing executable frame-selection skills. Starting from a small labelled source pool, LLM agents iteratively propose, implement, evaluate, and refine candidate skills. For a target benchmark, AutoSkill uses only unlabelled question and option text to induce a shared semantic taxonomy, rewrite labelled source examples into the target style, and estimate a category-to-skill mapping. Neither target videos nor target answers are used in this process. At inference time, each question is assigned one skill, which selects the frames used in a single inference of the frozen video MLLM. Across five long-video benchmark splits, AutoSkill improves Qwen2.5-VL-7B and Qwen3.5-4B by 2.4% and 1.2%, respectively, demonstrating the effectiveness of our AutoSkill.

Fri 11 SeptComputer Vision and Pattern Recognition
The gist
Long videos contain a vast amount of information, but answering questions about them quickly is hard because the system can only look at a few important frames. The authors found that different types of questions need different ways to choose these key frames. Their method, AutoSkill, automatically discovers and assigns the best frame-selection strategies for different question categories without needing direct video or answer information from the new videos. This improves the ability of video language models to answer questions on multiple long video datasets.
Open 2609.12517v1

Video model improves understanding by adaptive evidence gathering

VLX-VR: An Agentic-Aware Video Reasoning Model

Abstract: Real-world video understanding requires integrating visual, audio, textual, and temporal evidence distributed across a video. Yet many pipelines use a fixed video context and single-pass inference, limiting adaptive evidence acquisition when observations are incomplete, ambiguous, or conflicting. We present VLX-VR, an agentic-aware video reasoning model trained within a video reasoning framework defined by a Think--Memory--Observation loop. At each step, VLX-VR determines the needed evidence, invokes read_memory or write_memory, incorporates the returned Observation, and decides whether to continue or produce the task output. We train VLX-VR with multimodal data, including videos and agent trajectories, using reinforcement learning to learn evidence acquisition, memory use, and termination. On MINERVA, VLX-VR achieves state-of-the-art performance among the models included in our comparison, with 78.79% accuracy. Under the original three duration groups, its accuracies are 76.70%, 78.73%, and 80.92%, with a cross-duration accuracy variance of 2.97~$\mathrm{pp}^2$. On correctly answered samples, 96.20% of VLX-VR's reasoning traces are consistent with the MINERVA reference reasoning traces and the evidence described by them, while approximately 75.80% of all evaluated samples satisfy both answer correctness and this evidence-grounded trace criterion. These results show strong performance and broadly stable behavior across durations, while counting, state changes, causal reasoning, and spatial perception remain challenging.

Wed 9 SeptComputation and LanguageComputer Vision and Pattern Recognition
The gist
Understanding videos needs looking at pictures, sounds, words, and time clues together. Many video systems look at fixed parts once, missing important details when things are unclear. The authors created VLX-VR, a system that decides what information to check next and remembers what it learned, improving how it understands videos. It uses trial and error learning and works well on a video reasoning test with good accuracy and consistent reasoning.
Open 2609.09985v1

Input defenses fail against observation attacks on videollms

Do Input-Level Defenses Transfer to Observation-Level Attacks on VideoLLMs?

Abstract: Video Large Language Models (VideoLLMs) are increasingly deployed in safety-critical applications such as content moderation and video analytics. To process long videos efficiently, VideoLLMs rely on frame sampling, token compression, and modality fusion, which together form an observation pipeline that reduces the raw video to a compact internal representation. Recent observation-level attacks exploit this pipeline to prevent the model from perceiving harmful content, yet no defense has been explicitly designed for this threat. We introduce DefTEval, a controlled evaluation framework that systematically assesses whether input-level adversarial defenses, which operate on the pixel content of already-sampled frames, can mitigate observation-level attacks. Across five VideoLLMs, eleven representative defenses, and five attack types, we find that input-level defenses offer limited and inconsistent protection, with harmful detection rates frequently near zero. Critically, defenses fail even against attacks that embed harmful signals in every sampled frame, indicating that the bottleneck extends beyond sampling omission to the suppression of signals that do enter the model. Token compression discards localized features, and modality fusion systematically down-weights weakened visual signals. Furthermore, defense effectiveness is dominated by model architecture rather than by the defense method itself, and detection rates vary drastically across content categories, exposing structural weaknesses in temporal reasoning. These findings demonstrate that securing VideoLLMs requires system-level robustness mechanisms spanning sampling-aware coverage guarantees, token-level preservation of safety-relevant features, and modality-balanced fusion.

Tue 8 SeptComputer Vision and Pattern RecognitionCryptography and Security
The gist
Video Large Language Models (VideoLLMs) help analyze videos in important areas like content moderation but simplify their input by selecting and compressing video frames. The authors find that protections applied directly to these frames often fail to stop attacks that exploit how the model processes all video information together. These attacks hide harmful content from the model even when present in every sampled frame. The study shows that improving safety in VideoLLMs needs new protections across the entire system, not just at the input level.
Open 2609.08331v1