Papers for

security camera operators

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Video models keep object identities better under occlusion and overlaps

Learning to Reason with Persistent Object States for Video Instance Segmentation

Abstract: Video segmentation models maintain object identities by carrying instance information across frames. Under prolonged occlusion, reappearance, or interactions between similar instances, however, an unreliable update can overwrite a valid history and cause persistent identity drift. We introduce POSReasoner, a trainable, plug-and-play framework that explicitly decides when an observation should change an object's state. Each persistent state records identity, confidence, and absence history. A sparse state-observation graph supports Propose-Verify reasoning: provisional associations are revisited using object history, predicted presence, and competition among identities. The verified decisions determine whether to retain, update, reactivate, or suppress each state, while a learned gate controls the evidence written back to memory. Only verified transitions update the persistent state used in subsequent frames. POSReasoner uses standard video annotations and keeps the base model frozen, enabling integration with diverse VOS and VIS architectures. Experiments across long-term VOS and VIS benchmarks show consistent improvements over strong baselines, with the largest gains under occlusion and object reappearance.

Mon 28 SeptComputer Vision and Pattern Recognition
The gist
Tracking objects in videos can be tricky when they get hidden or look very similar to others. The authors introduce POSReasoner, a tool that carefully checks when to update the memory of an object's position and identity, avoiding mistakes that happen from confusing objects over time. It uses a system that proposes possible matches and then verifies them before confirming updates, helping to keep track of objects accurately even after they disappear and reappear. This approach works with different existing video segmentation models and improves their accuracy, especially in difficult cases.
Open → 2609.35539v1

Lightweight network improves brightness and color in low light images

IDM-Net: A Lightweight Illumination-Decoupled Modulation Network for Low-Light Image Enhancement

Abstract: Low-light image enhancement (LLIE) remains challenging for lightweight models because illumination restoration and color fidelity are difficult to optimize simultaneously in the RGB color space. Although recent color-decoupled methods separate luminance and chrominance representations, they primarily optimize luminance as an enhancement target, leaving its potential as an explicit guidance prior largely unexplored during feature reconstruction. To address this limitation, we propose IDM-Net, a lightweight Illumination-Decoupled Modulation Network for low-light image enhancement. IDM-Net adopts a dual-encoder architecture consisting of a structure encoder that extracts multi-scale appearance features from the RGB image and a lightweight illumination encoder that learns illumination priors from the decoupled luminance (Y) channel. To effectively exploit these priors, we introduce an Illumination-Guided Modulation (IGM) module that injects multi-scale illumination cues into the decoder through spatially adaptive affine modulation, enabling accurate brightness restoration while preserving natural color consistency. Furthermore, we design a lightweight Feature Refinement Block (FRB) to progressively suppress degradation artifacts and recover fine-grained image details during reconstruction. Extensive experiments on multiple standard low-light image enhancement benchmarks demonstrate that IDM-Net achieves competitive performance among lightweight LLIE methods while maintaining an excellent balance between restoration quality and computational efficiency.

Fri 25 SeptComputer Vision and Pattern Recognition
The gist
Photos taken in low light often look dark and have weird colors. The authors introduce IDM-Net, a small and fast computer program that uses two different ways to understand an image: one to see shapes and colors, and another to understand lighting. By combining these, their method brightens photos while keeping the colors looking natural and fixing small flaws in the image. Their tests show this approach works well compared to similar lightweight methods.
Open → 2609.30962v1

Deep vision systems warn of failure using temporal instability cues

Visual Tripwires: Anticipating Failure in Deep Vision Systems

Abstract: Deep vision systems remain vulnerable to corruption, occlusion, and distribution shift despite strong benchmark performance. Existing reliability methods typically evaluate uncertainty at individual time steps and do not explicitly model how a system progresses toward failure. We introduce Visual Tripwires, a predictive reliability framework that uses temporal instability in model behaviour to anticipate impending failure. Our central hypothesis is that predictive degradation develops progressively through measurable changes in latent representations, prediction trajectories, and attention structure. Visual Tripwires captures these changes using representation drift, prediction oscillation, trajectory curvature, and attention entropy. A lightweight tripwire predictor aggregates these signals over a temporal window to estimate the probability of failure within a future prediction horizon. Experiments across multiple datasets, architectures, and progressive perturbation settings show that the proposed instability signals emerge before predictive degradation and provide earlier and more accurate failure warnings than conventional uncertainty estimation methods. These results demonstrate that temporal instability contains useful information about future model reliability and provides a practical basis for early warning in deep vision systems.

Wed 23 SeptComputer Vision and Pattern RecognitionMachine Learning
The gist
Deep vision systems can sometimes fail due to changes like corruption or occlusion, and current methods only check for uncertainty at each moment without looking ahead. The authors show that tracking how the system's internal signals change over time can predict failures before they happen. They measure changes such as shifts in how the model represents images, unstable predictions, and changes in attention focus. Their approach gives earlier and more accurate warnings of impending failure than traditional methods, making it easier to catch problems in advance.
Open → 2609.28099v1

Spatiotemporal flux probing captures fast videos with few photons

Spatiotemporal Flux Probing for Single-Photon Videography

Abstract: We address the problem of recovering high-speed videos from dynamic scenes under extreme photon sparsity. Existing methods rely on aggregating photon detections in local spatiotemporal windows to improve signal-to-noise ratio; however, this local grouping discards global structure and fails in low-light regimes where photon detections are sparse in space and time. In this work, we show that the information needed to recover both motion and illumination is encoded in correlations over the full space-time pattern of photon arrivals. Building on this insight, we develop a spatiotemporal flux probing theory and an algorithm that estimates the Fourier coefficients of the underlying intensity directly from the photon stream. We demonstrate that our approach (1) recovers fast motion and temporal illumination dynamics with substantially fewer photons than prior methods, (2) enables velocity-selective videography that automatically refocuses video onto specific detected motions, and (3) generalizes across sensing modalities including single-photon, event, and spike cameras.

Fri 18 SeptComputer Vision and Pattern Recognition
The gist
Capturing videos in very dark or low-light places is hard because cameras only get a few photons, or particles of light. Existing methods group photons nearby in space and time to get clearer images, but they fail when photons are very sparse. The authors show that looking at the whole pattern of photon arrivals over space and time contains the information needed to recover detailed motion and changes in brightness. They developed a new way to directly estimate the important features of the video from photon data, enabling faster and clearer video with fewer photons, and even zoom in on specific movements.
Open → 2609.22479v1