Papers for

multimedia software developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

SyncRA improves timing connections between sounds and images in videos

SyncRA: Learning Temporal Correspondence in Omni-Modal Models

Abstract: Recent omni-modal models demonstrate strong perception of audio and visual inputs, yet often struggle to connect what they hear with what they see at the same moment. This weakness in temporal correspondence can cause models to associate spoken cues with the wrong visual scenes, producing plausible answers grounded in incorrect audio-visual pairings. We diagnose this problem through controlled temporal swaps, revealing that model answers do not reliably follow changes in these pairings. To address it, we propose Synchrony-Guided Representation Alignment (SyncRA), a lightweight method for strengthening temporal correspondence between audio and vision. Specifically, SyncRA contrasts intermediate audio-visual representations within each video, aligning matching moments while separating mismatched ones to capture local temporal correspondence within a shared global context. The objective derives supervision directly from existing input timing, requiring no additional annotations and leaving inference unchanged. We evaluate SyncRA across four open omni-modal models spanning different sizes and architectures on five public video benchmarks. SyncRA consistently outperforms answer-only fine-tuning across all model-benchmark combinations, while substantially improving the ability to track changing audio-visual pairings in controlled evaluations. These results demonstrate that lightweight, targeted supervision can effectively strengthen temporal correspondence and translate into broad improvements in audio-visual question answering.

Mon 28 SeptComputer Vision and Pattern RecognitionArtificial IntelligenceMultimedia
The gist
Many AI models that look at videos and listen to sounds have trouble matching the right sound to the right image at the same time. This can cause them to misunderstand what they see and hear together. The authors found that these mistakes happen because the models do not reliably link sound and visuals by their timing. They created SyncRA, a method that teaches models to better connect sounds and images happening together without needing extra labels or changing how the models work when watching videos. This method improved several popular AI models on tests where they had to answer questions about videos with both sound and pictures.
Open → 2609.34363v1

Multimodal transformers build virtual encoders inside their layers

Virtual Encoders in Multimodal Transformers

Abstract: Multimodal language models traditionally rely on dedicated perceptual encoders to construct task-usable representations. More integrated architectures have recently emerged, which instead expose the shared transformer to lightly projected patches, audio frames, or discrete visual tokens. Where does this encoding happen when such representations are not provided? We find that the transformer can internalize this missing computation, constructing task-usable perceptual representations within its own early-to-middle layers before the downstream language model. We call this computational structure a Virtual Encoder. Across linear probing, similarities to perceptual encoders, and causal analyses, we identify signatures of this structure in models that receive perceptual tokens without continuous encoder-derived features. These analyses also suggest that the boundary between perception and language processing need not coincide within an architectural module. Instead, encoder-like computation can emerge as a functional regime within a shared transformer, providing a new perspective for understanding where and how multimodal models process perception.

Tue 22 SeptComputer Vision and Pattern Recognition
The gist
Multimodal language models usually use separate parts to understand images, sounds, or other sensory data before using language. The authors found that some advanced models can do this sensory processing inside their own transformer layers without separate encoders. These internal computations, called Virtual Encoders, happen early inside the model and make the sensory data ready for language tasks. This changes how we think about where perception and language processing happen in multimodal AI.
Open → 2609.26513v1

Multimodal tokens unify images and text for better retrieval and generation

FLAT: Resampling Image and Text into 1D Flexible-Length Aligned Transmodal Tokens for Retrieval and Generation

Abstract: Traditional multimodal representation learning and generation are two stages: a contrastive or self-supervised visual encoder is trained first, followed by a separate downstream generative model. This setup bottlenecks generative performance behind frozen embeddings. To bridge this gap, we revisit joint multimodal representation learning and generation to produce linearly interpolatable embeddings that are directly consumable by generative decoders. We present FLAT (Flexible-Length Aligned Transmodal representations), a representation pre-training framework that jointly optimizes a shared multimodal encoder alongside downstream text-to-image (T2I) and image-to-text (I2T) decoders. By combining contrastive alignment with bidirectional cross-modal generative objectives, FLAT ensures its representations function as both discriminative semantic descriptors and generative conditions. Architecturally, FLAT maps visual and textual inputs into a unified continuous 1D sequence space, applying nested dropout over prefix-K tokens to enable dynamic output lengths. A single pre-training stage allows FLAT to perform cross-modal retrieval and generation across variable prefix K, achieving a T2I GenEval score of 71.1. Task-specific fine-tuning aligns model performance with state-of-the-art baselines: 83.1 GenEval on T2I generation; 40.5 BLEU-4 and 138.6 CIDEr on MS-COCO image captioning; and Recall@5 scores of 86.8 (I2T) / 75.8 (T2I) on MS-COCO alongside 98.3 (I2T) / 93.6 (T2I) on Flickr30K. Finally, qualitative evaluations demonstrate that FLAT representations natively support linear interpolation, latent space arithmetic, and zero-shot composed retrieval.

Tue 15 SeptComputer Vision and Pattern Recognition
The gist
Combining images and text for tasks like captioning or image generation is hard because the two are usually processed separately. The authors developed a way to turn both images and text into a shared 1D sequence of tokens that work well for both recognizing and creating content. This helps a single model handle searching for images from text, generating images from text, and describing images with text, all with a flexible number of tokens. Their approach improves the quality of these tasks and supports interesting features like blending two images or texts smoothly.
Open → 2609.16591v1

Real-time video compression using deformable 2D Gaussian splatting

Deformable 2D Gaussian Splatting for Efficient 4K Video Compression

Abstract: Ultra-High-Definition (UHD) video presents significant challenges for efficient storage and real-time decoding. Learning-based methods, such as Neural Video Compression (NVC) and Implicit Neural Representations (INR), achieve competitive rate-distortion performance but suffer from high decoding latency and excessive memory usage. Meanwhile, Gaussian Splatting has recently attracted attention in the computer graphics community due to its ultra-fast rendering and high-fidelity visual quality. Despite these advantages, its application in video compression remains largely unexplored. To bridge this gap, we propose a real-time video compression framework that represents and compresses a Group of Pictures (GOP) using a coarse-to-fine multi-scale 2D Gaussian Splatting (2DGS) structure coupled with a lightweight deformation network. Experiments demonstrate that our method delivers rate-distortion performance in LPIPS that surpasses H.265 and other state-of-the-art learning-based video compression methods. Our work demonstrates the potential of Gaussian Splatting as a practical solution for efficient high-resolution video compression.

Sat 12 SeptComputer Vision and Pattern Recognition
The gist
Ultra-high-definition videos require a lot of storage and are slow to decode in real time. The authors propose a new way to compress videos by using a 2D Gaussian Splatting method combined with a lightweight deformation network. This approach can compress groups of video frames quickly and efficiently, resulting in better video quality compared to the popular H.265 standard and other recent AI-based methods. Their method is especially suited for very high-resolution videos, like 4K.
Open → 2609.14129v1

Multimodal emotion recognition improves with adaptive fusion and facial geometry

Multimodal Emotion Recognition in Conversations via Class-Wise Adaptive Modality Fusion and Affective Geometry

Abstract: Emotion Recognition in Conversations (ERC) requires integrating heterogeneous textual, audio, and visual cues while accounting for conversational context and emotional dynamics. We extend the Self-Distillation Transformer architecture for ERC with appearance+geometry visual representations, class-wise adaptive modality fusion, and a valence-arousal prior for affective transitions. On the MELD and IEMOCAP datasets, geometry-enhanced visual representations improve weighted F1 by 0.27 and 4.36 points over appearance-only features, respectively, while class-wise adaptive fusion provides further gains of 0.17 and 0.25 points over the original softmax gate. The valence-arousal prior yields targeted improvements of 0.30 and 0.74 accuracy points on emotionally shifted utterances while preserving performance on stable turns. These results indicate that structured facial cues, emotion-dependent modality weighting, and affective geometry provide complementary benefits for multimodal ERC.

Wed 9 SeptComputer Vision and Pattern Recognition
The gist
Understanding emotions in conversations is hard because people express feelings through words, tone, and facial expressions. The authors improve emotion recognition by combining facial appearance and precise facial movements, adjusting how much each type of signal counts depending on the emotion, and using a model of how emotions change over time. They tested their approach on two conversation datasets and found better accuracy, especially when emotions shift during the conversation. This shows that using detailed facial cues and emotion-aware mixing of signals helps machines understand feelings more reliably.
Open → 2609.09924v1