Multimodal model splits task signals from noise for better prediction

Structured Latent Modeling for Supervised Multimodal Information Decomposition

Machine Learning

Summary

When computers look at multiple types of input, like images and text together, they need to figure out which parts of these inputs really help solve a problem and which parts are just noise or unrelated. The authors created a new method that breaks down the information into parts that are shared across all inputs, parts unique to each input type, and irrelevant parts, all in a structured way within the computer’s internal reasoning. This helps the computer understand how each kind of input adds to solving the problem, leading to better predictions. They showed this works well on various tests with different types of data.

What this means in practice

Authors

Wanting Huang, Sanvesh Srivastava, Weiran Wang

Abstract

Multimodal prediction relies on diverse forms of evidence: information repeated across modalities, cues specific to a single source, and complex cross-modal dependencies that emerge only when inputs are considered together. While recent methods promote richer interactions, they lack a principled way to isolate these target-relative contributions within learned continuous representations. We introduce a framework that applies contrastive or masked objectives at intermediate layers, coupled with source-wise invertible normalizing flows and a supervised, low-rank latent variable model. This architecture explicitly factorizes the joint distribution into shared task-relevant variation, modality-specific predictive variation, and task-irrelevant dependence. Drawing connections to prior multimodal learning assumptions, our approach evaluates how modalities independently and jointly contribute to the target. Ultimately, this framework unites intermediate representation learning with structured likelihood-based guidance, offering a practical latent-variable lens for characterizing continuous multimodal interactions. Empirically, we demonstrate the effectiveness of our approach across diverse multimodal benchmarks, showing robust improvements in predictive performance.