Papers for

multimodal machine learning engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Multimodal model splits task signals from noise for better prediction

Structured Latent Modeling for Supervised Multimodal Information Decomposition

Abstract: Multimodal prediction relies on diverse forms of evidence: information repeated across modalities, cues specific to a single source, and complex cross-modal dependencies that emerge only when inputs are considered together. While recent methods promote richer interactions, they lack a principled way to isolate these target-relative contributions within learned continuous representations. We introduce a framework that applies contrastive or masked objectives at intermediate layers, coupled with source-wise invertible normalizing flows and a supervised, low-rank latent variable model. This architecture explicitly factorizes the joint distribution into shared task-relevant variation, modality-specific predictive variation, and task-irrelevant dependence. Drawing connections to prior multimodal learning assumptions, our approach evaluates how modalities independently and jointly contribute to the target. Ultimately, this framework unites intermediate representation learning with structured likelihood-based guidance, offering a practical latent-variable lens for characterizing continuous multimodal interactions. Empirically, we demonstrate the effectiveness of our approach across diverse multimodal benchmarks, showing robust improvements in predictive performance.

Mon 28 SeptMachine Learning
The gist
When computers look at multiple types of input, like images and text together, they need to figure out which parts of these inputs really help solve a problem and which parts are just noise or unrelated. The authors created a new method that breaks down the information into parts that are shared across all inputs, parts unique to each input type, and irrelevant parts, all in a structured way within the computer’s internal reasoning. This helps the computer understand how each kind of input adds to solving the problem, leading to better predictions. They showed this works well on various tests with different types of data.
Open → 2609.35502v1