GUIDE: Guiding Internal Evidence with Language Instructions

2026-08-31Computation and Language

Computation and Language
AI summary

The authors introduce GUIDE, a method to help multimodal AI models better follow instructions about which information to use when making decisions. Instead of just controlling what the model outputs, GUIDE controls how the model relies on different types of evidence during reasoning and generation. They test GUIDE on tasks involving images, text, and audio, finding it makes the model more robust and able to adjust its information use as instructed. This approach helps models not just obey output commands but also manage their internal thought process about evidence.

multimodal modelsinstruction followingparameter-efficient adaptationevidence pathwaysgating mechanismrobustnessmultimodal reasoningimage-text-audio taskscontrolled perturbationautoregressive decoding
Authors
Soyeon Caren Han, Hyunsuk Chung, Jinwoo Kim, Seungyeon Ji, Kyungreem Han
Abstract
Large multimodal models follow instructions about what to generate, but not necessarily about what evidence to rely on. Hence, models may continue to depend on shortcut-associated cues even when instructions suggest otherwise. We introduce GUIDE, a framework for controlling internal evidence usage through language instructions. GUIDE combines grouped parameter-efficient adaptation with instruction-conditioned gating to modulate multimodal evidence pathways during reasoning and generation. We further introduce a pathway-level evaluation framework that characterizes instruction-conditioned evidence modulation through reliance sensitivity, controlled perturbation analysis, pathway modulation, and autoregressive decoding dynamics. Across multimodal reasoning, classification, and generation, GUIDE induces structured and instruction-aligned redistribution of evidence reliance while largely preserving task behavior. Experiments on GQA, TextVQA, MM-IMDb, CREMA-D, RAVDESS, and Flickr30K show that GUIDE improves robustness under targeted evidence perturbations and enables controllable modulation across diverse multimodal settings. This suggests that multimodal instruction following can extend beyond output control toward regulating how different evidence sources contribute to model predictions.