Benchmark tests multimodal image models with multiple visual instructions

VIF-Bench: Evaluating Visual Instruction Following in Multi-Reference Image Generation

Computer Vision and Pattern RecognitionArtificial Intelligence

Summary

Current AI systems that create images from multiple pictures and text instructions are not well tested when given many images and complex visual instructions at once. The authors created VIF-Bench, a set of 1,241 tasks that evaluate how well models combine several images and different visual instructions together. They found that models struggle to balance following instructions exactly with avoiding unwanted details in the images, and that giving instructions as visuals often works better than long text descriptions. This benchmark helps compare image generation tools fairly under these challenging conditions.

What this means in practice

  • For digital artists: Evaluate how image generation tools handle complex visual instructions combined with multiple reference images to improve creative control.
  • For software developers: Benchmark multimodal image generation models to select the best performing ones for applications requiring control from both images and visual cues.

Authors

Yuta Oshima, Masakazu Yoshimura, Masahiro Suzuki, Yutaka Matsuo, Hiroki Furuta

Abstract

Recent multimodal image generation models can take multiple images and textual instructions as input, enabling reference-based generation guided not only by text but also by visual instructions such as layouts, arrows, and pose cues. However, existing benchmarks do not evaluate the joint setting in which multiple references must be composed under multiple and heterogeneous visual-instruction images. To address this gap, we introduce VIF-Bench, a benchmark of 1,241 tasks designed to assess the edge of model capabilities in this joint setting by covering: (i) multi-reference generation (up to 7) under multiple heterogeneous visual instructions (up to 6), (ii) cases where reference images can potentially compete with visual instructions (e.g., a strongly posed subject vs. a target pose), and (iii) controlled comparison of visual instructions with text descriptions at different levels of specificity. Using these capabilities, we uncover three findings: (1) models face an adherence-artifact trade-off: once models reach stronger visual instruction adherence, stronger adherence tends to coincide with more instruction artifacts in generated images, (2) visual instruction adherence tends to be lower on tasks whose reference images carry a salient state of the controlled attribute (e.g., a neon-lit subject under a light-direction instruction), most consistently for light and wind, and (3) for models that can understand visual instructions, it is often better to provide visual constraints directly rather than describe them in text; when using text, a moderate level of detail works better than an exhaustive description. VIF-Bench is released as an open benchmark to establish a basis for fair comparison in controllable multi-reference image generation.