Visual hallucinations in diffusion models emerge early and persist through generation

Before the Token Commits: Trajectory-Level Benchmarking of Visual Hallucinations in Diffusion VLMs

Artificial Intelligence

Summary

Visual language models that generate image-based text answers do so step by step, but existing tests only check the final answer, missing when mistakes appear. The authors created a new way to track these models' answers at every step, revealing that false claims about images usually appear before the model commits to an answer and rarely get fixed later. This new test also finds hidden errors in counting and relationships that older methods miss. They propose a method to correct these mistakes early on, improving answer accuracy without hurting overall performance.

What this means in practice

  • For ai developers: Improve multimodal diffusion models’ accuracy by detecting and reducing hallucinations during intermediate generation steps rather than only at the final output.
  • For content moderation teams: Identify and address false visual claims in AI-generated image descriptions earlier in the generation process to enhance reliability.

Authors

Yadong Wang, Siping Yue, Yu Tian, Chuanxing Geng, Xiang Chen

Abstract

Multimodal diffusion language models generate responses by iteratively unmasking tokens, making each answer the endpoint of a multi-step trajectory rather than an immediate commitment. Hallucination benchmarks built for autoregressive models evaluate only the final output, and therefore cannot determine whether an unsupported claim in diffusion VLMs appears late or has already stabilized before any answer token is revealed. We introduce DynaHall, a trajectory-level benchmark of annotation-backed binary visual propositions covering object existence, counting, attributes, and relations, with controlled hard negatives graded by visual prior. DynaHall is paired with a commitment-aware protocol that records the intermediate answer tendency at every unmasking step alongside the committed output. Across five diffusion VLMs from three architecture families, visual hallucination is settled before commitment: an unsupported answer is already the preferred state while the answer position is still masked, and later unmasking steps rarely reverse it, so the failure is not introduced at the write step. This holds across decoding schedules, answer formats, and open-ended generation. DynaHall also exposes failures hidden by final-output metrics, including counting and relation collapse, prior-driven false positives, and attribute errors whose direction changes by type. Guided by this diagnosis, PGS (Pre-commitment Gradient Steering) edits still-masked answer states to reduce false positives, bringing the affirmation rate close to balance, and transfers to another architecture without degrading general ability. DynaHall and PGS suggest that hallucination should be measured and mitigated along the generation trajectory of diffusion VLMs, not only at the final answer.