Investigating Social Bias in Narrative Image Generation

2026-08-03Computer Vision and Pattern Recognition

Computer Vision and Pattern RecognitionArtificial Intelligence
AI summary

The authors studied how text-to-image models, which create pictures from descriptions, might show social biases. They found that these biases are not just in single photos but become stronger and more obvious in picture stories and comics, where the sequence and arrangement of images add more bias. Their work suggests it’s important to check these models with different types of visual stories, not just photos, because biases can be hidden or more visible depending on the format.

text-to-image generationsocial biasstoryboardscomicsbias evaluationimage generation modelsnarrative visual formatsevent sequencingcharacter positioning
Authors
Junyeong Park, Sowon Min, Euna Jang, Soobin Kim, Jiho Jin, Hyunseung Lim, Gahyeon Bae, Hwajung Hong
Abstract
Text-to-image (T2I) generation models are increasingly embedded in applications such as media content creation and education, raising concerns about how their outputs may reproduce social biases. Prior work has shown that T2I models exhibit social biases, yet existing evaluations largely focus on a photo generation task. As a result, it remains unclear whether and how such biases manifest in more narrative visual formats, such as storyboards and comics, where characters and events are presented across multiple panels. In this work, we compare bias expression across photo, storyboard, and comic generation in six T2I models by adapting BBG, a text-based bias evaluation framework, to image generation. Our results show that proprietary models generate 25.9% biased outputs in photo generation on average, with biased outputs increasing by 9.6pp in storyboard generation and 18.2pp in comic generation. We also find that photos mainly encode biases through subtle visual cues, while storyboards and comics reveal them more explicitly through event sequencing, character positioning, narrative resolution, and textual elements. These findings show that biases that remain less visible in photo generation may surface in narrative visual formats, highlighting the importance of evaluating T2I systems with diverse visual formats beyond photo generation.