Concept risk framework improves safety checks for image diffusion models
Concept-Level Risk and Calibration for Governance in Diffusion Foundation Models
Multimedia
Summary
Diffusion models create images from ideas but controlling them safely is complicated, especially when dealing with sensitive or private content. The authors developed a system to measure how risky these models are when generating certain concepts by turning the problem into a mathematical test of chance. They tested several popular diffusion models and found that the risks vary depending on how the model is accessed or prompted. Their approach helps compare different settings and improve the reliability of safety checks for image-generating AI.
diffusion modelsmultimedia generationsemantic controlprobabilistic auditconcept embeddingscalibrationrisk assessmentAI governanceimage generationmodel evaluation
Authors
Kun Xu, Yushu Zhang, Tao Wang, Shuren Qi, Barbara Carminati, Elena Ferrari, Yuming Fang
Abstract
Diffusion models have become a core paradigm for multimedia generation, offering powerful concept-driven controllability for personalization, semantic editing, and selective unlearning. However, as semantic control extends beyond natural-language prompts to learned embeddings and intervention pipelines, the safety and governance of these systems become increasingly difficult to evaluate in a unified manner, especially for safety-sensitive, identity-linked, and other privacy-relevant concepts. Existing studies mainly rely on heuristic audits, adversarial probing, or task-specific erasure benchmarks, and therefore provide limited support for systematic comparison across models, conditioning channels, and deployment conditions. We present a concept-level probabilistic audit and reporting framework for diffusion models. We formalize governance-relevant concept behaviors as Bernoulli semantic events induced by stochastic generation, and define a Concept Risk Operator that maps model-channel configurations to structured risk profiles, enabling comparison across prompting interfaces, learned embedding channels, models, and recorded conditions. We apply sample-level post-hoc calibration and configuration-level risk aggregation, and show that probability error can change thresholded actions near policy boundaries. Experiments on SD1.5, SD2.1, and SDXL reveal consistent yet non-uniform operational risk patterns across concept families, channels, recorded conditions, and shifted protocols. In particular, embedding-based access and obfuscated prompts expose risks often understated by standard-prompt evaluation. A pooled multi-protocol calibrator improves held-out probability reliability, but we do not claim transfer from a standard-only calibrator. CLRC provides a common audit schema for probabilistic and decision-aware governance of multimedia generation systems.