Papers for

model safety teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

RISE improves detecting unsafe prompts in text to image AI

RISE: Red-teaming via Iterative Strategy Evolution for Modern Text-to-Image Models

Abstract: On modern production text-to-image systems, successful policy violations are rare, and previously effective human-written seeds are often patched out. Current automated red-teamers are poorly matched to this regime in two ways: unreliable success measurement and poor exploration. First, we find that judges widely used in prior T2I red-teaming work are unreliable under vague unsafe-content targets: they either miss true violations or reward benign borderline images on hardened APIs. We therefore define strict category-specific success criteria and calibrate strong VLM judges against human labels. Second, we show that broadly used prompt-modification pipelines do not solve the exploration problem: on harder guardrail settings they remain tied to seed prompts, fail to transfer, or cannot bootstrap positive examples. We introduce RISE, which evolves reusable strategies used to generate prompts rather than rewriting them one by one. The best discovered strategies are then reused to generate attacks across new scenarios. On DALL-E 3, Nano Banana 2 (Google) and GPT-Image-2, RISE reaches up to 13% human-verified ASR; under the same calibrated evaluation, prior methods with reported ASR as high as roughly 30% fall to near zero.

Mon 28 SeptArtificial Intelligence
The gist
When AI systems create images from text, it's important to stop them from making unsafe or harmful pictures. The authors found that previous methods to test these AI were not very reliable and often missed problems or gave false alarms. They created a new method called RISE that better measures when the AI breaks safety rules and finds ways to trick the AI using evolving prompt strategies. RISE was tested on popular image generators and caught more real unsafe content slips than earlier methods could.
Open → 2609.34920v1