What Can I Edit? Open-Ended Strategy Discovery and the Emotion Editability Landscape

2026-07-27Computer Vision and Pattern Recognition

Computer Vision and Pattern Recognition
AI summary

The authors present EmoScope, a new system for editing images to express emotions in a way that fits each specific image rather than using fixed filters or templates. EmoScope first figures out what parts of an image can be changed to convey a target emotion, then balances keeping the image's content consistent while adding emotional expression. In tests with humans, EmoScope’s edits were preferred much more often than other methods, and it allows users to tweak the editing plan interactively. The authors also found that common emotion classifiers can miss subtle, context-aware edits, highlighting the importance of EmoScope’s approach.

emotional image editingaffordance reasoningmulti-agent frameworksemantic hierarchycontent consistencyemotion expressivenessMikels emotion categoriesclassifier-based metricsinteractive refinementemotion-conditioned strategies
Authors
Qing Li, Zeyu Dong, Yin Cui, Chuan Yan, Xiaojiang Peng
Abstract
Emotional image editing requires more than applying affective filters or modifying predefined visual factors: an effective edit must identify what a particular image can afford for a target emotion. Existing affective image manipulation methods, including recent agentic variants, largely operate within bounded strategy spaces based on predefined factor taxonomies, knowledge libraries, or conventional editing templates, and therefore often miss image-specific, context-grounded strategies. We introduce EmoScope, a multi-agent framework that reframes the task from "how should I edit?" to "what can I edit?" EmoScope first discovers an image-specific editable space through emotion-conditioned affordance reasoning, then uses a semantic hierarchy of anchors, variables, and context to balance content consistency and emotional expressiveness before executing and verifying the edit. Because its plans are expressed as image-specific affordances rather than retrieved templates, EmoScope also exposes the editing strategy as an interactive surface for user refinement at the plan level. In a large-scale human evaluation covering all eight Mikels emotion categories, with 4,693 valid responses across 1,824 pairwise questions, participants preferred EmoScope over two competitive baselines by 88.1% on average. Attribution analysis further shows that EmoScope selects target-emotion-adaptive strategies rather than applying a uniform template. The same affordance-level plan also supports lightweight user refinement in an interactive pilot. Finally, we show that classifier-based metrics exhibit emotion-conditional blind spots toward non-stereotypical, context-grounded edits, and present a relative content-emotion preference-affinity landscape showing that EmoScope's advantage varies systematically across image-emotion combinations.