Preserve and compose training improves zero shot composed image retrieval

Preserve-and-Compose Training for Composed Image Retrieval

Computer Vision and Pattern Recognition

Summary

Finding images that match a user’s description of how to change a picture is hard, especially when we lack examples of the final images. The authors propose a new training method that uses not only text descriptions of the target images but also visual details from the original picture. This helps the system remember important parts of the source image while applying changes described in text. They also introduce a new way to score matches that balances following the text and keeping the original picture’s look. Their tests show better image search across different settings without needing new example images.

What this means in practice

  • For e-commerce platform developers: Enhance product image search by enabling queries that modify existing images using natural language without needing example images for training.$Commercial implications: This enables new interactive shopping experiences where customers can find products by describing modifications to existing images, improving relevance and engagement.
  • For digital asset managers: Improve retrieval of images that are variations or edits of a reference image using combined visual and textual guidance without costly data collection.

Authors

Sehyun Kwon

Abstract

Composed image retrieval (CIR) aims to retrieve images that satisfy a user-specified modification while preserving relevant visual content from a reference image. Collecting target images for this purpose is costly, motivating zero-shot CIR methods that instead use target captions as supervision. However, target captions may omit source details that should be preserved. We therefore propose, Preserve-and-Compose Training, which complements target-caption supervision with visual evidence from the source image. PACT learns from image--text--text (ITT) triplets without target images or gallery updates, aligning composed queries with target captions while preserving source evidence through visual supervision. We further introduce Chord scoring, which combines target similarity with source-relative directional agreement in the frozen image space. Results across four ZS-CIR benchmarks show that combining target-caption supervision with source-image evidence leads to strong retrieval performance across datasets, backbone scales, and external galleries. The code is available on https://github.com/sehyunkwon/PACT.