Multi-agent ai creates editable product reviews from user images

From Visual Feedback to Textual Reviews: A Multi-Agent Vision-Language Framework for Image-Grounded Review Assistance

Computer Vision and Pattern RecognitionArtificial Intelligence

Summary

Shoppers often share pictures of products to show quality or defects, but don't always write detailed reviews. The authors introduce a system that looks at these photos and helps write draft reviews based on what it sees. Their approach breaks the task into steps, including recognizing parts of the product, guessing opinions, and putting together evidence for a review. Tests on real product photos show it can make sensible, opinion-aware reviews to help people write better feedback.

What this means in practice

  • For e-commerce platform developers: Generate editable review drafts automatically from customer product images to assist shoppers in providing richer feedback.$Commercial implications: This paper enables AI tools that help e-commerce sites boost quality and quantity of user reviews by translating images into text.
  • For customer service teams: Use AI-generated reviews from user images to understand product issues more quickly and improve response strategies.

Tested on one dataset.

Authors

Utsav Kumar Nareti, Ayush Bansal, Kumari Priya, Chandranath Adak, Soumi Chattopadhyay, Muhammad Saqib, Saeed Anwar

Abstract

Visual feedback in the form of user-uploaded images and videos is becoming increasingly common in e-commerce platforms because it provides authentic evidence of product quality, defects, packaging conditions, and real-world usage. However, visual feedback alone often lacks the contextual explanations and subjective opinions necessary for informed decision-making, while many users provide limited textual feedback due to the effort required to compose detailed reviews. To bridge this gap, we introduce image-grounded review assistance, a novel task that aims to generate editable review drafts from user-uploaded product images. Unlike conventional image captioning, which focuses on objective visual description, the proposed task requires product-specific understanding, sentiment estimation, and evidence-driven review composition under challenging real-world conditions, including degraded image quality, excessive zoom-in, target ambiguity, and partial product visibility. We propose a multi-agent vision-language framework consisting of four specialised roles: product grounding, visual sentiment estimation, visual evidence generation, and review synthesis. The framework employs explicit intermediate representations, including product entities, predicted ratings, and evidence summaries, to improve interpretability and visual grounding. Experiments on a curated subset of the Amazon Reviews Electronics dataset demonstrate the feasibility of generating coherent, product-aware, and sentiment-aware review drafts from visual feedback. To the best of our knowledge, this is the first study to formulate image-grounded review assistance as a multi-agent vision-language reasoning problem, providing a practical step toward AI-assisted review authoring in e-commerce systems.