Multimodal correction method improves attribute extraction from product data

Correcting to Predict: Pseudo-Value Correction for Multimodal Attribute Value Extraction

Information Retrieval

Summary

Extracting specific details like color or size from product descriptions and images is hard because some details are tricky and rely on understanding text and pictures together. The authors present a method called Correcting to Predict (C2P) that starts with a guess and then uses both text and images to fix that guess. This helps the system handle confusing or subtle cases better. Their approach works well on public and industry datasets and improved real online shopping experiences at AliExpress.

What this means in practice

  • For e-commerce platform engineers: Integrate C2P to improve automatic extraction of product features from combined text and images, especially for ambiguous attributes.
  • For online marketplace product teams: Deploy C2P-based attribute extraction to increase seller participation, enrich product listings, and boost user engagement.$Commercial implications: Provides a commercial tool for marketplaces to enhance product metadata quality, improving search and sales.

Authors

Junhao Zhang, Feiran Hu, Xiao Hu, Baoliang Cui, Xiaoyi Zeng

Abstract

Product attribute value extraction (AVE) is a fundamental task in e-commerce, aiming to identify specific values of predefined attributes from multimodal product profiles such as text and images. While multimodal large language models (MLLMs) have shown promise for AVE, they face challenges in extracting implicit attributes that require joint reasoning over visual and textual cues, often confusing semantically similar values. However, existing methods often fail to resolve such ambiguities because the correct value often depends on subtle multimodal cues that are easy to miss or override. To address this challenge, we propose Correcting to Predict (C2P), a framework that treats attribute extraction as a correction process. Given an initial pseudo-value such as a retrieved candidate or placeholder, the model learns to correct it using multimodal evidence. During training, diverse pseudo-values help the model learn evidence-based correction behavior, and a self-consistency refinement stage further reduces sensitivity to pseudo-value perturbations. At inference, a fixed placeholder triggers the learned correction behavior, enabling efficient single-pass prediction without online retrieval or iterative refinement. We evaluate C2P on a public benchmark and a large-scale industrial dataset. Offline results show that C2P outperforms strong baselines, with notable gains on ambiguous attributes. Online A/B tests on AliExpress further show consistent improvements in seller adoption, attribute completeness, and user engagement, validating C2P's effectiveness and efficiency in real-world deployment.