Papers for

e-commerce platform developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Session recommendation gets faster and better without attention models

No Attention, No Problem: Rethinking Session-based Recommendation with Pure Convolution

Abstract: Session-based recommendation (SBR) predicts the next choice in a session by analyzing recent interactions. Transformer-based models are widely used because of their ability to capture long-range dependencies through self-attention mechanisms. In contrast, traditional convolutional models, although more efficient, are often limited by their weak global modeling capabilities and are losing ground in SBR tasks. In this work, we propose a Next-generation Pure Convolutional Framework (NextConvRec) for SBR tasks, aiming to balance efficiency and performance. NextConvRec uses a Structural and Positional Convolutional Encoder (SPCE) for preprocessing, combining learnable convolutional positional biases with session-level structural signals extracted through GCN layers. Its backbone convolutional module effectively expands the effective receptive field through depthwise convolutions and pointwise convolutions, enabling robust long-range preference modeling without attention mechanisms. Extensive experiments on 4 benchmark datasets show that NextConvRec outperforms several state-of-the-art baselines by around 1.73% on average, and reduces the average inference time per session by 16.7%. The convolutional architectures remain a promising direction for efficient and accurate session-based recommendations.

Mon 28 SeptInformation Retrieval
The gist
Session-based recommendation systems try to predict what you will choose next based on your recent actions. Usually, complex models that use attention mechanisms help these systems understand long-term connections but are slower. The authors show that a new convolution-based method can do this faster while keeping or even improving accuracy. This method combines positional signals and structural information without using attention, making it more efficient for real-time predictions.
Open → 2609.34802v1

Recommendation models improve generalization beyond one epoch training

Beyond One Epoch: Uncertainty-Weighted Sensitivity Regularization for Recommendation Models

Abstract: Recommendation models with sparse embeddings and a shared consumer often exhibit the one-epoch phenomenon: a second epoch lowers training loss while sharply degrading generalization. We present a view based on the violation of the prequential principle. On the first epoch, an example's label has not affected the embedding rows used to score it. On later epochs, those rows contain a displacement induced by the labels earlier update. This creates an incentive for the shared consumer to exploit this displacement in subsequent epochs, which fails to generalize. We call this self-influence asymmetry. Using an exact scalar model and local influence analysis, we connect this mismatch to the uncertainty in the embeddings and the consumers incentive to exploit it in subsequent epochs. We verify this hypothesis using an embedding-consumer-update interventions in deep recommendation models and propose uncertainty-weighted sensitivity regularization (UWSR) which counteracts this mismatch by augmenting the loss function to penalize the consumer for relying on uncertain embeddings. Unlike existing remedies, UWSR preserves the learned embeddings and across three benchmarks, four-epoch UWSR reduces test cross-entropy by 1.38%-6.78% and improves AUC by 0.0058-0.0231 relative to one-epoch training.

Mon 28 SeptMachine Learning
The gist
Recommendation systems often train quickly in one pass but get worse when trained longer, a problem called the one-epoch phenomenon. The paper explains this happens because the system unintentionally uses information from earlier training in a way that does not work well for new data. The authors show that this issue is linked to uncertainty in the model’s internal parts and propose a new training method called uncertainty-weighted sensitivity regularization (UWSR). UWSR helps the model avoid relying too much on uncertain parts, improving recommendation quality over multiple training passes.
Open → 2609.34083v1

Preserve and compose training improves zero shot composed image retrieval

Preserve-and-Compose Training for Composed Image Retrieval

Abstract: Composed image retrieval (CIR) aims to retrieve images that satisfy a user-specified modification while preserving relevant visual content from a reference image. Collecting target images for this purpose is costly, motivating zero-shot CIR methods that instead use target captions as supervision. However, target captions may omit source details that should be preserved. We therefore propose, Preserve-and-Compose Training, which complements target-caption supervision with visual evidence from the source image. PACT learns from image--text--text (ITT) triplets without target images or gallery updates, aligning composed queries with target captions while preserving source evidence through visual supervision. We further introduce Chord scoring, which combines target similarity with source-relative directional agreement in the frozen image space. Results across four ZS-CIR benchmarks show that combining target-caption supervision with source-image evidence leads to strong retrieval performance across datasets, backbone scales, and external galleries. The code is available on https://github.com/sehyunkwon/PACT.

Fri 25 SeptComputer Vision and Pattern Recognition
The gist
Finding images that match a user’s description of how to change a picture is hard, especially when we lack examples of the final images. The authors propose a new training method that uses not only text descriptions of the target images but also visual details from the original picture. This helps the system remember important parts of the source image while applying changes described in text. They also introduce a new way to score matches that balances following the text and keeping the original picture’s look. Their tests show better image search across different settings without needing new example images.
Open → 2609.31202v1

Privacy and stability improve slate recommendation with noisy scores

Decoupled Learning and Selection in Slate Recommendation for Privacy and Stability Under Noisy Scores

Abstract: We formalize slate recommendation as a randomized score learner followed by deterministic selection. First, an appropriately scoped differential-privacy guarantee passes through selection and its audit trace by post-processing. End-to-end privacy holds only when selector inputs are public or independent, previous private outputs, or separately privacy-accounted; fixing raw state or candidate information instead yields only a conditional guarantee. Second, we derive a logged margin certificate: bounded score-induced objective movement below half the smallest greedy decision margin guarantees that the ordered slate is unchanged. Controlled fixed-margin tests show near-linear exponent scaling, with an empirical slope of $-0.220$ (95% CI $[-0.231,-0.210]$) against the independent-noise reference $-1/4$. Real-anchor experiments on OULAD, MovieLens-25M, and Amazon Musical Instruments show that greater anchor weight reduces score-noise-induced ranking churn. OULAD and EdNet certificate checks validate the implementation of the logged inequality, while closed-loop simulations show bounded target drift and setting-dependent downstream utility. The contribution is therefore a privacy-scope contract and a certifiable score-to-slate stability mechanism, not a universal utility claim.

Thu 24 SeptMachine LearningInformation Retrieval
The gist
Slate recommendation means picking a list of items to show to users, like movies or products, based on scores given by a model. The authors explore how privacy rules apply when the system both learns those scores privately and then selects items based on them. They also develop a way to check if small random changes in scores will change the chosen list, helping guarantee stability. Their tests on real datasets show their method reduces ranking changes caused by noise without hurting performance too much.
Open → 2609.29453v1

Verifiable agentic environments improve long horizon reinforcement learning

Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms

Abstract: Language-model agents increasingly face long-horizon tasks with evolving state, interdependent decisions, and delayed outcomes. Scaling their training requires diverse agentic environments, dependable outcome signals, and low extension cost. Existing generation pipelines commonly construct an environment before defining its outcome rule or annotating its trajectories, leaving dynamics and evaluation to be aligned post hoc. VHD-Play reverses this dependency by sampling and solving a mathematical model before a corpus-grounded setter renders its decision process as stateful tools. The executable dynamics and trajectory-scoring reference are inherited from the same solved model. The pipeline produces 3,300 diverse agentic environments at a cost of a few cents each. Training Qwen3.6-35B-A3B on three families raises its mean agentic score from 0.204 to 0.815 in a five-family diagnostic. Gains also appear on held-out instances from all three training families and eight unseen mechanism families, then extend beyond the generated substrate to external benchmarks for general function calling, travel planning, and 365-day e-commerce. On E-Commerce Bench, the trained checkpoint completes every run without bankruptcy and exceeds Qwen3.7-Max. We compare written-out problems with stateful versions that reveal or hide their parameters. The comparison shows that most of the learnable gap lies in stateful interaction rather than underlying problem solving. A frozen 35B setter realizes larger environments, and scale-matched training retains gains as mechanism size and horizon grow, indicating the potential for an evolving training substrate.

Wed 23 SeptArtificial IntelligenceComputation and Language
The gist
Training language-model agents to handle long, complex tasks is hard because you need many varied environments and clear signals about success. The authors developed VHD-Play, which creates environments by first solving mathematical models, then turning them into interactive setups agents can use. This method produces thousands of diverse and verifiable environments cheaply and helps train agents that perform much better on a range of tasks, including e-commerce and travel planning challenges. The results show that learning to interact with dynamic states is more important than just solving static problems.
Open → 2609.27321v1

Semantic user profiling methods compared for streaming recommendations

When LLM-Based User Profiling Adds Value in Production Streaming Recommendation

Abstract: Personalized recommendation depends critically on how user representations are constructed from historical behavior. Two paradigms have emerged for constructing semantic user profiles in content-based recommendation. First, aggregate methods derive user representations as numerical aggregates of semantic item embeddings. Second, LLM-based methods generate natural-language summaries of user preferences and encode them through a text encoder. Each paradigm can be combined with temporal disentanglement of recent versus historical behavior. LLM-based profile generation is significantly more expensive than aggregate approaches, raising the question of when this additional cost is justified. We present a systematic comparison of four semantic user-profiling strategies, factorially crossed across representation type and temporal handling, evaluated on a real-world production dataset. The comparison reveals how these strategies differ across user behavior types, across both accuracy and beyond-accuracy dimensions of recommendation quality, and across the temporal-window setting that governs the disentanglement.

Wed 23 SeptInformation Retrieval
The gist
Online services often suggest content based on what users liked before, but creating user profiles from past actions varies in complexity and cost. The authors compare two main ways to build these profiles: one that averages numerical representations of items, and another that uses large language models (LLMs) to write natural descriptions of user interests. They tested these methods on a real streaming service dataset, looking at how well each works depending on the user’s behavior and the timing of actions. Their findings help decide when the expensive LLM-based profiling is worth using compared to simpler methods.
Open → 2609.27183v1

System improves trustworthiness of images in entity matching

Knowing When to Trust Images: Reliability-Aware Multi-modal Entity Alignment

Abstract: The visual modality, i.e., images, plays a key role in multi-modal entity alignment (MMEA). Existing approaches often directly fuse the image with other modalities to align different entities. Although simple, such strategies overlook the potential noise in the images and their semantic misalignment with corresponding entities, resulting in suboptimal fusion and degraded performance. Addressing this, we propose a novel Reliability-Aware framework for MMEA (RA-MMEA), which assesses visual reliability and adaptively improves unreliable visual representations for robust entity alignment. The core lies in two modules, including dependency-aware visual reliability prediction (DA-VRP) and stability-regularized visual embedding generation (SR-VEG). The former aims to estimate the reliability of an image by leveraging multi-modal dependency within the entity, while the latter focuses on producing alternative visual representation conditioned on semantics encoded in textual modalities for multi-modal fusion. Compared to current methods, RA-MMEA enables more reliable visual representations for modality fusion, thereby improving performance. In extensive experiments, RA-MMEA achieves state-of-the-art results, verifying the importance of reliable visual modality for entity alignment and the effectiveness of RA-MMEA. The code and results will be released.

Sun 20 SeptComputation and LanguageArtificial IntelligenceComputer Vision and Pattern Recognition
The gist
Matching data about the same thing from different sources is hard when images are not reliable or don’t fit well with text. The authors developed a method that checks how trustworthy images are and adjusts their representations to better match text descriptions. This approach helps combine images and words more effectively, making it easier to correctly identify the same entities across sources. Their tests showed that considering image reliability improves matching accuracy.
Open → 2609.23267v1

Livestream videos get smarter at matching products with show moments

Grounded Product Understanding in Livestream Videos

Abstract: E-commerce livestreams have emerged as an important channel for presenting products to online consumers, containing multiple products whose information is scattered in different moments. This poses significant challenges for downstream product understanding applications, such as product-centric livestream clipping, where models need to identify the product and its relevant segments for information gathering. However, existing benchmarks for general product understanding typically evaluate product retrieval and temporal localization in isolation, leaving the critical correspondence between product identity and temporal evidence largely unassessed. To address this limitation, we introduce GPUB, a large-scale benchmark comprising 3,000 livestream instances with quality-controlled multi-moment temporal annotations and a catalog of over 31K fashion products. GPUB supports three evaluation tasks: the main task Grounded Product Understanding (GPrU) requires jointly identifying the target product and localizing its supporting moments from a livestream video and a candidate product set; Product Retrieval and Product Moment Localization serve as two complementary subtasks. Evaluation of existing multimodal models shows that GPrU remains highly challenging, with the best-performing baseline achieving only 10.13% Pair mAP@.3. To narrow the performance gap, we further develop UniPro, a unified product understanding model that derives product-aligned and temporally structured representations from shared multimodal encoding, improving Pair mAP@.3 to 21.53% while achieving 37.23% Joint R@1@.3 on GPrU.

Thu 17 SeptComputer Vision and Pattern Recognition
The gist
Online shopping livestreams often show many products at different times, making it hard to tell which product relates to which video moment. The authors created GPUB, a big new dataset that connects fashion products to specific moments in livestreams to help computers learn this connection. They also built a new model called UniPro that does better at identifying products and the exact video parts that show them. Despite improvement, the task remains difficult for current technology.
Open → 2609.20508v1

Efficient engine speeds up linking data for multi-step reasoning

Efficiently Linking Unstructured Data for Multi-step Reasoning

Abstract: Modern LLMs and AI agents increasingly support data engineering workflows that integrate evidence from unstructured sources. Such pipelines typically do data retrieval, integration, and ranking before proceeding to more complex agentic reasoning or actions, e.g., for scientific discovery. The core retrieval problem in these workflows jointly executes multi-attribute filtering, multi-vector search, exact relational joins, and thresholded embedding-similarity joins. Given a planned query and monotone scoring function, our DASE query engine constructs and ranks candidate evidence tuples. It comprises (i) a multi-step reasoning query model over structured predicates, multiple vectors, and relational links; (ii) SemJI, a sparse materialized embedding-similarity join index for rare near-neighbor pairs; and (iii) a co-designed execution layer that combines predicate-aware ANN traversal, batched access, and threshold-based score aggregation. On scientific-discovery workloads, DASE retrieves candidate evidence for multi-step reasoning queries 6x to 46x faster than strong RDBMS, rerank, and vector-database baselines at comparable recall; and for tasks that require semantic-operator post-processing, DASE acts as a high-recall prefilter that makes downstream LLM evaluation both cheaper and more accurate -- e.g., on SemBench E-Commerce it improves BigQuery quality from 0.67 to 0.80 while cutting cost from $2.42 to $0.54.

Wed 16 SeptDatabasesArtificial Intelligence
The gist
Working with unstructured data like text and images is hard when you want to combine multiple pieces of information for complex reasoning. The authors developed a system called DASE that quickly finds and ranks relevant pieces of evidence from scattered data by mixing different search techniques. This helps AI tools make better, faster decisions when using lots of messy data, like in scientific research or e-commerce. Their experiments showed DASE is much faster than popular databases and improves the quality and cost of AI evaluations.
Open → 2609.19491v1

Region aware retrieval improves image and text matching in large models

RegRet: Enhancing Region-Level Retrieval in Large Multimodal Models

Abstract: Region-level retrieval aims to align user-specified image regions with relevant regions or textual descriptions, playing a crucial role in realworld applications such as e-commerce product search and RAG. Although recent Large Multimodal Models (LMMs) have made significant strides in multimodal retrieval, they primarily focus on global-level tasks and struggle to capture effective region-level representations. To bridge this gap, we present RegRet, an LMM-based Region-level Retrieval framework that enhances the regional representations without compromising overall global retrieval performance. At its core, RegRet integrates a Region-Aware Encoder to capture detailed regional features while balancing them with the global background context. To further enhance the fine-grained understanding and discriminability of representations, we design a multi-stage training pipeline that includes detailed localized captioning and regional contrastive learning tasks. In addition, considering the absence of region-level contrastive training data and the limited diversity of evaluation tasks in current benchmarks, we introduce the REGMB benchmark. It comprises 225k contrastive pairs, covering four multimodal retrieval tasks. Extensive experiments validate the effectiveness of our approach. RegRet outperforms strong baselines in the zero-shot setting. Further training with contrastive learning leads to an average improvement of more than 20\% on both REGMB and public benchmarks, while achieving comparable or better results on global-level retrieval tasks.

Tue 15 SeptComputer Vision and Pattern RecognitionArtificial IntelligenceInformation Retrieval
The gist
Finding specific parts of an image and matching them to related words or other image parts is important but hard. The authors designed RegRet, which helps large multimodal models look more closely at image regions without losing the overall picture. They also created a big new dataset to train and test this regional matching. Their experiments show RegRet does a better job than previous methods, especially when it comes to detailed image parts.
Open → 2609.16847v1

Agent system diagnoses and improves recommendation algorithms at scale

AURA: Agentic Diagnosis and Refinement for Production Recommender Systems at Scale

Abstract: How and why does a recommender system fail the users it serves? Oftentimes, practitioners are left to improve their algorithms based on a combination of feedback from stakeholder teams, domain expertise, and insights from data analyses. Yet the nuances of how and where recommendations perform well or poorly for end users are difficult to discern from aggregate quantitative metrics. Whereas these metrics provide a high-level and incomplete picture, further granularity into the quality of recommendations and their patterns requires reasoning with domain understanding and objectivity, at scale. We contemplate this complex conundrum and describe a method and implementation that uses the latest AI agentic advances to provide actionable diagnoses and improvements for production recommender systems. We present AURA (Agentic Understanding and Refinement of recommender Algorithms), an end-to-end agentic system that performs qualitative evaluation at scale and can then generate improvements to our algorithms at the code level. Specialized agents read production engagement logs, from thousands of sessions to millions, and surface patterns and examples of how the recommender fails real users. The next step uses those diagnoses and context about the recommender's own code, data, and training pipeline to propose and implement refinements grounded in that codebase. We report the system design, initial tests on production data from two large consumer platforms at a major media-streaming company, safeguards, operational learnings, and early results toward a self-improving recommender system. Finally, the diagnostic gap AURA closes is not specific to streaming. The architecture is built to transfer: every domain-specific element enters through the configuration layer that already ported it between our two platforms. We map it concretely to e-commerce and online-retail recommendation.

Tue 15 SeptInformation RetrievalArtificial IntelligenceMachine Learning
The gist
Recommender systems suggest things like movies or products but sometimes fail to give good suggestions. The authors created AURA, an AI tool that looks deeply at how and where these systems make mistakes by checking large amounts of real user data. AURA can then suggest and even change the recommendation algorithms to fix problems in the code. This approach was tested on big real platforms and can be adapted for other areas like online shopping.
Open → 2609.16625v1

Multi-agent ai creates editable product reviews from user images

From Visual Feedback to Textual Reviews: A Multi-Agent Vision-Language Framework for Image-Grounded Review Assistance

Abstract: Visual feedback in the form of user-uploaded images and videos is becoming increasingly common in e-commerce platforms because it provides authentic evidence of product quality, defects, packaging conditions, and real-world usage. However, visual feedback alone often lacks the contextual explanations and subjective opinions necessary for informed decision-making, while many users provide limited textual feedback due to the effort required to compose detailed reviews. To bridge this gap, we introduce image-grounded review assistance, a novel task that aims to generate editable review drafts from user-uploaded product images. Unlike conventional image captioning, which focuses on objective visual description, the proposed task requires product-specific understanding, sentiment estimation, and evidence-driven review composition under challenging real-world conditions, including degraded image quality, excessive zoom-in, target ambiguity, and partial product visibility. We propose a multi-agent vision-language framework consisting of four specialised roles: product grounding, visual sentiment estimation, visual evidence generation, and review synthesis. The framework employs explicit intermediate representations, including product entities, predicted ratings, and evidence summaries, to improve interpretability and visual grounding. Experiments on a curated subset of the Amazon Reviews Electronics dataset demonstrate the feasibility of generating coherent, product-aware, and sentiment-aware review drafts from visual feedback. To the best of our knowledge, this is the first study to formulate image-grounded review assistance as a multi-agent vision-language reasoning problem, providing a practical step toward AI-assisted review authoring in e-commerce systems.

Sun 13 SeptComputer Vision and Pattern RecognitionArtificial Intelligence
The gist
Shoppers often share pictures of products to show quality or defects, but don't always write detailed reviews. The authors introduce a system that looks at these photos and helps write draft reviews based on what it sees. Their approach breaks the task into steps, including recognizing parts of the product, guessing opinions, and putting together evidence for a review. Tests on real product photos show it can make sensible, opinion-aware reviews to help people write better feedback.
Open → 2609.14761v1

ShopEase system improves customer support with dense retrieval AI

ShopEase: A Generative AI-Based Multi-Agent Framework for Intelligent Enterprise Customer Support Using Hybrid Retrieval-Augmented Generation

Abstract: Enterprise customer support systems must answer customer questions correctly, retrieve the right policy information, use customer context, and pass difficult cases to human agents when needed. This paper presents ShopEase, a Generative AI-based multi-agent framework for enterprise customer support. The system combines six components: Intent, CRM, Memory, Hybrid RAG, Escalation, and Supervisor, and uses LLaMA 3.2 running locally through Ollama for response generation. The retrieval module combines FAISS (dense retrieval) and BM25 (sparse retrieval), and six configurations are evaluated: BM25-only, FAISS-only, Fair RRF, Weighted RRF, RRF with Cross-Encoder, and Top-10 Hybrid with Cross-Encoder. Instead of using a fixed mapping between intent and policy, the policy category is decided directly from the retrieved documents. The system was evaluated on 2632 held-out customer queries across six categories: Refund, Return, Shipping, Cancellation, Damaged Product, and Unknown. FAISS-only achieved the highest accuracy of 85.37\% (2247 correct predictions), closely followed by Weighted RRF at 85.07\%. BM25-only achieved only 55.74\% accuracy. Adding cross-encoder reranking did not improve results: RRF with Cross-Encoder reached 83.24\%, and Top-10 Hybrid with Cross-Encoder reached 81.88\%, while also increasing response latency. Category-level analysis shows strong performance on Shipping, Cancellation, and Return, while Unknown queries remain the main source of errors. Statistical testing using McNemar's test shows no significant difference between FAISS-only and Weighted RRF, though both perform significantly better than Fair RRF and the cross-encoder configurations. Overall, dense retrieval gives the best accuracy on this dataset, and additional reranking adds processing time without improving classification performance.

Sat 12 SeptComputation and Language
The gist
Enterprise customer support systems need to answer questions accurately and know when to ask a human for help. The authors created ShopEase, which uses multiple AI tools working together to handle customer questions better. Their system finds the right policies using a combination of dense and sparse search techniques and decides responses based on retrieved documents. Tests showed that dense search techniques worked best for understanding customer issues, and adding extra steps slowed things down without improving results.
Open → 2609.13856v1

Generative transformer improves relighting for 3D object images

RelightFormer: Feed-forward Generative Transformer for Multiview Object Relighting

Abstract: Image relighting is traditionally tackled via complex inverse rendering pipelines, which suffer from ill-posed optimization, or single-image generative models that ignore crucial multi-view cues necessary for understanding 3D geometry and material interactions. To address these limitations, we introduce a feed-forward generative Transformer for direct single- and multi-view image relighting that entirely bypasses explicit intrinsic property estimation. Adapted from a video foundation model, our architecture features a latent illumination module that dynamically injects target environment maps into spatial features via cross-attention. Furthermore, we employ permutation-invariant positional encodings to symmetrically process unordered multi-view inputs without sequential bias. To train this robust data-driven model, we construct the massive Laval Objaverse Dataset (LOD), comprising 90K objects and 39K unique illuminations. Extensive experiments demonstrate state-of-the-art visual quality, photorealistic relighting quality, and strong zero-shot generalization across single-view, multi-view, and novel-view relighting tasks.

Mon 7 SeptComputer Vision and Pattern RecognitionArtificial IntelligenceGraphics
The gist
Adjusting the lighting in photos of objects from different viewpoints is usually tricky and slow. The authors created a new AI model that changes lighting directly on images without needing to figure out complicated 3D details first. They trained this model using a huge dataset of thousands of objects and lighting setups. Their approach works well on images from one or many angles and even for unseen views, producing realistic relighting.
Open → 2609.07414v1