Chatbots vary widely in product advice and sources given

"If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations

Computers and SocietyComputation and Language

Summary

People often ask AI chatbots for shopping advice, but these chatbots give very different answers and mention different websites as sources. The authors studied popular chatbots like ChatGPT and Google Gemini by asking them real questions and found that only ChatGPT often shared its own preferences. The chatbots’ recommendations and source websites changed a lot even when the same question was repeated. This means the shopping advice people get from these AIs can be biased and inconsistent depending on which chatbot or platform they use.

What this means in practice

  • For consumer protection teams: Audit chatbot product advice to detect bias and inconsistencies consumers might face across different AI platforms.
  • For ai platform developers: Design chatbots that provide more stable and transparent product recommendations by tracking multiple responses and clearly showing source origins.

Authors

Lucas G. Uberti-Bona Marin, Thales Bertaglia, Giovanni Astante, Bram Rijsbosch, Gijs van Dijck, Anikó Hannák, Gerasimos Spanakis, Konrad Kollnig

Abstract

Consumers increasingly use AI chatbots for advice on what to buy. With companies like OpenAI and Google monetising their AI through advertising, this raises difficult questions about the bias and impartiality of such advice. In response, we conduct an AI audit of popular chatbots using real commercial-advice queries. First, we curate a dataset of 2,528 real commercial-advice queries (ConsumerQ). Then, we evaluate 1,536 responses to product queries from popular AI chatbots: ChatGPT (chatbot and API), Google Gemini (chatbot and API), and Google Search (AI Overviews). We find that ChatGPT expresses a first-person product preference in 79% of product-recommending responses, compared with 7% for Gemini and 2% for AI Overviews, while the products recommended often change across repeated requests. Displayed sources vary strongly: for the same query, the ChatGPT and Gemini interfaces share only 5.4% of domains on average, with no domain in common in 76.7% of comparisons. APIs provide a different view from their corresponding interfaces, with mean domain overlaps of 12.0% for ChatGPT and 14.8% for Gemini, and also differ in the types and layers of source information they expose. Our findings show that neither isolated responses nor API observations can be assumed to represent the commercial advice consumers encounter. Independent audits of AI-mediated commercial advice should therefore account for repeated responses, consumer-facing conditions, and the source layer being observed.