Prompt relevance drives AI citations more than website features

What Drives Citations in Production Large Language Models? An Observational Multi-Method Study of Two Million AI Citations Across Ten Thousand Web Pages

Artificial Intelligence

Summary

Large language models (LLMs) like ChatGPT often cite web pages when they answer questions, but it wasn’t clear which features make some pages cited more than others. The authors studied about 2 million citations from four AI engines linked to 10,000 web pages and found that how closely a page’s content matches the prompts given to the AI is the strongest predictor of citation frequency. Other factors like web design or performance measures had less clear effects. They also found that the reputation or authority of a domain strongly influences citations. The authors share their analysis process for others to use.

What this means in practice

  • For website developers: Improve the relevance of webpage content to match typical AI prompts to increase citations from commercial language models.
  • For digital marketers: Focus on building domain-level authority to significantly boost the likelihood that large language models cite your site.

Authors

Ben Moore, Liam Dunne

Abstract

Production large language models retrieve and cite web pages alongside generated answers, yet the page-level features that predict citation frequency remain poorly characterised. We present an observational study of approximately 2 million LLM citations from four commercial engines (ChatGPT, Claude, Google AI, Gemini) over six months, joined to 10,000 crawled pages from nineteen B2B SaaS workspaces. Sixty-plus features are tested using a nine-method consensus framework combining mixed-effects regression with domain fixed effects, FDR correction, stability-selection Lasso, double machine learning, generalised additive models, and temporal hold-out replication. Four findings survive all checks. First, prompt-content alignment (Jaccard overlap between page tokens and the full workspace prompt corpus, including non-citing prompts) is the dominant page-level predictor (beta = +0.37, 95% CI [+0.33, +0.41], q ~ 10^-73). Second, the standard AEO checklist (FAQ blocks, structured data, Core Web Vitals) shows positive effects in pooled data that reverse or collapse to zero once domain fixed effects are applied: Simpson's paradox with practical consequences for the AEO literature. Third, domain-level AI authority exceeds the strongest non-alignment page-level feature by a factor of six in mean absolute SHAP value. We release the analytic pipeline as a methodological contribution.