Papers for

digital marketers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Prompt relevance drives AI citations more than website features

What Drives Citations in Production Large Language Models? An Observational Multi-Method Study of Two Million AI Citations Across Ten Thousand Web Pages

Abstract: Production large language models retrieve and cite web pages alongside generated answers, yet the page-level features that predict citation frequency remain poorly characterised. We present an observational study of approximately 2 million LLM citations from four commercial engines (ChatGPT, Claude, Google AI, Gemini) over six months, joined to 10,000 crawled pages from nineteen B2B SaaS workspaces. Sixty-plus features are tested using a nine-method consensus framework combining mixed-effects regression with domain fixed effects, FDR correction, stability-selection Lasso, double machine learning, generalised additive models, and temporal hold-out replication. Four findings survive all checks. First, prompt-content alignment (Jaccard overlap between page tokens and the full workspace prompt corpus, including non-citing prompts) is the dominant page-level predictor (beta = +0.37, 95% CI [+0.33, +0.41], q ~ 10^-73). Second, the standard AEO checklist (FAQ blocks, structured data, Core Web Vitals) shows positive effects in pooled data that reverse or collapse to zero once domain fixed effects are applied: Simpson's paradox with practical consequences for the AEO literature. Third, domain-level AI authority exceeds the strongest non-alignment page-level feature by a factor of six in mean absolute SHAP value. We release the analytic pipeline as a methodological contribution.

Mon 28 SeptArtificial Intelligence
The gist
Large language models (LLMs) like ChatGPT often cite web pages when they answer questions, but it wasn’t clear which features make some pages cited more than others. The authors studied about 2 million citations from four AI engines linked to 10,000 web pages and found that how closely a page’s content matches the prompts given to the AI is the strongest predictor of citation frequency. Other factors like web design or performance measures had less clear effects. They also found that the reputation or authority of a domain strongly influences citations. The authors share their analysis process for others to use.
Open → 2609.35077v1

Brand presence affects mention likelihood in AI search outputs

From Prompt to Recommendation: A Fitted Stage Model of Brand Visibility in AI Search

Abstract: We analyze 34,960 unbranded prompt-engine observations from 75 anonymized Aiso projects, covering 2,854 distinct monitored prompts and repeated GPT and Gemini runs from June-September 2026. When neither the target brand nor its own domain appears in the observable live retrieval path, target mention rates are 2.8% for GPT and 3.8% for Gemini. With an own-domain citation but no branded fan-out, they rise to 49.0% and 58.4%. When both own-domain exposure and a branded fan-out occur, mention rates reach 91.4% and 100%. The relationship persists within the same project, prompt, and engine across repeated runs: among prompt cells that vary in own-domain exposure while holding branded fan-out absent, exposure is associated with a mean mention-rate increase of 40.2 percentage points on GPT and 49.0 points on Gemini. Prior visibility is independently persistent. A previous non-mention plus no current own-domain exposure yields next-run mention rates of 1.6% and 1.9%; previous mention plus current exposure yields 80.5% and 83.7%. We fit a chronological diagnostic model using prior-run history and contemporaneous retrieval indicators: $ \operatorname{logit}P(M_t=1)=α_e+β_e\operatorname{logit}(\widetilde P_{t-1})+γ_e E_t+δ_e F_t+θ_e^\top X. $ On the latest 30% holdout, the full model achieves AUC 0.963 on GPT and 0.942 on Gemini, compared with 0.937/0.917 for prior history alone and 0.880/0.840 for live signals alone. A manually curated prompt sensitivity gives nearly identical AUCs (0.960 and 0.943). A separate 199-prompt page-corpus validation finds that prompt-page match predicts Gemini exposure (AUC 0.641) more clearly than GPT exposure (0.545), placing relevance upstream of a larger engine-mediated exposure effect. The equation is predictive and observational, not a causal description of proprietary engine internals.

Sat 19 SeptInformation Retrieval
The gist
When AI search engines like GPT or Gemini respond to prompts, whether a brand's own website or domain appears in the results greatly affects how likely that brand is mentioned. The authors found that without the brand’s domain visible, mentions are very low, but with it plus other signals, mentions become very likely. They created a mathematical model that predicts brand mentions based on previous mentions and current exposure signals, which performs well in tests. This model helps understand how brand visibility influences AI-generated recommendations.
Open → 2609.23162v1

Thompson sampling method improves bandit decisions with changing baselines

Odds-Ratio Thompson Sampling: A Specification and Design Guide for Contrast-Based Multi-Armed Bandits

Abstract: Batched multi-armed bandits update on a service's own schedule, and the usual implementation carries each arm's absolute reward rate from one update to the next. When the shared level moves between batches, that memory goes stale even though the comparisons between arms may not have. Odds-Ratio Thompson Sampling (OR-TS) instead carries the joint posterior over log-odds contrasts and fits the common level afresh in every batch, marginalizing it out. This paper specifies that update, places it inside a Bayesian bandit agent with two controls, decay for how much past evidence survives an update and aggressiveness for how sharply belief becomes allocation, and evaluates it against absolute-rate memory. Across 86 public A/B series the level varies about twenty-five times more than the contrast. In prespecified synthetic environments a moving level costs absolute-rate memory five times the regret and leaves the best arm below a majority of traffic in 7 of 20 runs, against none for OR-TS. In a policy simulation built from 71 real experiments, where the contrasts are too small to resolve, expected-click differences stay within 0.1% for 58 of them, yet contrast memory still ends on the better arm more than twice as often. Where the contrasts themselves move, the bet fails, and that case is reported too.

Thu 17 SeptMachine Learning
The gist
This paper looks at how computers decide between multiple options when the overall success rate changes over time. The usual methods remember each option’s absolute performance, which can get outdated if conditions shift. The authors propose a new method called Odds-Ratio Thompson Sampling that focuses on comparing differences between options instead, refreshing the shared baseline regularly. Their tests show this new method makes better decisions when conditions vary and avoids misleading memory problems. However, when the differences between options themselves change a lot, this method does not help.
Open → 2609.19709v1