Papers for
online platform security teams
Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.
Sustained identity participation as a measurable security cost resource
Sustained Participation as a Security Resource: The Bounded Participation Channel
Abstract: Can sustained, per-identity participation be engineered into a security resource? Most anti-Sybil defenses price identity creation rather than identity survival. Once admitted, an adversary may sustain many identities without paying a recurring cost. We introduce the Bounded Participation Channel (BPC), a formal primitive for repeatedly verifying participation window by window. BPC issues fresh, identity-bound challenges under a strict deadline and enforces four structural properties: identity binding, freshness, real-time response, and bounded per-channel throughput. Together, these yield a provable cost theorem: sustaining $s$ identities over $T$ windows requires $C(s,T) \geq sT/τ_h$ participation channel-windows. The guarantee is solver-agnostic: a channel may be operated by a human, an AI system, or a hybrid. We give a hash-based construction with publicly verifiable participation proofs, characterize four admissible challenge families, and evaluate two against GPT-4o, Gemini 2.5 Flash, and Claude Sonnet 4.5 across 600 trials. Despite near-perfect accuracy (97--100%) on the perceptual tasks, the evaluated automated channels remain throughput-bounded under the tested deployment conditions. The results illustrate a key distinction: solvability does not imply unlimited throughput. By requiring participation to be re-earned by every identity in every time window, BPC turns sustained participation into a measurable security resource with a linear structural cost floor, independent of whether the participation is supplied by humans, AI systems, or hybrids.
Llms vulnerable to misinformation when memory is wiped between talks
Benchmarking Factual Robustness of LLMs via Multi-conversation Persuasion
Abstract: As Large Language Models (LLMs) increasingly serve as primary knowledge retrieval interfaces, their robustness against \textit{persuasion attacks}---attempts to inject misinformation or enforce counterfactuals---has become a critical safety concern. Existing red-teaming frameworks typically evaluate models in multi-turn dialogues where the target model retains full conversation history. We identify a critical flaw in this setting termed \textbf{``Refusal Inertia''}: a model's initial refusal often propagates through subsequent turns largely to maintain contextual consistency, thereby masking its true vulnerability to sophisticated, isolated persuasion attempts. To rigorously evaluate the ``cold-start'' defense capabilities of SOTA models, we introduce the \textbf{SAST-IR} (Stateful Attacker, Stateless Target - Iterative Refinement) framework. By enforcing a memory wipe on the target while retaining the attacker's history, we simulate a worst-case adversarial setting using \textbf{multi-turn} (stateless) iterations. Leveraging \textbf{CP-Agent} (Cognitive Persuasion Agent), an enhanced diagnosis-guided agent, our experiments on the custom \textsc{CounterFact-Strict} dataset ($N=50$) yield alarming results: simple, diverse attack strategies achieved a staggering \textbf{96\%} success rate, exposing severe brittleness in memory-less defense. Furthermore, we reveal a \textbf{``Complexity Paradox''}: while complex, iteratively refined attacks are effective, they often trigger defensive compliance, whereas simple strategies achieve a higher rate of genuine persuasion (\textbf{84.7\%}). Our code and dataset are available at GitHub, https://github.com/cza1006/llm-persuasion-defense.
Detecting coordinated online manipulation by analyzing data distortions
Principled Detection of Coordinated Manipulation from Aggregate Distortion and Account Reuse
Abstract: Coordinated manipulation is collective: plausible accounts can jointly distort ratings, rankings, and engagement. Existing defenses primarily construct evidence from identities, graphs, content, or co-activity. We introduce an aggregate-first evidence layer that treats distortion of a context-level outcome distribution as the primary evidence object. The engine observes only a histogram, count, resolution, and reference distribution; identities are withheld until interval evidence is fixed. Because raw discrepancies have positive finite-sample expectation, we subtract a matched null expectation to obtain signed evidence and account for reference uncertainty. Participation logs then accumulate these fixed increments across accounts. We characterize matched-exposure divergence, bound self-influence, establish finite-horizon separation, and derive an exact linear reuse law for paired contexts. We evaluate the mechanism with controlled rotation experiments and paired counterfactual interventions on historical Amazon review streams. Historical reviews provide the behavioral background; synthetic identities provide known coalition membership, and exact clean twins provide counterfactual controls. In a fixed-attack sweep against historical non-donor comparison accounts, reassigning the same manipulated events across identities with increasing reuse raises account-score ROC-AUC from 0.500 to 0.797. With activity- and exposure-matched clean twins, frequency is at chance while counterfactual attribution achieves ROC-AUC 0.744. Under a mean-preserving shape intervention, Wasserstein-1 and Jensen-Shannon evidence achieve ROC-AUC 0.909 and 0.967, while frequency and mean-based attribution remain at chance. Aggregate evidence complements repeated co-activity, improving mixed-mechanism ROC-AUC from 0.750 to 0.874 with a simple untrained combination.
Agentic group attack improves fake profile shilling on recommender systems
An Efficient and Effective Agentic Group Shilling Attack on Recommender Systems
Abstract: Recommender systems have become core infrastructure for modern online platforms, personalizing content at scale and strongly influencing what users see, click on, and purchase. However, this dependence on user interaction also exposes them to shilling attacks, where malicious actors can inject fake profiles to distort item rankings and control visibility. Existing attacks often rely on target-specific fine-tuning or fixed profile templates, making them either difficult to adapt to different victims or easier to detect. To overcome these limitations, we propose the Agentic Group Attack System (AGAS), a coordinated shilling framework where a central Coordinator directs a group of role-switching worker agents to adaptively promote a target item across different victim families. The Coordinator dynamically adjusts the strategy when progress stalls or suppression signals increase, while workers pursue a shared objective and switch between active and inactive roles to avoid repetitive patterns. Under the same attack budgets and evaluation protocols, AGAS consistently surpasses strong baselines in target promotion while better preserving benign recommendation quality, weakening representative detectors, and achieving higher efficiency than prior attacks. These findings also emphasize that defending recommender systems may require mechanisms that can handle adaptive shilling campaigns, not just isolated fake-profile injections. Our code is available at https://github.com/phkhanhtrinh23/AGAS.