Papers for
social media platform moderators
Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.
Model improves detection of toxic language in gamer chat messages
In-game Toxic Detection: Bi-directional Representations with Attention Residuals
Abstract: In-game toxic language has emerged as a critical concern in the gaming industry and community. While several frameworks and models for online game toxicity analysis have been proposed, detecting toxicity in player chat utterances remains a formidable challenge: stemming not only from the extremely short length of such utterances but also from the heavy reliance on game slang, abbreviations, and domain-specific jargon, which generic language models are poorly suited to recognize. This paper presents a shared task for in-game toxic language detection built upon real-world in-game chat data, and proposes the best-preforming model for the toxic language slot filling: Bi-directional Representations with Attention Residuals (BRAR). Experimental results demonstrate that BRAR effectively captures the global context and outperforms the existing baselines on slot filling.
Visual system identifies deceptive patterns in ai generated images
ASAP: Visual Analytics for Identifying and Analyzing Image Patterns in AI-generated Images
Abstract: Generative image models can produce highly realistic images, raising concerns about potential misuse in creating deceptive content. Current deepfake approaches face several challenges, including limited generalizability, lack of interpretability, and poor actionability. To help address these, we present ASAP, an interactive visualization system designed to empower users in the analysis and summarization of deceptive patterns in AI-generated images. ASAP introduces a novel CLIP-adapted image encoder that generates interpretable representations, enabling the extraction of influential pixel regions via calculated masks. This approach facilitates the identification of key deceptive features through influence measurement techniques. These backend techniques are integrated into a visual analytics dashboard that allows users to quantify and analyze authenticity-indicative patterns in image collections containing both authentic and AI-generated images. This approach also supports the comparative analysis of various generative models, including GANs and diffusion models. We demonstrate ASAP's efficacy through a user study and two application scenarios using established fake image detection benchmarks, showcasing its ability to effectively extract and quantify deceptive patterns.
Public suspicion of bots rises with Twitter controversies and account traits
When Does the Public Become Suspicious of Bots? Demand-Side Evidence from Botometer Query Logs
Abstract: We study private bot-checking behavior from the demand side: when people suspect an account is automated, whom they suspect, and what follows. Using Botometer's server-side query logs, the most widely used bot-detection service, we treat each query as a behavioural trace of suspicion. We analyze over 1 million public checks of Twitter accounts from 2020 to 2023, enriched with 3.2 billion tweets from the contemporaneous 1% public stream. Collective suspicion spikes with platform crises, most sharply around the 2022 Musk-Twitter bot dispute. Checked accounts are older and more prolific, have more followers, and post promotional, political, and crypto content. Accounts that draw collective suspicion have higher bot scores and are more likely to be suspended. Bot-related public attention and Botometer activity are elevated during the same broad periods, although their short-run fluctuations are largely uncoupled. Public feedback focuses on first-person identity claims for humans, and evidence-based arguments citing posting rate, political content, and cross-account coordination for bots and cyborgs. Bot suspicion thus constitutes a mass, distributed form of platform auditing that tracks meaningful signals of automation, establishing audit-tool query logs as a novel lens on public responses to platform manipulation.