Papers for

social media platform moderators

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Model improves detection of toxic language in gamer chat messages

In-game Toxic Detection: Bi-directional Representations with Attention Residuals

Abstract: In-game toxic language has emerged as a critical concern in the gaming industry and community. While several frameworks and models for online game toxicity analysis have been proposed, detecting toxicity in player chat utterances remains a formidable challenge: stemming not only from the extremely short length of such utterances but also from the heavy reliance on game slang, abbreviations, and domain-specific jargon, which generic language models are poorly suited to recognize. This paper presents a shared task for in-game toxic language detection built upon real-world in-game chat data, and proposes the best-preforming model for the toxic language slot filling: Bi-directional Representations with Attention Residuals (BRAR). Experimental results demonstrate that BRAR effectively captures the global context and outperforms the existing baselines on slot filling.

Mon 28 SeptComputation and LanguageArtificial IntelligenceMachine Learning
The gist
Toxic language in gamer chat—like insults or mean words—is tricky to spot because players use lots of slang and abbreviations. The authors created a new detection model called Bi-directional Representations with Attention Residuals (BRAR) that understands the whole chat context better than older methods. They tested BRAR on real in-game chat data and found it works best for identifying toxic language. This model helps make online gaming conversations safer by catching harmful messages more accurately.
Open → 2609.34584v1

Visual system identifies deceptive patterns in ai generated images

ASAP: Visual Analytics for Identifying and Analyzing Image Patterns in AI-generated Images

Abstract: Generative image models can produce highly realistic images, raising concerns about potential misuse in creating deceptive content. Current deepfake approaches face several challenges, including limited generalizability, lack of interpretability, and poor actionability. To help address these, we present ASAP, an interactive visualization system designed to empower users in the analysis and summarization of deceptive patterns in AI-generated images. ASAP introduces a novel CLIP-adapted image encoder that generates interpretable representations, enabling the extraction of influential pixel regions via calculated masks. This approach facilitates the identification of key deceptive features through influence measurement techniques. These backend techniques are integrated into a visual analytics dashboard that allows users to quantify and analyze authenticity-indicative patterns in image collections containing both authentic and AI-generated images. This approach also supports the comparative analysis of various generative models, including GANs and diffusion models. We demonstrate ASAP's efficacy through a user study and two application scenarios using established fake image detection benchmarks, showcasing its ability to effectively extract and quantify deceptive patterns.

Wed 23 SeptHuman-Computer InteractionComputer Vision and Pattern Recognition
The gist
AI can create images that look very real, which raises worries about fake content. The authors built a tool called ASAP that helps people find and understand tricky patterns in AI-made images. ASAP uses special image analysis techniques to highlight important parts of images that reveal if they are fake. It shows these patterns visually to help users compare real and AI-generated images, and different AI models. The tool was tested with users and standard benchmarks, proving it helps spot and study deceptive image features.
Open → 2609.27371v1

Public suspicion of bots rises with Twitter controversies and account traits

When Does the Public Become Suspicious of Bots? Demand-Side Evidence from Botometer Query Logs

Abstract: We study private bot-checking behavior from the demand side: when people suspect an account is automated, whom they suspect, and what follows. Using Botometer's server-side query logs, the most widely used bot-detection service, we treat each query as a behavioural trace of suspicion. We analyze over 1 million public checks of Twitter accounts from 2020 to 2023, enriched with 3.2 billion tweets from the contemporaneous 1% public stream. Collective suspicion spikes with platform crises, most sharply around the 2022 Musk-Twitter bot dispute. Checked accounts are older and more prolific, have more followers, and post promotional, political, and crypto content. Accounts that draw collective suspicion have higher bot scores and are more likely to be suspended. Bot-related public attention and Botometer activity are elevated during the same broad periods, although their short-run fluctuations are largely uncoupled. Public feedback focuses on first-person identity claims for humans, and evidence-based arguments citing posting rate, political content, and cross-account coordination for bots and cyborgs. Bot suspicion thus constitutes a mass, distributed form of platform auditing that tracks meaningful signals of automation, establishing audit-tool query logs as a novel lens on public responses to platform manipulation.

Thu 17 SeptComputers and SocietySocial and Information Networks
The gist
People get suspicious about whether Twitter accounts are actually bots who post automatically. The authors looked at over a million checks made using a popular bot-detection tool called Botometer between 2020 and 2023. They found that suspicion spikes during big Twitter events like the 2022 dispute involving Elon Musk. Suspicious accounts tend to be older, have more followers, and post political, promotional, or crypto content. This shows that many people help catch bots by checking accounts, making the platform safer and more transparent.
Open → 2609.20661v1