Papers for

online community managers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Bluesky content moderation mixes AI and human review to catch harms

Characterizing Bluesky Content Moderation Service: From Automation of Service to Landscape of Harms

Abstract: Empirical research on content moderation is fundamentally constrained by the opaque deployment of moderation systems on major social media platforms. To this end, the recent emergence of decentralized platforms with transparent, public moderation logs presents an unprecedented opportunity for independent audits. In this work, we leverage this architectural transparency to conduct the first large-scale audit of the default moderation system on Bluesky, the Bluesky Moderation Service (BMS). Analyzing its 10.6M moderation labels from 2025, we investigate three foundational aspects: (i) its mechanism (the degree of automation versus human oversight), (ii) its efficacy (accuracy in detecting harms), and (iii) its purpose (the landscape of harms it identifies). Our findings reveal a human-AI collaborative system where labels for sexual and graphic content are applied automatically in seconds, while nuanced and high stakes labels require more human oversight, taking hours or days. Through a manual annotation study, we find the BMS operates with high precision (0.837), but struggles with low recall (0.222), with our annotators identifying 4.5$\times$ more harmful content than the moderation system in a random sample. Finally, unsupervised clustering of the most frequently applied labeled posts uncovers detected harms ranging from hostility in discourse toward protected groups to the spread of sexually explicit and other graphic content. Our work offers a look into the operational realities of a deployed moderation system, providing a concrete data-driven foundation for designing more effective and transparent moderation systems.

Thu 10 SeptComputers and SocietyArtificial Intelligence
The gist
Many social media platforms keep their content moderation methods secret, making it hard to study how well they work. The authors analyze Bluesky’s moderation system because it shares all its decisions publicly. They find that Bluesky uses AI to quickly flag obvious harmful content like sexual or graphic posts, but relies on humans for complex cases, which take longer. While Bluesky is precise in the harms it identifies, it misses many harmful posts that human reviewers found. This study helps show how practical moderation combines automation and human judgment.
Open 2609.11373v1

Foundation models improve online content moderation accuracy

Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization

Abstract: The growing complexity of content moderation policies presents a critical challenge for their consistent operationalization. While foundation models possess the basic capabilities needed to confront this challenge, whether they can reliably moderate online content remains an unanswered question. In this paper, we systematically compare two competing paradigms for Vision-Language Model (VLM) guidance: an instruction-driven approach where models reason from policy precepts, and an example-driven approach where they generalize from prior precedents. We ground this investigation in ModerationBench, a new benchmark of 4,000 manually annotated, in-the-wild posts from the Bluesky platform. Our experiments reveal that foundation models can substantially outperform Bluesky's deployed moderation system, nearly tripling its $F_1$ score (0.60 vs. 0.22) on Random Posts in the benchmark, with both instruction- and example-driven paradigms achieving comparable peak effectiveness. Our findings thus chart a path toward reliable and adaptable policy operationalization at scale.

Wed 9 SeptComputation and LanguageArtificial IntelligenceComputers and Society
The gist
Online content moderation is tricky because the rules can be complex and hard to apply consistently. The authors compare two ways of guiding AI models—one using instructions based on rules, and another using examples from previous decisions. They created a new test set from real online posts and found that big AI models can perform much better than an existing system at deciding what content should be allowed or removed. Both guiding methods worked about equally well, showing promise for more reliable moderation.
Open 2609.10410v1

TikTok community forms around sorority recruitment event videos

Ephemeral Feeds and Enduring Rituals: RushTok and the Formation of Event-Based Algorithmic Communities

Abstract: Each August, TikTok's For You page turns the University of Alabama's sorority recruitment into RushTok. We examine RushTok as an event-based algorithmic community: a collective assembled around a bounded offline ritual and sustained by recommendation. Using a mixed-methods survey (n=71) and a reflexive account of creator outreach, we ask who participates, how, and with what stakes. Findings show an ambiguous and entertainment based throughline; many called it a community (51/71) but few claimed membership (11/71). Affiliation centered on creators rather than shared practices, with parasocial attention clustering around a small set of potential new members (PNMs) and returning figures. Higher content exposure tracked with self-identification as a community member; those members commented, followed creators, and engaged across videos. Attempts to interview creators were met with silence or refusals, reflecting community boundary-work despite viral visibility. We outline implications for platform governance, including time-bounded context, graduated visibility, and aftercare.

Tue 8 SeptHuman-Computer InteractionComputers and SocietySocial and Information Networks
The gist
Every August, TikTok highlights videos about sorority recruitment at the University of Alabama, creating a temporary online community called RushTok. The authors studied who joins this community and how they interact with the content and each other. They found that many viewers see it as a community but few actually feel like members, with attention focused on certain creators and participants. Video engagement is linked to stronger feelings of belonging, but creators often avoid interviews, showing mixed boundaries despite the group’s popularity. The study discusses what this means for managing platforms like TikTok.
Open 2609.09331v1

Voting patterns overstate support in online community elections

Endorsement Without New Evidence: How Sequential Voting Inflates Mandates in Online Community Governance

Abstract: Online communities often treat large support margins in public elections as strong mandates. We argue that such margins can overstate the independent scrutiny behind a decision. Using 198,275 free-text rationales from Wikipedia admin elections, we introduce vote-text divergence, a measure that flags a decisive vote paired with a thin, deferential rationale. Divergence rises as voters arrive later, even after controlling for voter and election fixed effects. The pattern is consistent with information saturation: once prior text is accounted for, arrival order no longer predicts divergence, while accumulated prior evidence does. The effect is strongest among peripheral voters in the co-voting network. Yet divergence does not predict worse post-promotion outcomes, such as administrative activity or survival. Public tallies can therefore weaken the scrutiny signal even while selecting capable administrators: a margin may appear to reflect more consensus and support than it actually contains.

Tue 8 SeptSocial and Information NetworksComputers and SocietyHuman-Computer Interaction
The gist
Sometimes in online communities, a big win in a public vote is seen as strong support, but this study shows that later voters often repeat earlier opinions without adding new reasons. The researchers looked at Wikipedia admin elections and found that votes arriving later tend to have weaker explanations, suggesting people rely on previous comments more than their own judgment. This can create a false sense of overwhelming agreement even though the actual scrutiny might be limited. However, these large margins do not seem to harm the quality or activity of the elected admins.
Open 2609.09321v1