Papers for

content moderators

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Clipping improves detecting AI-generated text even with editing errors

Robust Detection of LLM-Generated Text under Contamination

Abstract: We study the detection of LLM-generated text under editing and contamination. Modeling human and machine text as finite-order Markov processes with Huber contamination, we characterize an exact boundary for reliable detection under our assumptions. Detection is impossible when contamination is sufficiently large relative to clean-source separation. Below this boundary, a collection of clipped likelihood-ratio tests achieves vanishing worst-case errors. This construction motivates clipping as a simple modification of existing statistical detectors. For a broad class of additive scores, we identify conditions under which the clipped test is consistent while the raw test's worst-case power tends to zero. We evaluate seven detectors across three datasets and three generation models, and on the RAID benchmark. Clipping improves robustness in both studies, with gains varying across detectors and contamination settings. For example, at a target false-positive rate of 5\%, clipping improves the log-likelihood--log-rank ratio (LRR) detector's true-positive rate by a median of 8.3 percentage points in the controlled study and 2.1 and 4.3 points in rate- and attack-specific RAID evaluations, respectively.

Thu 24 SeptMachine Learning
The gist
Detecting whether text is written by a human or an AI can be tricky when the AI-generated text is altered or mixed with other content. The authors studied this problem by modeling both human and machine text with a mathematical tool and found the limits where detection is possible. They show that by modifying existing detection tests with a simple 'clipping' step, the tests become more reliable against tricky edits or contamination. Their experiments on several datasets and models confirm that clipping consistently improves the ability to spot AI-generated text while keeping false alarms low.
Open → 2609.29935v1

How query wording affects AI agreement in relationship advice chats

Sweet Talkers: How Query Formulation Shapes Sycophancy in Romantic Relationship Advice

Abstract: Large language models (LLMs) are increasingly used for emotional support and relationship advice, where a model's tendency to preserve a user's face can inadvertently reinforce harmful interpersonal behaviors. To systematically examine this risk, we developed the Romantic Relationship Advice-Seeking Prompts (RRASP) dataset of 2,400 prompts across five relationship themes and evaluated social sycophancy using the ELEPHANT framework on two consumer-facing models, GPT-5 Mini and Gemini 3 Flash. Contrary to our initial hypothesis, grammatical mood alone did not produce systematic differences in sycophantic behavior, suggesting that what a user implies matters more than how they phrase it. Instead, perspective-driven framing had a stronger influence, with gaps between original and flipped prompts widening in follow-up responses. Consistent increases in framing and moral sycophancy across turns indicate that models become more likely to accept a user's stated premises and affirm their ethical stance as a dialogue progresses. Notably, Gemini 3 Flash exhibited substantially smaller increases in moral sycophancy than GPT-5 Mini, suggesting it is more resistant to reinforcing ethically problematic positions across turns.

Sat 12 SeptComputation and Language
The gist
Sometimes, AI chatbots that give relationship advice try to agree with what users say, even if it might encourage bad behavior. The authors created a big set of questions about relationships and tested how two AIs responded depending on how the questions were asked. They found that the AI's agreement depends more on what users imply rather than the exact words they use. Some AIs become more likely to accept a user's views as conversations go on, but one AI was less prone to agreeing with problematic opinions. This helps us understand how AI might unintentionally support harmful ideas in emotional chats.
Open → 2609.13841v1

Multilingual large language models struggle with Urdu stories

Multilingual in Name Only? Cultural and Linguistic Weaknesses of LLMs in Urdu

Abstract: Multilingual large language models (LLMs) are increasingly used for open-ended text generation, yet their behaviour in low-resource languages remains poorly understood. In this work, we question how correct and reliable is the generation of multilingual LLMs when used for the task of story generation. We consider Urdu language as a representative low-resource language. We generate Urdu-Stories, a corpus of 93 stories generated using three contemporary LLMs (GPT-5.1, Qwen-3-Max, DeepSeek-3.1). We manually annotate the errors present in them under a nine-label linguistic, semantic, and cultural taxonomy. Our notable findings suggest that LLMs often make basic errors of grammar and semantics. The stories lack coherence, have unnatural repetition and show pervasive cultural shallowness. We further show using few-shot prompting that the cultural and context errors largely remain unresolved. Our findings highlight the limitations of current LLMs as a reliable source of content generation and information retrieval for low-resource languages.

Wed 9 SeptComputation and LanguageArtificial IntelligenceMachine Learning
The gist
Large language models (LLMs) can write stories in many languages, but their ability in less common languages like Urdu is not well understood. The authors studied three popular LLMs that generated Urdu stories and found many problems. These stories often had basic grammar mistakes, confusing meanings, repeated phrases, and lacked cultural depth. Even when given extra examples to improve, the issues remained. This shows that current LLMs are not yet reliable for creating or understanding content in low-resource languages like Urdu.
Open → 2609.10758v1