ParsHate dataset enables Persian hate speech and target detection research
ParsHate: A Benchmark Dataset for Hate and Target Detection in Persian
Computation and LanguageDatabases
Summary
Detecting hate speech helps keep social media safe, but there isn’t enough data in Persian. The authors created ParsHate, a large dataset with 10,000 Persian tweets over ten years, carefully labeled for hateful content and who the hate is directed at. They also marked whether the hate and its targets were explicit or hidden and explained why. Tests with current hate speech detection methods showed there’s room to improve, especially for identifying targets. ParsHate is freely available to help build better tools.
What this means in practice
- •For social media moderation teams: Improve automated detection of hateful Persian content and identify who hate speech targets to enhance content moderation.
- •For language technology developers: Train and evaluate Persian hate speech detection models using ParsHate’s detailed annotations for both hate and target identification.
Authors
Zahra Bokaei, Walid Magdy, Bonnie Webber
Abstract
We introduce ParsHate, a manually annotated dataset of 10,000 Persian tweets spanning 2013-2022, representing the first decade-long benchmark for hate speech detection in Persian. The dataset contains 31% hateful content and supports both hate detection and multi-label fine-grained target identification across seven structured target categories. ParsHate also distinguishes explicit and implicit hate, marks explicit and implicit targets, and provides span-level rationales. Data collection combines random and score-stratified temporal sampling to reduce keyword-driven bias while preserving natural label distributions. Applying SOTA models for Persian hate-speech detection on ParsHate shows moderate performance (79% F1), especially with samples from earlier years, and low performance with target identification (25.5% macro-F1). This emphasizes the diverse sampling of hate speech in ParsHate and its challenging nature that requires more advanced methods for better performance. Dataset is made publicly available.