Mexican spanish video dataset advances hate speech detection research
MexHat: A Dataset for Hate Speech Detection in Mexican Spanish Videos
Computation and LanguageComputer Vision and Pattern Recognition
Summary
Detecting hate speech online is important but tricky because it depends on subtle cultural and language clues. The authors created MexHat, a collection of about 1,000 video clips in Mexican Spanish labeled to show whether they contain no negativity, offensive language, or hate speech. The dataset also breaks down hate speech into detailed categories. This resource helps build and test tools that understand hate speech better in Mexican Spanish videos.
What this means in practice
- •For content moderation teams: Improve detection systems for hate speech in Mexican Spanish video content by training models on culturally relevant examples with fine-grained labels.
- •For social media platform engineers: Develop better automatic filters to identify and remove hate speech in Mexican Spanish videos by using the dataset for training and evaluation.$Commercial implications: Enables commercial content moderation tools targeting Spanish-speaking markets to finely detect harmful speech in videos.
Authors
Itzel Tlelo-Coyotecatl, Hugo Jair Escalante
Abstract
Ensuring online safety through content monitoring had raised Hate Speech Detection as a crucial task to be addressed. By essence the task demands the capture of contextual cues, which are essential for a precise understanding of the content's intent. Although automated detection approaches for the task have advanced significantly, the scarcity of non-English resources persists, limiting the ability of models to adapt to the subtle, context-dependent, and culturally related nature of multimodal content. In this paper, we introduce MexHat, a video dataset designed to capture the linguistic and cultural cues for the hate-speech detection task in a Mexican Spanish context. Our dataset comprises around 1k video clips annotated across two tasks: a three-way class evaluation (no negative content, offensive content and hate-speech content), and a fine-grained class evaluation including three hate-speech sub-categories. The dataset statistics and the baseline results highlight the inherent challenges associated with the task. Disclaimer: This paper contains sensitive content that may be disturbing to some readers.