Transformer model improves verifying family relations from speech
A New Transformer-Based Approach for Audio-Based Kinship Verification and a New Uncontrolled Mandarin Kinship Speech Dataset
SoundArtificial IntelligenceMachine Learning
Summary
Determining if two people are closely related just by listening to their voices is a tricky task. The authors created a new method using a transformer-based model called CONVTRAP-TN to better recognize family connections from audio. They also collected a new Mandarin speech dataset recorded in everyday situations to better train and test such systems. Their experiments show current methods struggle when tested on different datasets, suggesting their approach and data reflect real-world use better.
What this means in practice
- •For voice assistant developers: Integrate kinship verification models for personalized authentication even under everyday noisy recording conditions.
- •For forensic audio analysts: Use the new transformer-based audio kinship verification to assess familial relationships from recordings collected in uncontrolled environments.
Authors
Qiyang Sun, Langqing Zhang, Yupei Li, Björn Schuller
Abstract
Kinship verification is a task involving determining whether two individuals share a first-order kin relation. To tackle this task, we propose CONVTRAP-TN, a new architecture for audio-based kinship verification, and conduct an ablation study on the proposed model. To the best of our knowledge, we are the first to apply the successful transformer architecture to the task of audio-based kinship verification. Furthermore, we also collect a custom speech dataset, ARKIN, which accurately reflects everyday recording conditions. We do this because only a few speech datasets with kinship labels currently exist, all of which either source extremely noisy in-the-wild data from the internet, or instruct speakers to record in specific environments. These settings fail to reflect real-world scenarios where users record on personal devices under unrestrained conditions. Additionally, we perform a series of preliminary baseline experiments on the collected dataset, including speaker verification and recognition, speech recognition, age estimation, and kinship verification, as well as cross-dataset kinship verification experiments to show that existing methods are not robust across datasets.