Fusealign improves word timing on noisy long audio recordings

FuseAlign: Forced Alignment in the Wild

Machine Learning

Summary

Figuring out exactly when each word happens in a spoken recording is important for subtitles and audio editing. The authors show that existing methods are tested on too-simple cases and don’t handle real-world problems like long recordings or imperfect transcripts. They created a new test set with varied conditions and introduced FuseAlign, a system that learns from large amounts of data and corrects itself while listening. FuseAlign finds words more accurately, even when the transcript has mistakes, without relying on traditional dictionary-based methods.

What this means in practice

  • For media production teams: Improve subtitle timing and editing automation on real-world recordings with noisy transcripts.
  • For speech data engineers: Enhance speech dataset curation by accurately aligning words in long and diverse audio with imperfect transcripts.

Authors

Mithilesh Vaidya, Stephen Bailey, Sumukh Badam, Matthew Bendel, Xingzhe He

Abstract

Word-level forced alignment estimates when each transcript word occurs in an audio recording. It underpins text-based media editing, subtitling, speech-data curation, and phonetic analysis. Existing evaluations understate the difficulty of forced alignment by relying on short, clean speech, perfect transcripts, and metrics that obscure consequential alignment errors. In contrast, real-world media and data-processing pipelines operate on long and diverse recordings. Additionally, forced aligners often operate on error-prone automatic speech recognition (ASR) output. We address these gaps with improved evaluation metrics, a scoring protocol for real ASR transcripts, and AlignBench, a benchmark spanning diverse speaker, acoustic, and text conditions. We further introduce FuseAlign, a transformer-based aligner trained on large-scale pseudo-labeled speech with online label correction. FuseAlign performs joint contextualization of audio and text for the localization of coarse words. The model then refines boundaries at millisecond resolution and detects missing transcript words in the audio without lexicon-based or Viterbi decoding. On AlignBench, FuseAlign substantially outperforms all baselines and remains robust under real ASR transcripts. Ablations show that convolutional upsampling and EMA-snapshot label correction matter more than model properties such as parameter count.