Papers for

speech data engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Fusealign improves word timing on noisy long audio recordings

FuseAlign: Forced Alignment in the Wild

Abstract: Word-level forced alignment estimates when each transcript word occurs in an audio recording. It underpins text-based media editing, subtitling, speech-data curation, and phonetic analysis. Existing evaluations understate the difficulty of forced alignment by relying on short, clean speech, perfect transcripts, and metrics that obscure consequential alignment errors. In contrast, real-world media and data-processing pipelines operate on long and diverse recordings. Additionally, forced aligners often operate on error-prone automatic speech recognition (ASR) output. We address these gaps with improved evaluation metrics, a scoring protocol for real ASR transcripts, and AlignBench, a benchmark spanning diverse speaker, acoustic, and text conditions. We further introduce FuseAlign, a transformer-based aligner trained on large-scale pseudo-labeled speech with online label correction. FuseAlign performs joint contextualization of audio and text for the localization of coarse words. The model then refines boundaries at millisecond resolution and detects missing transcript words in the audio without lexicon-based or Viterbi decoding. On AlignBench, FuseAlign substantially outperforms all baselines and remains robust under real ASR transcripts. Ablations show that convolutional upsampling and EMA-snapshot label correction matter more than model properties such as parameter count.

Sun 27 SeptMachine Learning
The gist
Figuring out exactly when each word happens in a spoken recording is important for subtitles and audio editing. The authors show that existing methods are tested on too-simple cases and don’t handle real-world problems like long recordings or imperfect transcripts. They created a new test set with varied conditions and introduced FuseAlign, a system that learns from large amounts of data and corrects itself while listening. FuseAlign finds words more accurately, even when the transcript has mistakes, without relying on traditional dictionary-based methods.
Open → 2609.33650v1

Audio encoders balance size and speed for better speech tasks

Joint Analysis of Latent Dimensionality and Frame Rate in Continuous Audio Encoders

Abstract: Continuous audio encoders compress audio along feature and time axes through latent width and frame rate, but their joint effect on downstream performance remains unclear. We train sixteen encoders spanning four widths and four frame rates, with downstream adapters and probes, using matched training protocols. Despite generally improved reconstruction at larger widths, automatic speech recognition (ASR) and spoken question answering (SQA) favor moderate widths at higher rates, with the best observed widths shifting toward larger values under stronger temporal compression. Frozen-model PCA interventions reveal distinct reconstruction and recognition sensitivities: removing the trailing half of the components substantially degrades ASR in high-rate 512-dimensional encoders with comparatively small reconstruction penalties, whereas 1024-dimensional encoders largely preserve both. Yet the projected 1024-dimensional model underperforms unmodified narrower models on ASR at 12.5Hz. These findings identify a width--rate interaction in downstream utility and suggest that how representations are organized during training matters beyond reconstruction fidelity and compressibility.

Thu 24 SeptSound
The gist
Audio encoders shrink sound data by changing two things: how many features they keep (width) and how often they sample in time (frame rate). The authors trained many versions with different widths and rates to see how these choices affect tasks like speech recognition and answering spoken questions. They found bigger isn’t always better; moderate sizes at higher speeds worked best for understanding speech, even if bigger sizes gave better sound reconstruction. This shows that just making the audio sound closer to the original doesn’t guarantee better performance in real tasks.
Open → 2609.29780v1