Papers for

audio engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Phoneme aware method improves extreme speech quality restoration

P2Flow: Phoneme-aware Progressive Flow Matching for Extreme Speech Super-Resolution

Abstract: Generative models have recently demonstrated considerable promise in speech super-resolution (SSR). Nevertheless, the majority of existing work has concentrated on standard or versatile SSR configurations, leaving the extreme setting with severely limited spectral inputs largely unexplored. In this regime, current approaches exhibit marked performance degradation, underscoring the need for dedicated solutions. To bridge this gap, we introduce P2Flow, a phoneme-aware progressive flow matching (FM) framework designed for extreme SSR with three main strategies. First, our model leverages phonetic information to reconstruct missing spectral components. Furthermore, it employs a progressive architectural design that hierarchically restores distinct frequency regions. Finally, we incorporate post-training of the vocoder to enhance overall waveform fidelity. Extensive experiments are conducted on the TIMIT and VCTK datasets under both 1 kHz to 16 kHz and 2 kHz to 16 kHz settings, demonstrating that P2Flow yields state-of-the-art results across multiple evaluation metrics.

Mon 21 SeptMachine LearningSound
The gist
Audio recordings often lose sound quality when their frequencies are greatly reduced. The authors address this problem by creating a new method that uses speech sounds (phonemes) to help fill in missing details. Their approach also restores different frequency ranges step-by-step to rebuild clearer speech. After training, they refine the sound-producing model to make the output more natural. They tested their method on well-known speech datasets and found it performed better than previous techniques.
Open 2609.24138v1

Equivariant deep learning improves sound event detection and localization efficiency

EquiSELD: Efficient training of equivariant sound event localization and detection networks

Abstract: First-order Ambisonics (FOA) signals exhibit exact O(3) symmetry: The rotation or reflection of the FOA signal modifies the direction of arrival of the sound sources, while preserving the sound sources themselves. Prior attempts to utilize this spatial symmetry of FOA to improve the efficiency and robustness of sound event detection and localization (SELD) systems either learned only an approximation of the symmetry through rotation-based augmentation or relied on computationally expensive methods to integrate equivariance. Furthermore, prior work focused exclusively on SO(3) equivariance , leaving the potential of incorporating O(3) equivariance for SELD tasks unclear. To address these limitations, we developed EquiSELD. This equivariant attention network processes first-order Ambisonics as paired streams of O(3)-invariant scalars and equivariant intensity vectors, producing an invariant activity magnitude and an equivariant DOA with a Multi-ACCDOA readout. To compare the impact of O(3) versus SO(3)-equivariance, we designed a matched SO(3)-only variant. EquiSELD outperforms prior equivariant networks on both simulated scenes with measured RIRs and recordings of real-world sound scenes at a fraction of the training cost. EquiSELD additionally surpasses the performance of non-equivariant SELD networks of a similar size on the simulated real-world sound scenes and achieves competitive performance on the real-world sound scenes.

Sat 19 SeptSoundArtificial Intelligence
The gist
Sound event detection and localization involves figuring out what sounds are happening and where they come from. The paper introduces a new method called EquiSELD that uses a mathematical symmetry property of sound signals to process them more efficiently and accurately. This approach splits the sound data into parts that stay the same under rotation and parts that change predictably, helping the model understand sounds regardless of their direction. Their method works better and trains faster than previous models using similar ideas, and also beats non-specialized models on simulated real-world sounds.
Open 2609.23156v1

Soundfield embeddings add spatial info to audio event detection

Bearings: Self-Supervised Soundfield Embeddings from First-Order Ambisonics

Abstract: Recently proposed self-supervised audio encoders learn powerful general-purpose representations of sound scenes, yet they are spatially blind. To supply the missing spatial representation of sound scenes, we introduce Bearings. Bearings is a self-supervised framework that learns soundfield embeddings from unlabeled first-order Ambisonics. We pre-train a masked auto-encoder paired with a decoder conditioned on frozen acoustic embeddings from an off-the-shelf single-channel audio encoder. Our results show that the resulting soundfield embeddings form a reusable stream that can be attached to frozen acoustic encoders with a lightweight trainable fusion head. On sound event localization and detection, concatenating our soundfield embeddings with acoustic representations provides the missing spatial information and enables joint detection and localization, raising the location-dependent F-score from below 4 to 50 on TAU-NIGENS 2021 and 39 on STARSS23. To our knowledge, Bearings is the first self-supervised soundfield encoder whose embeddings plug into frozen acoustic encoders without retraining either model.

Sat 19 SeptSoundArtificial Intelligence
The gist
Current systems that learn from audio scenes do not understand where sounds come from in space. The authors introduce Bearings, a way to teach computers to understand the spatial layout of sounds using special audio recordings called first-order Ambisonics. Bearings creates compact representations of sound fields that can combine with existing audio encoders without changing those encoders. This combination improves the ability to find and locate sounds in an environment, making detection much more accurate.
Open 2609.23152v1

Replicating band split RNN reveals energy costs in music separation

Investigating the Performance and Energy Costs of Replicating Band-Split RNN for Music Source Separation

Abstract: Band-split recurrent neural network (BSRNN) is a popular music source separation model that yields close to state-of-the-art results using reasonable computational resources and public datasets. It is therefore interesting from a reproducible research perspective, but achieving its performance is not straightforward since its full code is not available. In this paper, we conduct a replication of BSRNN via implementing the full pipeline. We extend the original paper's analysis by experimentally studying various design choices about data preprocessing, the optimization protocol, and architectural parameters. We report and discuss this project's energy cost, and we underline how its footprint could have been substantial lower upon availability of the full pipeline, which advocates for more reproducible research practices. To comply with this objective, we publicly release our code and pre-trained models.

Fri 18 SeptSound
The gist
Separating different sounds from music recordings is tricky, and one popular method uses a type of AI called band-split recurrent neural networks (BSRNN). The original method was partly hard to copy because its full code wasn’t shared. The paper recreates the entire process, tests different setups, and measures how much energy the training uses. The authors also share their code and models to help others avoid high energy costs and make such research easier to repeat.
Open 2609.21918v1

Minimax-H3 shows limited physical reasoning using multiple input types

Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model

Abstract: Recent Omni-Modal Generative Models (Omni-Models) have advanced content generation toward unified modeling of text, images, video, and audio. MiniMax-H3 exemplifies this transition by combining multimodal context understanding with joint audio-visual generation in a shared latent framework. Its unified architecture raises a fundamental question: Can multimodal alignment improve the model's world reasoning, and what new evaluation paradigms do omni-modal inputs enable? To investigate this question, this work introduces a comprehensive evaluation framework organized around four complementary dimensions of physical world reasoning. Unlike existing evaluation frameworks for video generation and world models, which are often constrained by limited input modalities and evaluation settings where prompts closely match the target video content, our evaluation is specifically designed to exploit the multimodal inputs of Omni-Model. We construct a diverse set of novel tasks that require models to integrate complementary information across modalities. Specifically, we consider four scenarios, including implicit prompts paired with multiple frames, audio-image, prefix-videos, and audio-video inputs. Every single modality provides only partial evidence about the underlying event, requiring the model to jointly reason over the complementary semantic cues to infer latent event states and future dynamics. Across 517 evaluation instances, MiniMax-H3 achieves an overall success rate of 41.97%. Video-based Decision Reasoning yields the highest success rate at 56.00%, while Audio-based Disambiguation Reasoning is the weakest, reaching only 27.40%. These results indicate that effective multimodal integration remains key to fully exploiting the benefits of diverse input modalities. The project is available at https://github.com/gulucaptain/MiniMax-H3-Reason.

Wed 16 SeptComputer Vision and Pattern Recognition
The gist
Minimax-H3 is a computer model that tries to understand and predict events by looking at different kinds of information, like pictures, videos, and sounds together. The paper asks whether using all these kinds together helps the model better understand the real world. To test this, the authors created tricky tasks where each kind of input only gives part of the story, so the model must combine clues to guess what happens next. They found that while the model is better at some tasks, like using video to make decisions, it struggles more with tasks relying on audio clues. This shows there is still work to do to improve how models mix different types of information to understand the world.
Open 2609.18323v1

Musical recordings made together have more precise timing

Musical Timing in Studio Recordings

Abstract: Studio recording techniques vary considerably, ranging from live recordings in a shared acoustic space, to isolated booth recording and overdubbing, where musicians record their parts separately, while listening to previously recorded material. These approaches differ in terms of physical co-presence, visual contact, sound leakage, and acoustic isolation, factors that may influence musical coordination. This study investigates whether the recording style affects timing precision; specifically, whether musicians performing together in the same space achieve tighter temporal coordination. To address these questions, two datasets with a total of 391 multi-track recordings were analyzed to measure the timing relationships between musicians. In addition, an automated sound-leakage detection method was employed to infer the recording conditions of each song: pairwise Mel-spectrogram similarities between the tracks were computed, and summary statistics from the resulting similarity matrices were used to identify recordings with shared acoustic content. The results suggest that recordings showing evidence of shared room acoustics and sound leakage tend to exhibit lower timing variability than highly isolated recordings, whereas overdubbing is associated with increased timing variability.

Sat 12 SeptSound
The gist
This study looked at how musicians' timing differs depending on how they record music. When musicians record together in the same room, their timing tends to be more precise than when they record their parts separately with overdubbing. The researchers used sound analysis to figure out whether recordings shared room sounds or were isolated. They found that shared room acoustics usually lead to better timing coordination, while overdubbing can increase timing variability.
Open 2609.13881v1