Papers for
hearing aid designers
Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.
Bathroom sound recognition model learns to identify activities from audio
SincDPNet: Interpretable Raw-Waveform Bathroom Activity Recognition for Assistive Living
Abstract: Bathroom acoustic-event recognition can support ambient assisted living in settings where continuous video monitoring is undesirable. However, practical deployment requires models that are compact, interpretable, and robust to changes in the recording environment. This work introduces \dataset{}, a seven-class bathroom acoustic-event dataset containing 21{,}387 annotated clips recorded across five environments, and proposes SincDPNet, a compact raw-waveform classifier with a learnable sinc filter bank followed by a depthwise-separable convolutional body. Each sinc filter is controlled by two frequency parameters, allowing the learned passbands to be inspected directly in hertz while keeping the front end small. To reduce room-specific leakage, recording sessions and environments are separated before overlapping windows are assigned to the training, validation, and test partitions. We further use multi-objective Bayesian optimization as a design tool to examine the validation performance--model-size trade-off across 24 configurations. The selected designs span different operating points: the best-performing model achieves 80.2\% accuracy and 0.760 macro-F1 with 14{,}040 parameters, while the compact $N_f=25$ configuration uses only 2{,}848 parameters and achieves 75.7\% accuracy, 0.661 macro-F1, and 0.716 MCC on the held-out environment. Analysis of the learned filters and confusion patterns shows that spectral overlap contributes to confusion among water-related events, while the \textit{Door}/\textit{Walker/Crutch} errors also reflect similarities in their transient temporal structure.
New way to improve audio signals by focusing on sound levels
Prox-Friendly Log-Magnitude Prior on Complex-Valued Signal
Abstract: The logarithmic transform is essential in audio signal processing since human auditory perception is approximately logarithmic with respect to magnitude. However, directly incorporating prior knowledge about signals (e.g., harmonic structure) in the log-magnitude domain into optimization problems solved by standard proximal splitting algorithms remains challenging. To address this issue, this paper proposes a novel regularizer termed EPILOG (Exponential Penalty for Imposing priors on LOG-magnitude). EPILOG indirectly imposes prior knowledge on the log-magnitude of a complex-valued signal through regularization of an auxiliary variable that is shown to be linked with the log-magnitude. Furthermore, we derive its variable-wise proximity operators and develop a proximal splitting algorithm using these operators. Experiments on speech dereverberation demonstrate the effectiveness of the proposed regularizer, particularly in promoting cepstral-domain sparsity.
WhisperVC-AV improves noisy whisper to normal voice conversion
WhisperVC-AV: Audio-Visual Content Restoration for Noise-Robust Whisper-to-Normal Voice Conversion
Abstract: Environmental noise obscures the content cues needed for whisper-to-normal conversion. We propose WhisperVC-AV, which restores acoustic content features while leaving the original WhisperVC conversion module unchanged. Its context-guided restoration module combines temporal acoustic context with synchronized lip features through attention and gated residual correction. Experiments on AISHELL6-Whisper show lower character error rates (CERs) than WhisperVC across three ASR systems, on clean speech and under all six signal-to-noise ratio (SNR) conditions with MUSAN noise. The largest gains occur at 0 dB SNR, where Qwen3-ASR CER falls from 36.29% to 28.88%. WhisperVC-AV also improves predicted speech quality while maintaining speaker similarity. The CER gains extend to unseen background noise without retraining, while visual controls support the use of utterance-specific lip cues. Audio examples are available on our demo page.
Multimodal dataset improves attention tracking with brain and eye signals
MAESTRO: a Multimodal Auditory-attention Egocentric Speech-TRacking Open corpus
Abstract: Humans rely on gaze, head movements, and visual cues to attend to speakers in noisy environments, yet auditory attention decoding (AAD) has been studied primarily using electroencephalography (EEG). We introduce the Multimodal Auditory-attention Egocentric Speech-TRacking Open (MAESTRO) corpus, the first AAD dataset to simultaneously record EEG, eye gaze, pupillometry, egocentric video, and head inertial measurement unit (IMU) data. MAESTRO includes four competing speakers and background noise across multiple signal-to-noise ratio (SNR) conditions, enabling attention decoding under realistic listening scenarios. Through a four-speaker attention decoding benchmark, we show that combining behavioral and physiological signals improves decoding performance over EEG-only approaches, enabling future advances in multimodal auditory attention decoding. These findings open the door to new applications, analyses, and methodological advances in multimodal AAD. The complete dataset is publicly available at https://huggingface.co/datasets/aspire-osu/maestro-eeg-dataset . The official code repository is available at https://github.com/ASPIRE-OSU/MAESTRO .
Audio-visual method enhances target speaker voice with low delay
Adapting Personalized Speech Enhancement for Low-Latency Audio-Visual Target-Speaker Extraction
Abstract: Online audio-visual target-speaker extraction aims to remove competing voices while preserving speech quality and bounding lookahead. Existing extractors are built and evaluated for separation on synthetic mixtures, leaving listening quality and meeting behavior largely untested. We introduce Audio-Visual Personalized Voice Quality Enhancement (AV-PVQE), which approaches these requirements from the other direction. We start from a personalized speech enhancement model that reconstructs a requested voice at high quality but confuses the target in 46% of two-speaker mixtures despite clean enrollment. Adding mouth features at its speaker-conditioning input and jointly fine-tuning the visual and reconstruction networks reduces this rate to 1.6%, with no future frames and 20 ms of algorithmic delay. Compared with an online autoregressive audio-visual extractor, AV-PVQE yields separation gains on two synthetic benchmarks and larger gains on recorded meetings, and keeps its advantage on excerpts with more speakers than the fine-tuning mixtures. In personalized P.835 listening tests on two meeting corpora, it improves overall quality over this extractor by 0.57 and 0.63 MOS, with similar mean rating relative to the starting model. Preservation and rejection tests show that it keeps the target intact when no competing voice is present and suppresses competing speech when the target is absent.
Sparse graph method improves audio-visual speech cleaning efficiency
G-Mamba: Sparse Graph-Guided Mamba for Audio-Visual Speech Enhancement
Abstract: Lightweight audio-visual speech enhancement (AVSE) models face a critical trade-off between computational efficiency and cross-modal alignment accuracy. While simple concatenation lacks relational expressiveness, dense cross-attention incurs computational overhead and is prone to unreliable cross-modal correspondence under strong acoustic interference. We propose Sparse Graph-Guided Mamba (SG-Mamba), a lightweight AVSE framework that integrates a sparse heterogeneous graph with a linear-complexity Mamba backbone. The graph explicitly models modality-specific relations through content-adaptive attention and cross-frame audio-visual connections, while Mamba captures long-range temporal context. We further introduce an audio skip connection to preserve spectral detail without sacrificing noise suppression. Evaluated on LRS3, SG-Mamba achieves competitive or superior performance against strong lightweight baselines and reaches 13.091 dB SI-SDR under noise-only condition. It also remains robust in cluttered multi-speaker conditions with a competitive cost of 3.45 G MACs (or 6.90 G FLOPs). Results on VoxCeleb2 further suggest that explicit structural priors improve robustness, generalizability, and computational efficiency in lightweight AVSE.