Papers for

hearing aid designers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Bathroom sound recognition model learns to identify activities from audio

SincDPNet: Interpretable Raw-Waveform Bathroom Activity Recognition for Assistive Living

Abstract: Bathroom acoustic-event recognition can support ambient assisted living in settings where continuous video monitoring is undesirable. However, practical deployment requires models that are compact, interpretable, and robust to changes in the recording environment. This work introduces \dataset{}, a seven-class bathroom acoustic-event dataset containing 21{,}387 annotated clips recorded across five environments, and proposes SincDPNet, a compact raw-waveform classifier with a learnable sinc filter bank followed by a depthwise-separable convolutional body. Each sinc filter is controlled by two frequency parameters, allowing the learned passbands to be inspected directly in hertz while keeping the front end small. To reduce room-specific leakage, recording sessions and environments are separated before overlapping windows are assigned to the training, validation, and test partitions. We further use multi-objective Bayesian optimization as a design tool to examine the validation performance--model-size trade-off across 24 configurations. The selected designs span different operating points: the best-performing model achieves 80.2\% accuracy and 0.760 macro-F1 with 14{,}040 parameters, while the compact $N_f=25$ configuration uses only 2{,}848 parameters and achieves 75.7\% accuracy, 0.661 macro-F1, and 0.716 MCC on the held-out environment. Analysis of the learned filters and confusion patterns shows that spectral overlap contributes to confusion among water-related events, while the \textit{Door}/\textit{Walker/Crutch} errors also reflect similarities in their transient temporal structure.

Mon 28 SeptSoundArtificial IntelligenceMachine Learning
The gist
Listening to sounds in a bathroom can help assistive living by monitoring activities without video. The authors created a new dataset of bathroom sounds and designed a small, easy-to-understand computer program that listens to raw audio and recognizes seven different bathroom activities. Their model uses a special type of filter that can be directly inspected by frequency, making the model more interpretable. They showed their model can work well even when tested in different rooms.
Open → 2609.34907v1

New way to improve audio signals by focusing on sound levels

Prox-Friendly Log-Magnitude Prior on Complex-Valued Signal

Abstract: The logarithmic transform is essential in audio signal processing since human auditory perception is approximately logarithmic with respect to magnitude. However, directly incorporating prior knowledge about signals (e.g., harmonic structure) in the log-magnitude domain into optimization problems solved by standard proximal splitting algorithms remains challenging. To address this issue, this paper proposes a novel regularizer termed EPILOG (Exponential Penalty for Imposing priors on LOG-magnitude). EPILOG indirectly imposes prior knowledge on the log-magnitude of a complex-valued signal through regularization of an auxiliary variable that is shown to be linked with the log-magnitude. Furthermore, we derive its variable-wise proximity operators and develop a proximal splitting algorithm using these operators. Experiments on speech dereverberation demonstrate the effectiveness of the proposed regularizer, particularly in promoting cepstral-domain sparsity.

Mon 28 SeptSound
The gist
Audio sounds are naturally understood by humans through changes in loudness on a logarithmic scale. The authors found it hard to directly use this loudness information when improving audio using existing math tools. They introduced a new technique called EPILOG that helps apply prior knowledge about sound patterns in the loudness scale indirectly, making the math easier to solve. They showed it works well on improving recorded speech by reducing echoes.
Open → 2609.34445v1

WhisperVC-AV improves noisy whisper to normal voice conversion

WhisperVC-AV: Audio-Visual Content Restoration for Noise-Robust Whisper-to-Normal Voice Conversion

Abstract: Environmental noise obscures the content cues needed for whisper-to-normal conversion. We propose WhisperVC-AV, which restores acoustic content features while leaving the original WhisperVC conversion module unchanged. Its context-guided restoration module combines temporal acoustic context with synchronized lip features through attention and gated residual correction. Experiments on AISHELL6-Whisper show lower character error rates (CERs) than WhisperVC across three ASR systems, on clean speech and under all six signal-to-noise ratio (SNR) conditions with MUSAN noise. The largest gains occur at 0 dB SNR, where Qwen3-ASR CER falls from 36.29% to 28.88%. WhisperVC-AV also improves predicted speech quality while maintaining speaker similarity. The CER gains extend to unseen background noise without retraining, while visual controls support the use of utterance-specific lip cues. Audio examples are available on our demo page.

Sat 26 SeptSound
The gist
Turning whispered speech into normal talking voice is hard when there’s background noise because the sounds are harder to understand. The authors created WhisperVC-AV, which uses both the sound and lip movements to better guess what was said. This method helps computers recognize whispered words more accurately, even with lots of noise, without changing the main conversion process. It also keeps the speaker’s voice sounding natural.
Open → 2609.32843v1

Multimodal dataset improves attention tracking with brain and eye signals

MAESTRO: a Multimodal Auditory-attention Egocentric Speech-TRacking Open corpus

Abstract: Humans rely on gaze, head movements, and visual cues to attend to speakers in noisy environments, yet auditory attention decoding (AAD) has been studied primarily using electroencephalography (EEG). We introduce the Multimodal Auditory-attention Egocentric Speech-TRacking Open (MAESTRO) corpus, the first AAD dataset to simultaneously record EEG, eye gaze, pupillometry, egocentric video, and head inertial measurement unit (IMU) data. MAESTRO includes four competing speakers and background noise across multiple signal-to-noise ratio (SNR) conditions, enabling attention decoding under realistic listening scenarios. Through a four-speaker attention decoding benchmark, we show that combining behavioral and physiological signals improves decoding performance over EEG-only approaches, enabling future advances in multimodal auditory attention decoding. These findings open the door to new applications, analyses, and methodological advances in multimodal AAD. The complete dataset is publicly available at https://huggingface.co/datasets/aspire-osu/maestro-eeg-dataset . The official code repository is available at https://github.com/ASPIRE-OSU/MAESTRO .

Fri 25 SeptHuman-Computer InteractionMachine LearningSound
The gist
People use where they look, head movements, and seeing cues to focus on someone talking in noisy places, but most studies only use brain waves to follow attention. The authors created a new dataset called MAESTRO that records brain activity, eye gaze, pupil size, video from a wearer’s perspective, and head motion all at once with multiple people talking. They found that combining these different signals helps track who a person is listening to better than just using brain waves alone. This dataset and code are publicly available for others to build better listening technology.
Open → 2609.31898v1

Audio-visual method enhances target speaker voice with low delay

Adapting Personalized Speech Enhancement for Low-Latency Audio-Visual Target-Speaker Extraction

Abstract: Online audio-visual target-speaker extraction aims to remove competing voices while preserving speech quality and bounding lookahead. Existing extractors are built and evaluated for separation on synthetic mixtures, leaving listening quality and meeting behavior largely untested. We introduce Audio-Visual Personalized Voice Quality Enhancement (AV-PVQE), which approaches these requirements from the other direction. We start from a personalized speech enhancement model that reconstructs a requested voice at high quality but confuses the target in 46% of two-speaker mixtures despite clean enrollment. Adding mouth features at its speaker-conditioning input and jointly fine-tuning the visual and reconstruction networks reduces this rate to 1.6%, with no future frames and 20 ms of algorithmic delay. Compared with an online autoregressive audio-visual extractor, AV-PVQE yields separation gains on two synthetic benchmarks and larger gains on recorded meetings, and keeps its advantage on excerpts with more speakers than the fine-tuning mixtures. In personalized P.835 listening tests on two meeting corpora, it improves overall quality over this extractor by 0.57 and 0.63 MOS, with similar mean rating relative to the starting model. Preservation and rejection tests show that it keeps the target intact when no competing voice is present and suppresses competing speech when the target is absent.

Thu 24 SeptComputer Vision and Pattern RecognitionSound
The gist
When there are multiple people talking, it can be hard to focus on just one person's voice clearly. The authors designed a system that uses both sound and video (like watching mouth movements) to cleanly extract one speaker's voice from a mix of voices in real time. Their method improves the quality of the chosen speaker's voice and reduces confusion from other voices, doing this quickly with minimal delay. They tested it on real meeting recordings and listening tests, showing better results than previous approaches.
Open → 2609.30631v1

Sparse graph method improves audio-visual speech cleaning efficiency

G-Mamba: Sparse Graph-Guided Mamba for Audio-Visual Speech Enhancement

Abstract: Lightweight audio-visual speech enhancement (AVSE) models face a critical trade-off between computational efficiency and cross-modal alignment accuracy. While simple concatenation lacks relational expressiveness, dense cross-attention incurs computational overhead and is prone to unreliable cross-modal correspondence under strong acoustic interference. We propose Sparse Graph-Guided Mamba (SG-Mamba), a lightweight AVSE framework that integrates a sparse heterogeneous graph with a linear-complexity Mamba backbone. The graph explicitly models modality-specific relations through content-adaptive attention and cross-frame audio-visual connections, while Mamba captures long-range temporal context. We further introduce an audio skip connection to preserve spectral detail without sacrificing noise suppression. Evaluated on LRS3, SG-Mamba achieves competitive or superior performance against strong lightweight baselines and reaches 13.091 dB SI-SDR under noise-only condition. It also remains robust in cluttered multi-speaker conditions with a competitive cost of 3.45 G MACs (or 6.90 G FLOPs). Results on VoxCeleb2 further suggest that explicit structural priors improve robustness, generalizability, and computational efficiency in lightweight AVSE.

Wed 16 SeptComputation and LanguageSound
The gist
Cleaning up speech recorded in noisy places is hard, especially when using both sound and lip movement to understand speech. The authors created a lightweight system called SG-Mamba that uses a smart graph to connect sounds and images in a way that saves computing power while keeping good accuracy. Their method also keeps some original sound details to avoid losing voice quality. Tests show SG-Mamba works well even in tricky multi-speaker noisy settings and runs efficiently on modest hardware.
Open → 2609.18009v1