New way to improve audio signals by focusing on sound levels
Prox-Friendly Log-Magnitude Prior on Complex-Valued Signal
Sound
Summary
Audio sounds are naturally understood by humans through changes in loudness on a logarithmic scale. The authors found it hard to directly use this loudness information when improving audio using existing math tools. They introduced a new technique called EPILOG that helps apply prior knowledge about sound patterns in the loudness scale indirectly, making the math easier to solve. They showed it works well on improving recorded speech by reducing echoes.
What this means in practice
- •For audio engineers: Improve software-based reduction of reverberation in recorded speech to enhance clarity in real-world environments.
- •For hearing aid designers: Refine algorithms that process sound magnitude logarithmically to better mimic human hearing perception and improve speech intelligibility.
Authors
Kazuki Matsumoto, Keidai Arai, Kohei Yatabe
Abstract
The logarithmic transform is essential in audio signal processing since human auditory perception is approximately logarithmic with respect to magnitude. However, directly incorporating prior knowledge about signals (e.g., harmonic structure) in the log-magnitude domain into optimization problems solved by standard proximal splitting algorithms remains challenging. To address this issue, this paper proposes a novel regularizer termed EPILOG (Exponential Penalty for Imposing priors on LOG-magnitude). EPILOG indirectly imposes prior knowledge on the log-magnitude of a complex-valued signal through regularization of an auxiliary variable that is shown to be linked with the log-magnitude. Furthermore, we derive its variable-wise proximity operators and develop a proximal splitting algorithm using these operators. Experiments on speech dereverberation demonstrate the effectiveness of the proposed regularizer, particularly in promoting cepstral-domain sparsity.