Exploring Time and Frequency Domain Loss Functions in Real-Time Speech Enhancement Using an Efficient Neural Network: A Comparative Study
Amir Rajabi, Mohammed Krini · 2024
Speech enhancement is a fundamental objective in audio signal processing, aiming to improve the clarity and intelligibility of noisy signals. The primary goal is to enhance speech signals while carefully considering the trade-off between algorithmic complexity and performance improvement. The computational complexity of speech enhancement algorithms is a fundamental concern, especially in real-time applications due to limited hardware computational capacity. Numerous deep learning methods have been suggested to address this challenge. Despite their computational complexity, these algorithms, when combined with novel loss functions, tend to exhibit superior performance. This is attributed to their ability to learn intricate relationships between noisy and clean speech signals. In conclusion, the selection of appropriate deep learning-based enhancement algorithms and suitable loss functions, whether in the time- or frequency-domain, is essential for achieving optimal speech enhancement results. This paper emphasizes the significance of making informed choices in methods and loss functions to accommodate computational advancements and meet the demands for improved speech quality. In this study, we adopted a complexity-efficient model from our previous work and trained it with loss functions selected for their demonstrated success in improving speech quality. Furthermore, experimental findings demonstrate that the frequency-domain-based loss function, which combines the absolute and complex magnitudes, achieves outstanding performance across the evaluation metrics defined in this paper, such as Perceptual Evaluation of Speech Quality (PESQ), Deep Noise Suppression Mean Opinion Score (DNSMOS), and others. It significantly outperforms alternative time-domain-based approaches.