Spectral Mapping of Lombard Speech to Normal Speech: A Deep Learning Approach to Improved Lombard Effect Compensation
S. Uma Maheswari, A. Shahina, A. Nayeemulla Khan · Fluctuation and Noise Letters · 2025
Lombard speech (LS) causes a degradation in the performance of speech-based systems that are built using Normal Speech (NS). This loss in recognition rate can be attributed to the acoustic-phonetic differences that exist between Lombard speech and normal speech. In order to improve the performance of the existing NS-based speech systems, while using Lombard speech, these differences need to be compensated. Several techniques have been proposed in the literature. This work proposes a method for improving the Lombard effect compensation by modeling the nonlinear functional relationship between the spectral features of weighted Linear Prediction Cepstral Coefficients (wLPCCs) extracted from LS and NS. The wLPPCCs of LS are mapped to those of the NS using three different deep learning architectures, namely Deep Neural Network (DNN), Long Short-Term Memory (LSTM) and Bidirectional LSTM (BLSTM), with different network sizes. The effectiveness of the spectral mapping is objectively evaluated using (a) The Itakura–Saito distance metric and (b) a speaker identification system (trained using NS features and tested using mapped features). The speaker recognition systems based on the mapping outputs from DNN, LSTM and BLSTM achieved an accuracy of 98.6%, 99.7% and 99.6%, respectively. Among different network architectures of different sizes, LSTM and BLSTM-based spectral mapping give the best Lombard effect compensation.