An Innovative Method for Improving Speech Intelligibility in Automatic Sound Classification Based on Relative-CNN- RNN
Alkawati Magadum, Monica Goud, Gopinath V, P. Prabbu Sankar, B Sivadharshini, M. Ananthi · 2023
Sound engineers and producers establish the speech-to-background ratio (SBR) during mixing depending on the rules of thumb and their own ears. However, there is no assurance that the general public will be able to understand the voice content. This research introduces a method for automatically selecting the optimal SBR for a scenario based on an objective intelligibility metric. For several types of ambient noise, the SBR estimated by the model is necessary to achieve a minimal intelligibility level when compared to the SBR selected by listeners. The model estimation was found to benefit from an additional gain even for normally hearing listeners. The proposed method involves three stages: preprocessing, feature extraction, and model training. It employs a FIR filter for preprocessing and further the spectral centroid, roll-off, flux, and ZCR for feature extraction. R-CNN -RNN is used for training the model. The proposed method outperforms the CNN and CNN-RNN models.