Cognitively Robust Speech Diarization and Overlap Detection: An Advanced Model for Granular Analysis of Audio Segments
Ashin K Ajay, M. Prathilothamai · 2023
Multiple speakers conversing at once, or overlapping speech, poses a significant challenge in various speech analysis applications. The accuracy of handling this phenomenon greatly affects tasks such as speaker identification and speech recognition. Current methodologies categorize overlapping speech using individual techniques like MFCC (Mel Frequency Cepstral Coefficients), LFCC (Linear Frequency Cepstral Coefficients), and STFT (Short-time Fourier Transform), yielding certain results. In this research, we propose a comparative study that refines features using MFCC, LFCC, STFT, and wavelet transform, combined with a CNN (Convolutional Neural Network) model for classification. By capturing both time and frequency data, the Wavelet Transform holds promise for improving the precision of overlapping speech classification. To capture complex patterns and variations, the CNN model develops hierarchical representations. Our findings indicate that implementing MFCC in the CNN model yields superior results compared to LFCC. Additionally, the Wavelet transform outperforms the STFT transform in terms of performance and effectiveness. Furthermore, this paper places emphasis on the implementation of an algorithm that focuses on voice diarization and overlap detection. The objective is to accurately detect and categorize overlapping voice segments within an audio file, enabling a better understanding of its structure. The algorithm estimates the overall duration of each group by combining consecutive chunks with the same categorization, providing valuable insights into the temporal distribution of overlapping speech. Leveraging this data, the algorithm predicts the starting and ending periods of overlapping voice segments, facilitating the localization of overlaps.