Multimodal Speaker Recognition: Combining FFT, CNN, Speech-to-Text, BERT-Based Punctuation Restoration and Sentence Correction
Kasula Pavan Sai, K. Ajay, Ramesh Mande · 2023
In this paper, we look at technological solutions to problems with speech recognition, text-to-speech conversion, sentence correction, and punctuation restoration. Convolutional neural networks (CNNs) for model development, the Fast Fourier transform (FFT) for feature extraction from data, and the potent Bidirectional Encoder Representations from Transformers (BERT) model for context-aware sentence correction are all included into our methodology. Our study intends to enhance the precision, context sensitivity, and usability of language processing systems by merging CNNs, FFT, and BERT. By providing excellent transcription and precise sentence correction, this holistic approach pushes the limits of audio-based applications.