An Efficient Deep Learning framework with CNN and RBM for Native Speech to Text Translation
Sadineni Neelima, Kalyankumar Dasari, Annemneedi Lakshmanarao, Peluru Janardhana Rao, Madhan Kumar Jetty · 2024
This paper presents a novel approach to native speech-to-English text translation using deep learning techniques. Traditional speech-to-text systems often face challenges with accurate translation, especially in noisy environments. Our proposed method leverages Convolutional Neural Networks (CNNs) and Restricted Boltzmann Machines (RBMs) to enhance translation accuracy. Initially, raw audio data is collected under diverse conditions and converted into spectrograms, which are then preprocessed and normalized. The spectrograms are transformed into matrices for feature extraction. The CNN is used to recognize and classify these features, while the RBM extracts high-level features from the matrices, improving the model’s ability to interpret complex audio patterns. The features are fed back into the CNN for refined analysis. The model is evaluated on the THUYG-20 SRE dataset under various noise conditions, demonstrating significant improvements in translation accuracy. The final output includes transcribed English text alongside speaker recognition results and performance metrics, offering a robust solution for effective communication across languages.