Audio Signal Analysis and Recognition of Bengali Alphabets: A Comparative Study of Machine Learning Approaches
Jubaida Quader Jerin, M. Akhtaruzzaman · 2024
This research focuses on signal pattern analysis for Bengali alphabet recognition in audio, utilizing established machine-learning approaches. It addresses inherent linguistic challenges, presents insights derived from the collected data, and discusses benchmark results. The primary focus is on processing raw Bengali audio files through advanced signal pattern analysis for efficient feature extraction and adapting machine learning models for Bengali alphabet recognition. Specific techniques, such as windowing and overlap-add, were applied to address the unique characteristics of Bengali vowels. Feature extraction methods include Root Mean Square Energy, Zero Crossing Rate, and Mel-frequency Cepstral Coefficients (MFCCs). In experimental settings, MFCCs consistently demonstrated superior performance compared to other methods. Various machine learning models, including Linear Regression, MLP Classifier, SVM, and LSTM, were employed, with MFCCs consistently showing enhanced performance for Bengali alphabet recognition. Future research will focus on advancing automatic speech recognition for Bengali alphabets, with the goal of seamless integration into embedded systems, such as Arduino, for practical applications based on raw audio data. Additionally, this study explores ensemble learning techniques for Bangla phoneme identification, aiming to improve the robustness and accuracy of classification systems.