Comparative Study of Various Machine Learning Approaches for Marathi Dialect Recognition
Ashwini G Pawar, Ajay S. Patil, Nita V Patil · Cureus Journal of Computer Science. · 2025
Identifying dialects accurately in the digital age to preserve linguistic diversity has become increasingly important. This research focuses on developing a strong Marathi dialect identification system that uses sound data and artificial intelligence techniques. Various acoustic features such as spectral centroid, bandwidth, roll-off, root mean square energy, zero-crossing rate, and mel-frequency cepstral coefficients were extracted from audio recordings representing distinct Marathi dialects spoken in different regions. The dataset was preprocessed by dividing the recordings into frames that can be analyzed and normalizing audio formats. Several machine learning algorithms were utilized to classify the dialects, such as K-nearest neighbors, decision trees, naïve Bayes, AdaBoost, and gradient boosting. Multiple metrics were used to compare the performance of the algorithms, which resulted in the best output at 94.56% accuracy using the gradient boosting algorithm. Precision, recall, and F1-score were also used to evaluate the model's performance, where gradient boosting outperformed other methods in all cases. This work demonstrates the potential of advanced machine learning techniques in dialect recognition, especially for resource-scarce languages like Marathi. It opens the way for further improvements in natural language processing applications.