MapReduce based for speech classification
Quang Trung Nguyen, The Duy Bui · 2016
Speech classification is one of the most vital problems in speech processing as well as spoken word recognition. Although, there have been many studies on the classification of speech signals, the results are still limited on both accuracy and the size of the vocabulary. When classifying a huge volumes vocabulary, the speech classification becomes more and more difficult. Today, there are some frameworks that allow working with big data. One of these is a data mining utility. It can perform supervised classification procedures on very large amounts of data, usually named as big data, on a distributed infrastructure by using the MapReduce framework of Hadoop clusters. This tool has four classification approaches implemented. These are Random Forest, Naïve Bayes, Decision Trees and Support Vector Machines (SVM). All these approaches require input data having the same size, so the input data must be quantized before using. This leads to decrease the accuracy in the classification stage. In this paper, we propose an implementation of Local Naïve Bayes Nearest Neighbor based on Hadoop framework, which allows input data with different sizes and works well with huge training data.