Research of Text Categorization Based on Sparse Autoencoder Algorithm
Qin Sheng-ju · Science Technology and Engineering · 2013
Tradition text classification algorithms use the expected cross entropy,information gain and mutual information statistical method to get the feature set,but these methods require setting thresholds.If the training data set is large which prone to feature items is not clear,the feature information loss and other defects.In order to solve the above problem,the sparse autoencoder algorithm is used which belongs to learningautomatically extracts text features,and then combines with the deep belief networks to form SD algorithm for text classification.Experiments show that,in the case of small training set,SD algorithm performs lower than traditional support vector machines,but when dealing with high-dimensional data,SD has higher accuracy and recall rate than support vector machine algorithm.