Research on the Application of Artificial Intelligence and Distributed Parallel Computing in Archives Classification
Erbo Shang, Xiaohua Liu, Hailong Wang, Yangfeng Rong, Yuerong Liu · 2019
In the traditional mode, the classification of archives often has large workload, long time consumption and high labor cost. With the development of archives management towards informationization and paperlessness, it has become an important research topic to classify archives by using artificial intelligence classification method based on machine learning. In this paper, an archives classification method based on XGBoost and Spark distributed parallel computing is proposed. XGBoost algorithm can continuously improve the classification accuracy of small class samples during training rounds, which makes XGBoost algorithm has certain advantage in the classification of archives data with large samples and unbalanced class distribution. Spark distributed parallel computing system can greatly improve the computational efficiency of XGBoost algorithm and reduce the training time of the archives classification model. The simulation results show that our method has the advantages of faster model training and higher classification accuracy than the traditional method, and can satisfy the complex and huge archives classification task.