A Metadata Extraction Approach from Papers Based on Meta-learning
Fuzhi Zhang · Journal of Information and Computational Science · 2013
To improve the accuracy of paper metadata extraction, a paper metadata extraction approach based on meta-learning is presented. Firstly, we propose a construction method of base-classifiers, which combines the Support Vector Machine (SVM) with the created diverse base-level training sets to construct some base-classifiers with larger diversity. The training sets are created according to the Open Access (OA) journal categories. Secondly, we present a paper metadata extraction algorithm based on meta-learning, which uses the meta-classifier to integrate the classification results of the base-classifiers and generates the final extraction results. The proposed approach guarantees extraction accuracy and independence of the base-classifiers, increases the diversities of them, and improves the classification performance of the meta-classifier, as well as reduces the degree of correlation of misclassifications. The experimental results show that our approach is superior to other single machine learning algorithm and the accuracy of paper metadata extraction is improved.