Automatic extraction of paper research methods based on multi-strategy

Zuo Yang, Zhenghan Chen, Hua Zhao · 2019

With the deeply advancement of information technology, the number of research papers has become increasingly prominent with the trend of geometric growth. How to obtain effective information in a million-level papers has become a primary problem faced by researchers. The paper propose to extract the research methods used in these papers based on the analysis of paper titles, abstracts and keywords. Firstly, using crawler to crawl the data of the doctoral thesis and preprocess it. Secondly, adopting the multi-strategy methods which are naive Bayesian and the regular expression to extract the research methods. When naive Bayes is used to extract the paper, the research method is extracted to form a model. Through the calculation of the paper text by model, the most probable research method is output. When using Regular Expressions for extraction, the Chinese text segmentation is performed, and then matched by a custom vocabulary. Experimental results show that when the multi-strategy extraction methods are used to extract the research method, the accuracy rate is satisfactory.

Read the paper · More papers on PaperTik