The study on automitic classification of digital documents of scientific papers
Sen Li · Journal of Shandong University · 2006
Since scientific papers are usually semi-structural documents,a hierarchy classification model based on the metadata of scientific papers is proposed,where the metadata include the titles,sets,abstracts and so on.Experiments show the precision of the classification based on the metadata of papers is close to that of the classification based on the full text of papers.Furthermore,the classification precisions are better than the best known classification algorithm if the papers are classified based on taxonomy of application domains as follows: first,the metadata are used to classify paper roughly based on the higher levels of taxonomy,then full texts are utilized to classify these papers on the lower levels of taxonomy.Since the size of metadata is less than that of full text and the number of papers classified in a subclass is less than that of total number of papers,the new model enhances the efficiency of paper classification when the number of classes is bigger and the documents are distributed averagely in the given taxonomy