Multilabel Associative Classification Categorization of MEDLINE Articles into MeSH Keywords An Intelligent Data Mining Technique to More Accurately Classify Large Volumes of Documents
Rafał Rak, Lukasz Kurgan, Marek Reformat · 2007
The constantly increasing numbers of scientific documents and the extensive manual work associated with their description and classification requires intelligent classification capabilities for users to find required information. The article discusses an automated method for classification of medical articles into the structure of document repositories, which would support currently performed extensive manual work. Understanding the Problem Exponential growth of the number of scientific documents results in increased difficulty in their categorization. This motivates our research in providing intelligent categorization methods. MEDLINE is the National Library of Medicine’s (NLM) database consisting of approximately 13 million article references to biomedical journal articles dating back to 1966. NLM employees add approximately 1,500 to 3,500 new article references every day to this database. This includes manual assignment of each article to the corresponding entries in the medical subject headings (MeSH). MeSH is NLM’s controlled vocabulary thesaurus consisting of medical terms at various levels of specificity. The rapidly growing number of incoming documents together with error-prone manual work may make this task difficult.