Using Google's Keyword Relation in Multidomain Document Classification
Ping-I Chen, Shi-Jen Lin · International Journal of Data Mining & Knowledge Management Process · 2013
People can collect all kinds of knowledge from search engines to improve the quality of decision making, and use document classification systems to manage the knowledge repository.Document classification systems always need to construct a keyword vector, which always contains thousands of words, to represent the knowledge domain.Thus, the computation complexity of the classification algorithm is very high.Also, users need to download all the documents before extracting the keywords and classifying the documents.In our previous work, we described a new algorithm called "Word AdHoc Network" (WANET) and used it to extract the most important sequences of keywords for each document.In this paper, we adapt the WANET system to make it more precise.We will also use a new similarity measurement algorithm, called "Google Purity," to calculate the similarity between the extracted keyword sequences to classify similar documents together.By using this system, we can easily classify the information in different knowledge domains at the same time, and all the executions are without any pre-established keyword repository.Our experiments show that the classification results are very accurate and useful.This new system can improve the efficiency of document classification and make it more usable in Web-based information management.