Knowledge Representation, Learning, and Reasoning in WebDoc - A Web Document Classification System
Bo Tang, Julia E. Hodges · 2000
This paper describe a novel approach to knowledge representation, learning, and reasoning in WebDoc, a system that classifies Web documents according to the Library of Congress classification system. We argue that an automatically constructed domain-independent knowledge base is indispensable. The WebDoc system builds a knowledge base (represented as a semantic network) that contains the Library of Congress subject headings and their relationships. Through training on human-indexed and NLP-parsed Web documents, WebDoc modifies the semantic network and generates rules for future index generation tasks. Introduction The rapid growth of the World Wide Web makes a tremendous amount of information available to people who have access to a computer connected to the Internet. However, there is still a long way to go from simply having access to really taking advantage of the information. People often get lost rather than enlightened due to the lack of efficient Web information ret...