Categorizing text and detecting passages and category relationships
Nazli Goharian, Saket S. R. Mengle · 2009
With the increasing number of digital documents, the ability to automatically classify those documents both efficiently and accurately is becoming more critical and difficult. Text classification is used to automatically assign predefined categories to documents that reflect the overall contents of the document. We explore various issues related to text classification. A key problem in text classification is the high dimensionality of feature space. We propose the Ambiguity Measure (AM) feature selection algorithm, which selects features whose presence in a document indicate a strong degree of confidence that a document belongs to only one specific category. We also evaluated our methodology to detect passages of specified categories within documents. Passages can be hidden within a text to circumvent their disallowed transfer. Hence, detecting the presence of such passages in documents is of concern to all corporate and governmental organizations. The proposed keyword based dynamic passage methodology to split documents into passages is favorably compared to three existing document splitting techniques. Furthermore, we evaluated our approaches for effectively identifying relationships among document categories. Knowledge of relationships among categories is useful in various domains such as recommendation systems, news feeds and information security. Our novel method capitalizes on the misclassification results of both document and passage classifiers to identify potential relationships among categories. We also utilize association rule mining to identify relationships among categories.