A Ontology-based Document Feature Extraction

Lin Dong · 2008

To effectively reduce the dimension of document vectors,we introduce a novel method employing domain ontology to extract feature concept. For all document categories,all raw words in each category are mapped to concepts in their relative concept tree derived from the domain ontology. At the same time the frequency of raw words is transformed into the frequency of concepts. Experimental results show that this method can effectively reduce the dimension of document vectors without loss of categorization accuracy,compared with traditional document vectors.

Read the paper · More papers on PaperTik