Recording word position information for improved document categorization

Piotr Gawrysiak, Łukasz Gancarz, Michal J. Okoniewski · 2002

In this paper, which is a report from work in progress, we briefly present the new document representation that could be used in classic text mining applications, such as document categorization. We briefly present most popular unigram and n-gram document representations that are used frequently in text mining research, mention their shortcomings, and present the idea of representation that records not only word count information, but also word position within document. As experiments are still underway, we do not present final result, but only mention types of tests that are being done.

Read the paper · More papers on PaperTik