On Metrics for Keyword-Based Document Selection and Classification
Mikhail Alexandrov, Alexander F. Gelbukh, Pavel Makagonov · 2010
For a large set of documents, certain numerical characteristics (metrics) are discussed that allow to select the documents relevant to a given topic and divide the set of the relevant documents into several groups (clusters) reflecting various subtopics of the given topic. The choice of the metrics is justified by expected results for known examples. A given topic is defined by a domain-oriented keyword dictionary. The results are implemented in a program Text Classifier.