Density estimation-based document categorization using von Mises-Fisher kernels
Andrew A. Skabar, Shahmeer Rafiq Memon · 2010
Although classifiers such as Support Vector Machines (SVMs) and k Nearest Neighbors (kNN) are able to achieve excellent document classification performance as measured by common information retrieval measures such as F1 score and breakeven point, they do not produce output values that can reliably be interpreted as probabilities. This means that without additional post-processing of their outputs, these classifiers are not capable of answering a question as simple as “To which classes does a document belong with a minimum probability of 80%?” This paper presents a density estimation-based classification technique which outputs probabilities directly. The technique is based on estimating densities using von Mises-Fisher kernels, and combining these under Bayes' Theorem to arrive at posterior probabilities of class membership. Results of applying the technique to the Reuters-21578 dataset show that the technique is computationally feasible, that its classification performance as measured by F1 score is comparable to that of SVMs and better than that of kNN classifiers, and that the output values can be interpreted as well-calibrated probabilities.