Assessing the quality of classification models: Performance measures and evaluation procedures
Paweł Cichosz · Open Engineering · 2011
Abstract This article systematically reviews techniques used for the evaluation of classification models and provides guidelines for their proper application. This includes performance measures assessing the model’s performance on a particular dataset and evaluation procedures applying the former to appropriately selected data subsets to produce estimates of their expected values on new data. Their common purpose is to assess model generalization capabilities, which are crucial for judging the applicability and usefulness of both classification and any other data mining models. The review presented in this article is expected to be sufficiently in-depth and complete for most practical needs, while remaining clear and easy to follow with little prior knowledge. Issues that receive special attention include incorporating instance weights to performance measures, combining the same set of evaluation procedures with arbitrary performance measures, and avoiding pitfalls related to separating data subsets used for evaluation from those used for model creation. With the classification task unquestionably being one of the central data mining tasks and the vastly increasing number of data mining applications — not only in business, but also in engineering and research — this is expected to be interesting and useful for a wide audience. All presented techniques are accompanied by simple R language implementations and usage examples, which — whereas created to serve the illustration purpose mostly — can be actually used in practice.