Utility-based Performance Measures for Regression
Rita P. Ribeiro, Porto La, R. Ceuta, R. Campo Alegre, Lu ́ õs Torgo, Roberto Frías · 2008
Costs and benefits are of key importance for most real-world data mining applications. The main work carried out within ML/DM on cost-sensitive learning has been centered on classification tasks. Nevertheless, there are many real-world regression applications where costs and/or benefits of the predictions play an important role. On such applica-tions, the costs and/or benefits of a predic-tion can vary across the domain of the target variable. For instance, it may be much more relevant to be accurate at a certain range of values than on other parts of the target vari-able domain. In this context, it is impor-tant to address regression tasks from a cost-sensitive perspective. Using as an example the case of rare extreme values prediction, we show why the standard regression error metrics are no longer effective for this type of applications. We propose an utility-based evaluation framework, which allows for dif-ferentiated scores to be assigned to the pre-dictions based on their cost/benefit in con-formity with the application preference bi-ases. Based on this utility concept, new per-formance metrics can be developed for a more reliable evaluation/comparison of models and also for model development. 1.