The Effect of the Distribution of Predictions of User Models.
Eric Van Inwegen, Yan Wang, Seth Adjei, Neil Thomas Heffernan · Educational Data Mining · 2015
We hypothesize that there are two basic ways that a user model can perform better than another: 1.) having test data averages that match the prediction values (we call this the coherence of the model) and 2.) having fewer instances near the mean prediction (we call this the differentiation of the model). There are several common metrics used to determine the goodness of user models; these metrics conflate coherence and differentiation. We believe that user model analyses will be improved if authors report the differentiation, as well as to include an ordering metric (e.g. AUC/A’ or R) and an error measurement (Efron’s R, RMSE or MAE). Lastly, we share a simplified spreadsheet that enables readers to examine these effects on their own datasets and models.