Crowdsourced data analytics: A case study of a predictive modeling competition
Yukino Baba, Nozomi Nori, Shigeru Saito, Hisashi Kashima · 2014
Predictive modeling competitions provide a new data mining approach that leverages crowds of data scientists to examine a wide variety of predictive models and build the best performance model. Competition hosts, who provide their own dataset and specify the problem to be solved, are not only able to obtain the best model from among those submitted but also to aggregate the submitted models to obtain one that outperforms the rest. In this paper, we report the results of a study conducted on CrowdSolving, a platform for predictive modeling competitions in Japan. We hosted a competition on a link prediction task and observed that (i) the prediction performance of the winner significantly outperformed that of a state-of-the-art method, (ii) the aggregated model constructed from all submitted models further improved the final performance, and (iii) the performance of the aggregated model built only from early submissions nevertheless overtook the final performance of the winner. Our results show the power of crowds for predictive modeling, not only in the quality of the obtained model, but also in its speed to achieve it. Furthermore, they demonstrate the possibilities of combining human insights and machine learning in data analytics.