On the impact of data sampling on hyper-parameter optimisation of recommendation algorithms

Matteo Montanari, Cesare Bernardis, Paolo Cremonesi · Proceedings of the 37th ACM/SIGAPP Symposium on Applied Computing · 2022

Hyper-parameter optimisation (HPO) is a fundamental task that must be performed in order to achieve the highest accuracy performance that a recommendation algorithm can provide. In the recent past, with the the growth of dataset sizes, the amount of resources and time needed to perform the optimisation dramatically increased. Sampling the data used during the HPO procedure allows reducing the required resources, but it impacts the accuracy metric score. In this paper, we study the effects of optimising the hyper-parameters through a random search, sampling the users in a dataset. The results of our experiments show that sampling reduces the amount of time needed to conduct HPO, but it also influences differently the accuracy of the best configuration found by HPO, depending on the algorithm optimised and the dataset selected.

Read the paper · More papers on PaperTik