Why Simple Models Perform Well in Predicting Popularity for Caching?

Jiajun Wu, Chenyang Yang · 2019

Predicting file popularity is an important task for proactive caching, which has been shown a promising way to support the explosive increase of data traffic. Both linear and non-linear predictors have been proposed in literature to predict file popularity, and shallow neural networks have been shown to achieve good performance for caching. In this paper, we strive to explain why simple models perform well in predicting dynamic popularity with the historical number of requests. We consider linear regression, deep neural networks, random forest, support vector regression, and a persistence model for predicting popularity. We employ MovieLens and Youku datasets for analyzing the time-varying pattern of popularity and for evaluating the caching performance. We show that the cache hit ratios achieved by the caching policy with predicted popularity using linear models are close to that using non-linear models. We interpret the observation by first proving that deterministic popularity with typical profiles can be predicted with linear models and then showing that majority of the popular files are with these profiles in the real dataset.

Read the paper · More papers on PaperTik