Predicting the Number of Downloads of Open Datasets by Naïve Bayes Classifier

Barbara Šlibar · TEM Journal · 2019

Nowadays, the use of Open Data has become more common and prominent, but there are a lot of questions regarding its quality. Most of the revised researches deal with the quality of Open Data portals, rather than estimation of the open datasets quality. Therefore, the main idea of this research is lowering to the level of the dataset itself in order to assess how much such data is downloaded by end users of Open Data portals on the basis of general dataset characteristics. A model for predicting the number of downloads of open datasets based on their general characteristics was constructed using the Naïve Bayes Classifier. Based on the obtained results, it is discussed if the certain dataset character is good predictor of open dataset downloading and to what extent.

Read the paper · More papers on PaperTik