Efficient Data Security Using Predictions of File Availability on the Web

Kevin Saric, Gowri Sankar Ramachandran, Raja Jurdak, ‪Surya Nepal‬ · 2024

As we approach the physical limits of storage density, digital storage prices are no longer plummeting, despite the lingering belief that they still are. Meanwhile, data production continues to grow, making it harder to securely manage the data we produce. Typical digital storage media is often consumed by a small number of large files that are widely available on the web. If the availability of files on the web could be predicted, the choice between consuming local storage resources or simply redownloading the file in the future could be automated, thus increasing the efficiency of backup and encryption workflows. Through a large-scale analysis of hundreds of billions of crawl URLs spanning 8 years, as well as over 60 million HTTP header request responses from web servers, we explore the requirements and design of a framework for such predictions. It includes a data structure for efficiently representing the lateral/longitudinal availability of files and an extensible mathematical model for fast and adaptable prediction calculations. Additionally, we contribute novel observations about file availability on the web, including the identification of a period of initial volatility in their lifespans. Analysis indicates that a pool of 2,500TB of distributed, popular files is freely and predictably available to users, offering opportunities to reduce the storage and computational costs of both backup and encryption.

Read the paper · More papers on PaperTik