Big data skipping in the cloud

Oshrit Feder, Guy Khazma, Gal Lushi, Yosef Moatti, Paula Ta-Shma · 2019

According to today's best practices, cloud compute and storage services should be deployed and managed independently. However, this generates a problem for big data analytics in the cloud: potentially huge datasets need to be shipped from the storage service to the compute service to analyse the data. To address this, minimizing the amount of data sent across the network is critical to achieve good performance and low cost. Data skipping is a technique which achieves this for SQL style analytics on structured data.

Read the paper · More papers on PaperTik