Big data skipping in the cloud
Oshrit Feder, Guy Khazma, Gal Lushi, Yosef Moatti, Paula Ta-Shma · 2019
According to today's best practices, cloud compute and storage services should be deployed and managed independently. However, this generates a problem for big data analytics in the cloud: potentially huge datasets need to be shipped from the storage service to the compute service to analyse the data. To address this, minimizing the amount of data sent across the network is critical to achieve good performance and low cost. Data skipping is a technique which achieves this for SQL style analytics on structured data.