Fine-grained partitioning for aggressive data skipping
Liwen Sun, Michael J. Franklin, Sanjay Krishnan, Reynold Xin · 2014
Modern query engines are increasingly being required to process enormous datasets in near real-time. While much can be done to speed up the data access, a promising technique is to reduce the need to access data through data skipping. By maintaining some metadata for each block of tuples, a query may skip a data block if the metadata indicates that the block does not contain relevant data. The effectiveness of data skipping, however, depends on how well the blocking scheme matches the query filters.