Turbo Table: A Semantic-Aware Cache Acceleration System
Hao Yuan, Chentao Wu, Jie Li, Minyi Guo · 2024
In the current storage disaggregation architecture, the challenge of quickly retrieving data from storage clusters is typically addressed using caching or data pushdown strategies to accelerate data access and reduce data movement. However, most caching systems still manage data at the page granularity level. For compute-intensive applications, such as analytical applications, this coarse data management approach leads to underutilized computational resources and read-write amplification issues. Additionally, insufficient cache utilization results in inefficient data flow. By managing data at the schema granularity level and modifying the data flow path, we alleviate the read-write amplification problem and improve data transfer speeds. Managing data at the schema level also enables semantic awareness, allowing us to proactively analyze semantics for more precise cache management strategies and accurate prefetching, rather than passively waiting for cache request sequences. We also observed that native compute caches waste valuable cache space and complicate the association between original and result data. To address this, we propose a multi-grained caching model to avoid these limitations. Compared to the baseline, Turbo Table reduces computation time by 2.4% to 36.8%.