GAP-DCCS: A Generic Acceleration Paradigm for Data-Intensive Applications With Efficient Data Compression and Caching Strategy Over CPU-GPU Clusters
Jiangwei Xiao, Yingzhe Bai, Hanfei Diao, Guofeng Liu, Yuzhu Wang · IEEE Transactions on Parallel and Distributed Systems · 2025
Seismic exploration is a geophysical method used for imaging subsurface structures, capable of providing high-resolution images of the underground. In seismic data processing, Kirchhoff Pre-Stack Depth Migration (KPSDM) serves as one of the key techniques, playing a critical role in significantly enhancing the lateral resolution of imaging and providing accurate characterization of subsurface media. However, with the continuous growth in high-density seismic data volumes, the computational efficiency of KPSDM is primarily constrained by substantial computational loads, end-to-end I/O bottlenecks, and data storage pressures. To address the performance optimization challenges of computation-intensive applications that require frequent large-scale data transfers between the host and accelerator devices, this paper proposes GAP-DCCS, a GPU-based Generic Acceleration Paradigm with efficient Data Compression and Caching Strategy, which includes the following core strategies: (1) For compute-intensive modules, a GPU-based three-dimensional parallel acceleration is implemented, combined with memory access optimization techniques and overlapping strategies for data transfer and computation, to improve GPU resource utilization; (2) To alleviate the storage pressure of large-scale datasets, the BitComp compression algorithm is introduced to efficiently compress task data while maintaining output stability, significantly reducing storage requirements and end-to-end data transfer volume; (3) To tackle the I/O bottleneck caused by frequent large-scale data transfers between the host and devices, an adaptive dynamic caching data management mechanism is designed to effectively increase data reuse rates and markedly reduce end-to-end transfer frequency. Experimental results demonstrate that the proposed optimization method significantly enhances the computational performance of KPSDM, achieving a speedup of 123.51× on a single NVIDIA Tesla A800 GPU compared to a 16-core CPU. This optimization paradigm has not only been effectively validated in KPSDM but also offers a referable high-performance computing solution for other large-scale data processing tasks.