A middle-ware approach to leverage the distributed data de-duplication capability on HPC and Cloud storage systems
Hsing‐Bung Chen, Sihai Tang, Song Fu · 2020
The unprecedented growth in the volume and diversity of the data in today's HPC and Enterprise computing environment has posted challenging problems on data management and data space reduction. More than 71% of enterprise and HPC communities are seeking de-duplication technologies to reduce cost and increase the storage efficiency. The importance of applying data de-duplication techniques is critical for active research and development. Current implementations of data de-duplication systems are mainly hardware dependent, system dependent, and platform dependent. Also, most of these implementations are proprietary software and not in open source domains. In this paper, we present a new middle-ware design and implementation approach, named D3M, to support distributed data de-duplication feature on existing file and object storage systems. We also incorporate this proposed D3M middle-ware with the Redhat's Linux device layer de-duplication and compression driver, called VDO (Virtual Data Optimizer). With these two layers of data de-duplication support, we accommodate both client side and server-side data de-duplication features. Finally, we conduct various testing cases on HPC data sets and Enterprise data sets to illustrate the benefits and advantages of applying our bilayer data de-duplication middle-ware solution.