POFFO: A Perceptual Online File Fingerprint Offloading Strategy for Effective Data Deduplication at Cloud-Edge Systems
Hexian Lu, Yuhui Deng, Jiande Huang · 2024
Edge servers usually store collected data in cloud servers and use deduplication techniques to remove redundancy. However, edge servers can also perform deduplication during data collection. This requires transferring fingerprints from the cloud servers to the edge servers for assistance. Since the large volume of fingerprint data on the cloud server, for example, 1 PB of data corresponds to 8 TB of fingerprints, transferring all fingerprints to the edge servers is impractical. Therefore, we propose a fingerprint offloading strategy. Only a small amount of fingerprints and data chunks needed for edge deduplication are offloaded from the cloud server to the edge server, enabling cloud-edge collaborative deduplication. The general process is as follows: First, the edge server collect a large amount of data from various devices, divides the data into chunks, and calculates fingerprints to identify unique data chunks and fingerprints. Then, the edge server upload the fingerprints to the cloud server. The cloud server check the fingerprints and offload the data chunks corresponding to existing fingerprints back to the edge server. Upon receiving these data chunks, the edge server performs thorough deduplication. Finally, the deduplicated data chunks are uploaded to the cloud server, ensuring that only unique data is transmitted to the cloud server. Experiments used chunks ranging from 1 KB to 16 KB, with an average size of 4 KB, and employed three real backup datasets. The results showed that the edge server computation time was reduced by 48.1%, metadata storage was reduced by 98.1%, and the upload volume from the edge server to the cloud server decreased by 87.1%. The size of fingerprints and data chunks offloaded to the edge server ranged from 9.4 MB to 1624.2 MB.