Deduplication of Image Files to Reduce the Redundancy in Cloud Storage

G Ragul., N Nithyashree., K K Harini, A. Prabhu Chakkaravarthy · 2024

Efficient cloud storage management is critical as the volume of digital content grows exponentially. Redundant image files contribute significantly to storage inefficiencies, increasing costs and affecting retrieval performance. This project proposes a scalable deduplication system for cloud storage, focusing on reducing redundancy, optimizing storage space, and enhancing operational efficiency. The objective is to ensure that only unique image files are stored, while duplicates are replaced with references to existing files, thereby conserving resources and improving data management. The proposed system utilizes a combination of SHA-1024 hashing for precise detection of identical files and perceptual hashing algorithms to identify visually similar images. These algorithms ensure comprehensive identification of duplicates, including nearidentical files. The process begins with image ingestion, where unique hashes are generated for each file. These hashes are compared against an existing metadata store to detect duplicates. Identified duplicates are stored as references, reducing storage space requirements, improving retrieval speed, and minimizing bandwidth usage by avoiding repeated transmissions. The deduplication system also enhances backup and disaster recovery processes by reducing the volume of files to be managed, resulting in faster backups and recovery times. By integrating this framework with cloud environments, such as Azure, the solution delivers scalability and real-time operation. This work addresses the growing challenges of data redundancy in cloud storage, contributing to efficient and sustainable storage practices.

Read the paper · More papers on PaperTik