Data Deduplication Using Python

Harshada Deshingkar, Preeti Desai, Shrinidhi Deshmane, Aarya Deshmukh, Tanuja Desai, Sankhya Desai, Ganesh Ubale · 2023

This paper aims to develop a data deduplication technique for optimizing storage utilization in enterprise environments and for general users. Data deduplication is a process that identifies and eliminates duplicate data from storage systems, resulting in reduced requirements and lower costs. The project involves working on duplicate files from the folders present in storage systems and processing them to reduce the redundancy of data. The results of the project will provide valuable insights on usage of data deduplication technique to delete duplicate files. This method can be used by users to optimize their storage utilization and reduce their storage costs. Additionally, the project will contribute to the field of data storage and management by providing a way to reduce the constant storage of duplicate files.

Read the paper · More papers on PaperTik