Efficient and Secure Data Deduplication in Distributed Cloud Storage: A Fault-Tolerant Model Using Active Learning for Big Data Management
S Seethalakshmi, B. Balakumar · International Journal of Enhanced Research In Science Technology & Engineering · 2024
As the big data era advances, more and more redundant data are being expressed in various ways. Data deduplication knowledge has never been more important than it is now for lowering redundant data storage and enhancing data quality. Connecting several data tables and identifying distinct entries that point to the same item is typically required, particularly when multi-source data deduplication is involved. Active learning minimises the amount of data that needs to be annotated and trains the classical by choosing the data pieces with the greatest evidence divergence. This approach offers special benefits when handling massive data annotations. Unfortunately, the majority of existing active learning techniques are rarely used for data deduplication jobs and only use classical entity matching. The research effort suggests a model, a distributed cloud storage method that ensures effectiveness, security, and high availability, to close this research gap. By intelligently distributing data subsets among servers, the suggested approach achieves fault tolerance while preserving redundancy for high availability. This suggested model reduces the amount of storage space and upkeep required for the data while eliminating duplicate copies. Additionally, the data is kept in a highly secure and effective manner. The suggested model's efficacy in fault tolerance, cost reduction, batch auditing, and block- and file-level deduplication is demonstrated by the experimental findings. It performs better than current systems thanks to its robust fault tolerance, minimal time complexity, and excellent deduplication capabilities.