Detecting of Alike Data for Information Recognition and Storing With Low Charges

K M Siva Krishna K Somasekhar · Zenodo (CERN European Organization for Nuclear Research) · 2018

Cloud computing greatly facilitates information suppliers who need to source their information to the cloud while not revealing their sensitive information to external parties and would love users with sure credentials to be ready to access the info. Data reduction has become more and more vital in storage systems because of the explosive growth of digital information within the world that has ushered within the huge information era. one amongst the most challenges facing large-scale information reduction is a way to maximally notice and eliminate redundancy at terribly low overheads. during this paper, we tend to gift DARE, a low-overhead Deduplication-Aware alikeness detection and Elimination theme that effectively exploits existing duplicate-adjacency info for extremely economical alikeness detection in information deduplication based mostly backup/archiving storage systems. the most plan behind DARE is to use a theme, decision Duplicate-Adjacency based mostly alikeness Detection (DupAdj), by considering any 2 information chunks to be similar (i.e., candidates for delta compression) if their several adjacent information chunks are duplicate in an exceedingly deduplication system, and so additional enhance the alikeness detection potency by Associate in Nursing improved super-feature approach. Our experimental results supported real-world and artificial backup datasets show that DARE solely consumes regarding 1/4 and 1/2 severally of the computation and categorization overheads needed by the normal super-feature approaches whereas police investigation 2-10% additional redundancy and achieving the next outturn, by exploiting existing duplicate-adjacency info for alikeness detection and finding the "sweet spot" for the super-feature approach.

Read the paper · More papers on PaperTik