Prevention Of Arbitrary Data Tuples In Open Nets Using Apache Spark

K. Yacob Raju, T. Srinivasarao · IJITR · 2016

Because the progress of internet application, the regularity of storing the great deal of information is growing day-to-day. Generally, file systems are made to handle file procedures like store, edit, retrieve etc. But, the primary crisis would be to enhance the storage efficiency, not understanding the information and semantics of files you're storing exactly the same file multiples occasions i.e., duplication. To solve such problems, a note digest formula (MD-5) produced hash code can be used to check on if the file already is available or otherwise to lessen storage and improve resource utilization. The goal of the paper would be to generate Hadoop Distributed File System with Deduplication method to maintain bulk of information. For this reason, you are able to increase storage efficiency and improve network bandwidth for cloud atmosphere also.

Read the paper · More papers on PaperTik