DEDUPLICATION OF DATA IN TEXT FILES IN CLOUD STORAGE
D. K. G. Mallika · Zenodo (CERN European Organization for Nuclear Research) · 2017
The world around us consists of so much information which has to be interpreted. When this interpretation takes place we obtain something called data. Data about this interpreted data is called Metadata. For example, File consists of Information so we can say File is data and details like File name, file size etc. are Metadata. Data is stored in storage spaces. Currently the most dominantly used storage method is Cloud Storage. One of the major issues faced by Cloud storage is data redundancy. To solve this issue one should achieve Data Deduplication. Data Deduplication is a concept where unique data or a particular pattern of data is identified and stored. This is then compared to other data available in the system. If a match is found, then the data is replaced by a link or a reference to the stored data. The proposed system aims at building a data deduplication mechanism which provides a practical solution that is more secure than previous techniques in some cases.