SDIRTCC: Secure duplicate Identification and removal from textual data in cloud computing
T.M.S. Mekalarani, M. Giri, Dimas Subhan, B. Sai Kumar, K. Sunilkumar Reddy, B. Shilpa · 2025
In the world day by day textual information size is increased exponential and in particular, vision, analysis and interpretation of language force a challenge of management of data, in a big storage with heavy storage space. In order to minimize data size eliminate duplicate is one of the way to follow and duplicate types are also forces security problems. In this research a new method secure duplicate packet and elimination from textual dataset in a cloud computing (SDIRTCC). It is working in hybrid model by combining eliminate duplication from user side and as well as from server side to get good rate of compression and provided full security to data. SDIRTCC is small techniques preprocess data at user side to identify and eliminate duplicate entries and this is well suitable method for devices like IoT working under constraints, and also designed to protect from side channel attack. Touchdown(TD) dataset is used to conduct experiment. It consists of running the written by peoples. SDIRTCC shows 72% of compression rate that in turn shows the impact of duplicate data present in the dataset, and reduced TD dataset occupies less storage when compared with original TD dataset. Finally, reducing dataset size demands less storage space in cloud, saved lot of space in cloud, easy to maintain security, and enhanced efficiency and management of big data in cloud computing.