Enhancing Cloud Security by Performing Deduplication Using Serial Cascaded Autoencoder With GRU and Optimal Key‐Based Data Sanitization

J. K. Periasamy, Chin‐Shiuh Shieh, Mong‐Fong Horng · Computational Intelligence · 2025

ABSTRACT De‐duplication is critically important for cloud computing since it permits the detection of repeated data within the cloud system under fewer resources and expenses. De‐duplication removes unnecessary data from the cloud centers, which helps to identify the appropriate owner of cloud material. Each piece of data saved in the cloud is owned by a large number of cloud users, even though it contains just a single copy of the data. The dynamic nature of the cloud resources is not handled by the prior deduplication models and the existing models require more computing power for accurately determining the presence of duplicate files in the cloud system. In addition, the prior models split the files into chunks for determining the similar files in the cloud system which affects the quality of the data. To conquer these difficulties, an adaptive deep learning‐based data deduplication model is developed using an optimization algorithm. The main innovation of the proposed research is to rapidly detect duplicate records in the cloud data and also provide high‐level security while maintaining the operational efficiency of the cloud system. The proposed model acts as an efficient attack resistance system and it also ensures the data availability of the cloud system more rapidly. This data deduplication implies in detecting and examining the patterns inside records of information to precisely notice and eliminate repeated identical information. Hence, the data connected to the input pattern is given to the Serial Cascaded Autoencoder with Gated Recurrent Unit (SCA‐GRU) for the deduplication process. After deduplication, the unnecessary data are removed for the precise consumption of resources to store exclusive data. To maintain the security of data, the optimal key‐based data sanitization process is performed, in which the key is optimally generated with the aid of a Mutated Fitness‐Based Krill Herd Optimization Algorithm (MF‐KHO). This encoded data is then safely kept in the cloud, which protects the data from illegal access and possible defense breaches. The outcome of the suggested approach is validated with the previous data deduplication system to show the efficiency of the developed model. The experimental results showed that the recommended deduplication approach reaches an accuracy of 95.37%. Through efficient data deduplication, the storage requirement of the data is greatly reduced, which facilitates cost reduction and resource optimization within the cloud system and also the storage capacity utilization of the cloud system is greatly improved.

Read the paper · More papers on PaperTik