Utilizing reinforcement learning to optimize data archiving strategy for application server

Andrii Harasivka, Anatolii Lupenko · Scientific journal of the Ternopil national technical university · 2025

Efficient data compression is a critical component of modern data backup systems, particularly in environments with diverse file types and different performance features. Backups safeguard critical application server data against unexpected failures, such as hardware malfunctions, software bugs, or cyberattacks, ensuring business continuity. Many industries maintain secure data backups to meet legal and regulatory requirements, ensuring loyalty to data protection and privacy laws. However regular backup solutions are often unable to adapt effectively to the changing data of the application server. This paper proposes a novel reinforcement learning (RL) approach using the Proximal Policy Optimization (PPO) model to dynamically optimize data archiving strategy, which is part of backup systems. The model is trained to predict the most efficient combination of compression parameters based on the attributes of files in the target file system. By learning from file-specific observations and rewards, it adapts to work in a backup-specific environment to minimize consumed disk storage and backup time. This adaptive approach enables real-time decision-making tailored to workload variations of the application server environment. The proposed solution performs backup operations across various file types and configurations for the learning phase, where the model evaluates and adjusts policy to maximize its efficiency. To evaluate the results another client should perform all possible combinations of action parameters to determine all possible observations and rewards. The rewards are compared to decisions made by the proposed solution to ensure PPO model has correct and best possible predictions during the evaluation phase. This study highlights the potential of RL in automating and optimizing data backup tasks, providing a scalable solution for high-performance systems or environments with frequent data writes. The results obtained contribute to improving the software backup systems and DevOps specialists' work and reduce disk storage consumption and time elapsed for backup tasks for the application server.

Read the paper · More papers on PaperTik