Reliability Quantitation of Private Cloud Data Center Considering Device Online Availability

Haotong Qiu, Chen Ma, Jingqi Wu, Peng Liu · 2025

The rapid evolution of cloud computing has under-scored the critical need for reliable cloud data centers to ensure uninterrupted service and high performance. This study presents a unified reliability analysis framework that integrates device online rate as a core component of system evaluation, with Mean Time Between Critical Failures (MTBCF) serving as the primary reliability metric. By systematically modeling the impacts of hard-ware, software, and network failures, the framework captures the complex interactions inherent to cloud environments and accounts for fault-tolerant mechanisms and redundant configurations. The research also tackles resource allocation in cloud data centers as an optimization problem. It uses metaheuristic algorithms such as Genetic Algorithm (GA), Grasshopper Optimization Algorithm (GOA), Whale Optimization Algorithm (WOA), and a new hybrid method called the Whale Optimization Simulated Annealing Algorithm (WOSAA) to manage the complexity of exploring different network topologies. Experimental results show that WOSAA nearly reaches the optimal reliability and it performs much better than the basic Equal Allocation Strategy (EAS). These findings provide valuable insights into reliability-oriented resource management and lay a robust foundation for future research in dynamic fault management and optimization within cloud computing infrastructures.

Read the paper · More papers on PaperTik