Failure-Aware Virtual Machine Configuration for Cloud Computing
Yaqin Luo, Qi Li · 2012
Failure occurrence and its impact on system performance have become an increasingly important concern in cloud computing. Most techniques in today's system are reactive schemes to recover after failure which could lead to major cost and significantly affect system performance. Instead, we propose a proactive failure aware virtual machine infrastructure for cloud computing. Our approach takes both the performance and reliability status of a node into account to forecast failure in a given time window. We leverage failure prediction techniques to mitigate the potential failure impact on system reliability and productivity. The mechanism can also reschedule the running job in case the failures occurred during execution. The experiment results show the enhancement of system productivity and reliability significantly by using the proposed strategy.