Improving cluster computing performance based on job futurity prediction
Hossein Salami, Hamid Saadatfar, Farhad Rahmani Fard, S.‐Kazem Shekofteh, Hossein Deldari · 2010
By recognizing the necessity for preventative and proactive management for today's large scale and fault prone distributed systems, a tendency for these mechanisms has been appeared in recent researchers' efforts. From the birth point of this opinion, event prediction has been known as an effective approach to manage errors preventively. In this work, we attempt to make system performance better by using a job futurity predictor as an advisor for system scheduler. We have a clear sighted vision of failure sources and therefore use a comprehensive data set of system, user and job domains. Using predictor results leads to a significant improvement in performance metrics.