loading...
 This Article 
   
 Share 
   
 Bibliographic References 
   
 Add to: 
 
Digg
Furl
Spurl
Blink
Simpy
Google
Del.icio.us
Y!MyWeb
 
 Search 
   
Eighth ACIS International Conference on Software Engineering, Artificial Intelligence, Networking, and Parallel/Distributed Computing (SNPD 2007)
A Survey on Failure Prediction of Large-Scale Server Clusters
Haier International Training Center, Qingdao, China
July 30-August 01
ISBN: 0-7695-2909-7
Zhenghua Xue, Xi'an Jiaotong University, China
Xiaoshe Dong, Xi'an Jiaotong University, China
Siyuan Ma, Xi'an Jiaotong University, China
Weiqing Dong, Xi'an Jiaotong University, China
As the size and complexity of cluster systems grows, failure rates accelerate dramatically. To reduce the disaster caused by failures, it is desirable to identify the potential failures ahead of their occurrence. In this paper, we survey the state of the art in failure prediction of cluster systems. The characteristic of failures in cluster systems are addressed, and some statistic results are shown. We explore the ways of the collection and preprocessing of data for failure prediction, and suggest a procedure for preprocessing the records in automatically generated log files. Focused on the main idea of five prediction methods, including statistic based threshold, time series analysis, rule-based classification, Bayesian network models and semi-Markov process models, are analyzed respectively. In addition, concerning the accuracy and practicality, we present five metrics for evaluating the failure prediction techniques and compare the five techniques with the five metrics.
Citation:
Zhenghua Xue, Xiaoshe Dong, Siyuan Ma, Weiqing Dong, "A Survey on Failure Prediction of Large-Scale Server Clusters," snpd, vol. 2, pp.733-738, Eighth ACIS International Conference on Software Engineering, Artificial Intelligence, Networking, and Parallel/Distributed Computing (SNPD 2007), 2007
Usage of this product signifies your acceptance of the Terms of Use.