loading...
 This Article 
   
 Share 
   
 Bibliographic References 
   
 Add to: 
 
Digg
Furl
Spurl
Blink
Simpy
Google
Del.icio.us
Y!MyWeb
 
 Search 
   
13th IEEE International Symposium on Modeling, Analysis, and Simulation of Computer and Telecommunication Systems
Disk Infant Mortality in Large Storage Systems
Atlanta, Georgia
September 27-September 29
ISBN: 0-7695-2458-3
Qin Xin, Storage Systems Research Center, University of California, Santa Cruz, CA 95064
J. E. Thomas, Storage Systems Research Center, University of California, Santa Cruz, CA 95064
S. J. Schwarz, Computer Engineering Department, Santa Clara University, Santa Clara, CA 95053
Ethan L. Miller, Storage Systems Research Center, University of California, Santa Cruz, CA 95064

As disk drives have dropped in price relative to tape, the desire for the convenience and speed of online access to large data repositories has led to the deployment of petabyte-scale disk farms with thousands of disks. Unfortunately, the very large size of these repositories renders them vulnerable to previously rare failure modes such as multiple, unrelated disk failures leading to data loss. While some business models, such as free email servers, may be able to tolerate some occurrence of data loss, others, including premium online services and storage of simulation results at a national laboratory, cannot.

This paper describes the effect of infant mortality on long-term failure rates of systems that must preserve their data for decades. Our failure models incorporate the well-known bathtub curve, which reflects the higher failure rates of new disk drives, a lower, constant failure rate during the remainder of the design life span, and increased failure rates as components wear out. Large systems are vulnerable to the "cohort effect" that occurs when many disks are simultaneously replaced by new disks. Our more accurate disk models and simulations have yielded predictions of system lifetimes that are more pessimistic than existing models that assume a constant disk failure rate. Thus, larger system scale requires designers to take disk infant mortality into account.

Citation:
Qin Xin, J. E. Thomas, S. J. Schwarz, Ethan L. Miller, "Disk Infant Mortality in Large Storage Systems," mascots, pp.125-134, 13th IEEE International Symposium on Modeling, Analysis, and Simulation of Computer and Telecommunication Systems, 2005
Usage of this product signifies your acceptance of the Terms of Use.