Many current approaches to software-implemented fault tolerance (SIFT) rely on process replication, which is often prohibitively expensive for practical use due to its high performance overhead and cost. The Adaptive Reconfigurable Mobile Objects of Reliability (Armor) middleware architecture offers a scalable low-overhead way to provide high-dependability services to applications. It uses coordinated multithreaded processes to manage redundant resources acrossinterconnected nodes, detect errors in user applications and infrastructural components, and provide failure recovery. The authors describe their experiences and lessons learned in deploying Armor in several diverse fields.
Index Terms:
software-implemented fault tolerance, high-dependability services, failure recovery, middleware
Citation:
Zbigniew Kalbarczyk, Ravishankar K. Iyer, Long Wang, "Application Fault Tolerance with Armor Middleware," IEEE Internet Computing, vol. 9, no. 2, pp. 28-37, Mar./Apr. 2005, doi:10.1109/MIC.2005.31