The Community for Technology Leaders
RSS Icon
Subscribe
Issue No.12 - December (1997 vol.46)
pp: 1381-1386
ABSTRACT
<p><b>Abstract</b>—In checkpointing schemes with task duplication, checkpointing serves two purposes: detecting faults by comparing the processors' states at checkpoints, and reducing fault recovery time by supplying a safe point to rollback to. In this paper, we show that, by tuning the checkpointing schemes to a given architecture, a significant reduction in the execution time can be achieved. The main idea is to use two types of checkpoints: compare-checkpoints (comparing the states of the redundant processes to detect faults) and store-checkpoints (storing the states to reduce recovery time). With two types of checkpoints, we can use both the comparison and storage operations in an efficient way and improve the performance of checkpointing schemes. Results we obtained show that, in some cases, using compare and store checkpoints can reduce the overhead of DMR checkpointing schemes by as much as 30 percent.</p>
INDEX TERMS
Fault-tolerant computing, checkpointing, task duplication, parallel computing, performance optimization.
CITATION
Avi Ziv, Jehoshua Bruck, "Performance Optimization of Checkpointing Schemes with Task Duplication", IEEE Transactions on Computers, vol.46, no. 12, pp. 1381-1386, December 1997, doi:10.1109/12.641939
23 ms
(Ver 2.0)

Marketing Automation Platform Marketing Automation Tool