loading...
 This Article 
   
 Share 
   
 Bibliographic References 
   
 Add to: 
 
Digg
Furl
Spurl
Blink
Simpy
Google
Del.icio.us
Y!MyWeb
 
 Search 
   
17th International Symposium on Computer Architecture and High Performance Computing (SBAC-PAD'05)
VRM: A Failure-Aware Grid Resource Management System
Rio de Janeiro, Brazil
October 24-October 27
ISBN: 0-7695-2446-X
Lars-Olof Burchard, Technische Universitaet Berlin, GERMANY
Cesar A. F. De Rose, PUCRS, Porto Alegre, BRASIL
Hans-Ulrich Heiss, Technische Universitaet Berlin, GERMANY
Barry Linnert, Technische Universitaet Berlin, GERMANY
Jorg Schneider, Technische Universitaet Berlin, GERMANY
For resource management in Grid environments, advance reservations turned out to be very useful and hence are supported by a variety of Grid toolkits. However, failure recovery for such systems has not yet received the attention it deserves. In this paper, we address the problem of remapping reservations to other resources, when the originally selected resource fails. Instead of dealing with jobs already running, which usually means checkpointing and migration, our focus is on jobs that are scheduled on the failed resource for a specific future period of time but not started yet. The most critical factor when solving this problem is the estimation of the downtime. We avoid the drawbacks of under- or overestimating the downtime by a dynamic load-based approach that is evaluated by extensive simulations in a Grid environment and shows superior performance compared to estimationbased approaches.
Citation:
Lars-Olof Burchard, Cesar A. F. De Rose, Hans-Ulrich Heiss, Barry Linnert, Jorg Schneider, "VRM: A Failure-Aware Grid Resource Management System," sbac-pad, pp.218-227, 17th International Symposium on Computer Architecture and High Performance Computing (SBAC-PAD'05), 2005
Usage of this product signifies your acceptance of the Terms of Use.