Parallel and Distributed Processing Symposium, International (2006)
Rhodes Island, Greece
Apr. 25, 2006 to Apr. 29, 2006
K. Malkowski , Dept. of Comput. Sci.&Eng., Pennsylvania State Univ., University Park, PA, USA
Ingyu Lee , Dept. of Comput. Sci.&Eng., Pennsylvania State Univ., University Park, PA, USA
P. Raghavan , Dept. of Comput. Sci.&Eng., Pennsylvania State Univ., University Park, PA, USA
M.J. Irwin , Dept. of Comput. Sci.&Eng., Pennsylvania State Univ., University Park, PA, USA
We characterize the performance and power attributes of the conjugate gradient (CG) sparse solver which is widely used in scientific applications. We use cycle-accurate simulations with SimpleScalar and Wattch, on a processor and memory architecture similar to the configuration of a node of the BlueGene/L. We first demonstrate that substantial power savings can be obtained without performance degradation if low power modes of caches can be utilized. We next show that if Dynamic Voltage Scaling (DVS) can be used, power and energy savings are possible, but these are realized only at the expense of performance penalties. We then consider two simple memory subsystem optimizations, namely memory and level-2 cache prefetching. We demonstrate that when DVS and low power modes of caches are used with these optimizations, performance can be improved significantly with reductions in power and energy. For example, execution time is reduced by 23%, power by 55% and energy by 65% in the final configuration at 500 MHz relative to the original at 1 GHz. We also use our codes and the CG NAS benchmark code to demonstrate that performance and power profiles can vary significantly depending on matrix properties and the level of code tuning. These results indicate that architectural evaluations can benefit if traditional benchmarks are augmented with codes more representative of tuned scientific applications.
level-2 cache prefetching, conjugate gradient sparse solver, performance-power characteristics, cycle-accurate simulation, SimpleScalar, Wattch, dynamic voltage scaling, memory subsystem optimization
P. Raghavan, M. Irwin, K. Malkowski and Ingyu Lee, "Conjugate gradient sparse solvers: performance-power characteristics," Parallel and Distributed Processing Symposium, International(IPDPS), Rhodes Island, Greece, 2006, pp. 338.