| 2008 | Optimizing irregular shared-memory applications for clusters. | Seung-Jai Min, Rudolf Eigenmann |
| 2008 | Focused prefetching: performance oriented prefetching based on commit stalls. | R. Manikantan, R. Govindarajan |
| 2008 | Exploiting idle register classes for fast spill destination. | Fang Lu, Lei Wang, Xiaobing Feng, Zhiyuan Li, Zhaoqing Zhang |
| 2008 | An approach for adaptive DRAM temperature and power management. | Song Liu, Seda Ogrenci Memik, Yu Zhang, Gokhan Memik |
| 2008 | Analyzing memory access intensity in parallel programs on multicore. | Lixia Liu, Zhiyuan Li, Ahmed H. Sameh |
| 2008 | Adaptive runtime tuning of parallel sparse matrix-vector multiplication on distributed memory systems. | Seyong Lee, Rudolf Eigenmann |
| 2008 | The deep computing messaging framework: generalized scalable message passing on the blue gene/P supercomputer. | Sameer Kumar, Gbor Dzsa, Gheorghe Almsi, Philip Heidelberger, Dong Chen, Mark Giampapa, Michael Blocksome, Ahmad Faraj, Jeff Parker, Joe Ratterman, Brian E. Smith, Charles Archer |
| 2008 | Can software reliability outperform hardware reliability on high performance interconnects?: a case study with MPI over infiniband. | Matthew J. Koop, Rahul Kumar, Dhabaleswar K. Panda |
| 2008 | Rotating register allocation with multiple rotating branches. | Suhyun Kim, Soo-Mook Moon |
| 2008 | Petaflop/s, seriously. | David E. Keyes |
| 2008 | Implementing Wilson-Dirac operator on the cell broadband engine. | Khaled Z. Ibrahim, Franois Bodin |
| 2008 | Performance portable optimizations for loops containing communication operations. | Costin Iancu, Wei Chen, Katherine A. Yelick |
| 2008 | Biomedical image analysis on a cooperative cluster of GPUs and multicores. | Timothy D. R. Hartley, mit V. atalyrek, Antonio Ruiz, Francisco D. Igual, Rafael Mayo, Manuel Ujaldon |
| 2008 | Many-core GPU computing with NVIDIA CUDA. | Mark J. Harris |
| 2008 | CUBA: an architecture for efficient CPU/co-processor data communication. | Isaac Gelado, John H. Kelm, Shane Ryoo, Steven S. Lumetta, Nacho Navarro, Wen-mei W. Hwu |
| 2008 | Fast scan algorithms on graphics processors. | Yuri Dotsenko, Naga K. Govindaraju, Peter-Pike J. Sloan, Charles Boyd, John Manferdelli |
| 2008 | Autonomous learning for efficient resource utilization of dynamic VM migration. | Hyung Won Choi, Hukeun Kwak, Andrew Sohn, Kyusik Chung |
| 2008 | Three-dimensional delaunay refinement for multi-core processors. | Andrey N. Chernikov, Nikos Chrisochoides |
| 2008 | Orchestrating data transfer for the cell/B.E. processor. | Tong Chen, Haibo Lin, Tao Zhang |
| 2008 | Automatic analysis of speedup of MPI applications. | Marc Casas, Rosa M. Badia, Jess Labarta |
| 2008 | Data mining on the cell broadband engine. | Gregory Buehrer, Srinivasan Parthasarathy, Matthew Goyder |
| 2008 | The shared-thread multiprocessor. | Jeffery A. Brown, Dean M. Tullsen |
| 2008 | Soft error vulnerability of iterative linear algebra methods. | Greg Bronevetsky, Bronis R. de Supinski |
| 2008 | Analysis of dynamic power management on multi-core processors. | William Lloyd Bircher, Lizy K. John |
| 2008 | A compiler framework for optimization of affine loop nests for gpgpus. | Muthu Manikandan Baskaran, Uday Bondhugula, Sriram Krishnamoorthy, J. Ramanujam, Atanas Rountev, P. Sadayappan |