| 2008 | Efficient auction-based grid reservations using dynamic programming. | Andrew Mutz, Richard Wolski |
| 2008 | Floating point based Cellular Automata simulations using a dual FPGA-enabled system. | S. Murtaza, Alfons G. Hoekstra, Peter M. A. Sloot |
| 2008 | Programming the Intel 80-core network-on-a-chip terascale processor. | Timothy G. Mattson, Rob F. Van der Wijngaart, Michael A. Frumkin |
| 2008 | A multi-level parallel simulation approach to electron transport in nano-scale transistors. | Mathieu Luisier, Gerhard Klimeck |
| 2008 | A lightweight execution framework for massive independent tasks. | Hui Li, Huashan Yu, Xiaoming Li |
| 2008 | Massively parallel genomic sequence search on the Blue Gene/P architecture. | Heshan Lin, Pavan Balaji, Ruth Poole, Carlos P. Sosa, Xiaosong Ma, Wu-chun Feng |
| 2008 | Dynamically adapting file domain partitioning methods for collective I/O based on underlying parallel file system locking protocols. | Wei-keng Liao, Alok N. Choudhary |
| 2008 | Lessons learned at 208K: towards debugging millions of cores. | Gregory L. Lee, Dong H. Ahn, Dorian C. Arnold, Bronis R. de Supinski, Matthew P. LeGendre, Barton P. Miller, Martin Schulz, Ben Liblit |
| 2008 | Global trees: a framework for linked data structures on distributed memory parallel systems. | D. Brian Larkins, James Dinan, Sriram Krishnamoorthy, Srinivasan Parthasarathy, Atanas Rountev, P. Sadayappan |
| 2008 | Using overlays for efficient data transfer over shared wide-area networks. | Gaurav Khanna, mit V. atalyrek, Tahsin M. Kur, Rajkumar Kettimuthu, P. Sadayappan, Ian T. Foster, Joel H. Saltz |
| 2008 | A novel migration-based NUCA design for chip multiprocessors. | Mahmut T. Kandemir, Feihui Li, Mary Jane Irwin, Seung Woo Son |
| 2008 | Capturing performance knowledge for automated analysis. | Kevin A. Huck, Oscar R. Hernandez, Van Bui, Sunita Chandrasekaran, Barbara M. Chapman, Allen D. Malony, Lois C. McInnes, Boyana Norris |
| 2008 | Hardware task scheduling optimizations for reconfigurable computing. | Miaoqing Huang, Harald Simmler, Proshanta Saha, Tarek A. El-Ghazawi |
| 2008 | The role of MPI in development time: a case study. | Lorin Hochstein, Forrest Shull, Lynn B. Reid |
| 2008 | System support for many task computing. | Eric Van Hensbergen, Ronald G. Minnich |
| 2008 | Exploring data parallelism and locality in wide area networks. | Yunhong Gu, Robert L. Grossman |
| 2008 | Communication avoiding Gaussian elimination. | Laura Grigori, James Demmel, Hua Xiang |
| 2008 | High performance discrete Fourier transforms on graphics processors. | Naga K. Govindaraju, Brandon Lloyd, Yuri Dotsenko, Burton Smith, John Manferdelli |
| 2008 | Scalable load-balance measurement for SPMD codes. | Todd Gamblin, Bronis R. de Supinski, Martin Schulz, Robert J. Fowler, Daniel A. Reed |
| 2008 | Characterizing application sensitivity to OS interference using kernel-level noise injection. | Kurt B. Ferreira, Patrick G. Bridges, Ron Brightwell |
| 2008 | BitDew: a programmable environment for large-scale data management and distribution. | Gilles Fedak, Haiwu He, Franck Cappello |
| 2008 | Virtualizing and sharing reconfigurable resources in High-Performance Reconfigurable Computing systems. | Esam El-Araby, Ivn Gonzlez, Tarek A. El-Ghazawi |
| 2008 | High-radix crossbar switches enabled by proximity communication. | Hans Eberle, Pedro Javier Garca, Jos Flich, Jos Duato, Robert J. Drost, Nils Gura, David Hopkins, Wladek Olesinski |
| 2008 | An adaptive cut-off for task parallelism. | Alejandro Duran, Julita Corbaln, Eduard Ayguad |
| 2008 | The cost of doing science on the cloud: the Montage example. | Ewa Deelman, Gurmeet Singh, Miron Livny, G. Bruce Berriman, John Good |