| 2009 | Implementation of a wide-angle lens distortion correction algorithm on the cell broadband engine. | Konstantis Daloukas, Christos D. Antonopoulos, Nikolaos Bellas |
| 2009 | Fast memory snapshot for concurrent programmingwithout synchronization. | JaeWoong Chung, Woongki Baek, Christos Kozyrakis |
| 2009 | Design of a novel SIMD architecture by fusing operations and registers. | Jih-Ching Chiu, Kai-Ming Yang, Yu-Liang Chou |
| 2009 | Computer generation of fast fourier transforms for the cell broadband engine. | Srinivas Chellappa, Franz Franchetti, Markus Pschel |
| 2009 | A parallel levenberg-marquardt algorithm. | Jun Cao, Krista A. Novstrup, Ayush Goyal, Samuel P. Midkiff, James M. Caruthers |
| 2009 | EpiFast: a fast algorithm for large scale realistic epidemic simulations on distributed memory systems. | Keith R. Bisset, Jiangzhuo Chen, Xizhou Feng, V. S. Anil Kumar, Madhav V. Marathe |
| 2009 | Dynamic topology aware load balancing algorithms for molecular dynamics applications. | Abhinav Bhatele, Laxmikant V. Kal, Sameer Kumar |
| 2009 | PARSEC: hardware profiling of emerging workloads for CMP design. | Major Bhadauria, Vincent M. Weaver, Sally A. McKee |
| 2009 | Pattern-based sparse matrix representation for memory-efficient SMVM kernels. | Mehmet Belgin, Godmar Back, Calvin J. Ribbens |
| 2009 | Designing multi-socket systems using silicon photonics. | Scott Beamer, Krste Asanovic, Christopher Batten, Ajay Joshi, Vladimir Stojanovic |
| 2009 | Dynamic task set partitioning based on balancing memory requirements to reduce power consumption. | Diana Bautista, Julio Sahuquillo, Houcine Hassan, Salvador Petit, Jos Duato |
| 2009 | A comprehensive power-performance model for NoCs with multi-flit channel buffers. | Mohammad Arjomand, Hamid Sarbazi-Azad |
| 2009 | Efficient high performance collective communication for the cell blade. | Qasim Ali, Samuel P. Midkiff, Vijay S. Pai |
| 2008 | Shifted declustering: a placement-ideal layout scheme for multi-way replication storage architecture. | Huijun Zhu, Peng Gu, Jun Wang |
| 2008 | CprFS: a user-level file system to support consistent file states for checkpoint and restart. | Ruini Xue, Wenguang Chen, Weimin Zheng |
| 2008 | A projection-based optimization framework for abstractions with application to the unstructured mesh domain. | Brian S. White, Sally A. McKee, Daniel J. Quinlan |
| 2008 | Accurate memory signatures and synthetic address traces for HPC applications. | Jonathan Weinberg, Allan Snavely |
| 2008 | A freespace crossbar for multi-core processors. | Michel N. Victor, Aris K. Silzars, Edward S. Davidson |
| 2008 | Power-aware dynamic placement of HPC applications. | Akshat Verma, Puneet Ahuja, Anindya Neogi |
| 2008 | Efficient computation of sum-products on GPUs through software-managed cache. | Mark Silberstein, Assaf Schuster, Dan Geiger, Anjul Patney, John D. Owens |
| 2008 | Automatic SIMD vectorization of chains of recurrences. | Yixin Shou, Robert A. van Engelen |
| 2008 | Evaluating the effect of replacing CNK with linux on the compute-nodes of blue gene/l. | Edi Shmueli, George Almsi, Jos R. Brunheroto, Jos G. Castaos, Gbor Dzsa, Sameer Kumar, Derek Lieber |
| 2008 | Phasers: a unified deadlock-free construct for collective and point-to-point synchronization. | Jun Shirako, David M. Peixotto, Vivek Sarkar, William N. Scherer III |
| 2008 | Preserving time in large-scale communication traces. | Prasun Ratn, Frank Mueller, Bronis R. de Supinski, Martin Schulz |
| 2008 | Timely offloading of result-data in HPC centers. | Henry M. Monti, Ali Raza Butt, Sudharshan S. Vazhkudai |