| 2009 | Performance modeling and automatic ghost zone optimization for iterative stencil loops on GPUs. | Jiayuan Meng, Kevin Skadron |
| 2009 | A translation system for enabling data mining applications on GPUs. | Wenjing Ma, Gagan Agrawal |
| 2009 | An infrastructure for scalable and portable parallel programs for computational chemistry. | Victor Lotrich, Norbert Flocke, Mark Ponton, Beverly A. Sanders, Erik Deumens, Rodney J. Bartlett, Ajith Perera |
| 2009 | Thrifty interconnection network for HPC systems. | Jian Li, Lixin Zhang, Charles Lefurgy, Richard R. Treumann, Wolfgang E. Denzel |
| 2009 | Evaluating high performance communication: a power perspective. | Jiuxing Liu, Dan E. Poff, Blent Abali |
| 2009 | DBDB: optimizing DMATransfer for the cell be architecture. | Tao Liu, Haibo Lin, Tong Chen, Kevin O'Brien, Ling Shao |
| 2009 | R-ADMAD: high reliability provision for large-scale de-duplication archival storage systems. | Chuanyi Liu, Yu Gu, Linchun Sun, Bin Yan, Dongsheng Wang |
| 2009 | Virtualization polling engine (VPE): using dedicated CPU cores to accelerate I/O virtualization. | Jiuxing Liu, Blent Abali |
| 2009 | Prefetch optimizations on large-scale applications via parameter value prediction. | Shih-wei Liao, Tzu-Han Hung, Donald Nguyen, Hucheng Zhou, Chinyen Chou, Chia-Heng Tu |
| 2009 | P-Code: a new RAID-6 code with optimal properties. | Chao Jin, Hong Jiang, Dan Feng, Lei Tian |
| 2009 | Cancellation of loads that return zero using zero-value caches. | Md. Mafijul Islam, Sally A. McKee, Per Stenstrm |
| 2009 | Access map pattern matching for data cache prefetch. | Yasuo Ishii, Mary Inaba, Kei Hiraki |
| 2009 | Approximate kernel matrix computation on GPUs forlarge scale learning applications. | Mohamed E. Hussein, Wael Abd-Almageed |
| 2009 | A graph based approach for MPI deadlock detection. | Tobias Hilbrich, Bronis R. de Supinski, Martin Schulz, Matthias S. Mller |
| 2009 | Rate-based QoS techniques for cache/memory in CMP platforms. | Andrew Herdrich, Ramesh Illikkal, Ravi R. Iyer, Donald Newell, Vineet Chadha, Jaideep Moses |
| 2009 | Parametric multi-level tiling of imperfectly nested loops. | Albert Hartono, Muthu Manikandan Baskaran, Cdric Bastoul, Albert Cohen, Sriram Krishnamoorthy, Boyana Norris, J. Ramanujam, P. Sadayappan |
| 2009 | Dynamic cache clustering for chip multiprocessors. | Mohammad Hammoud, Sangyeun Cho, Rami G. Melhem |
| 2009 | Performance modeling for DFT algorithms in FFTW. | Liang Gu, Xiaoming Li |
| 2009 | The roadrunner project and the importance of energy efficiency on the road to exascale computing. | Don G. Grice |
| 2009 | QuakeTM: parallelizing a complex sequential application using transactional memory. | Vladimir Gajinov, Ferad Zyulkyarov, Osman S. Unsal, Adrin Cristal, Eduard Ayguad, Tim Harris, Mateo Valero |
| 2009 | Computing outside the box. | Ian T. Foster |
| 2009 | How GPUs can outperform ASICs for fast LDPC decoding. | Gabriel Falco Paiva Fernandes, Vtor Manuel Mendes da Silva, Leonel Sousa |
| 2009 | MPI collective communications on the blue gene/p supercomputer: algorithms and optimizations. | Ahmad Faraj, Sameer Kumar, Brian E. Smith, Amith R. Mamidala, John A. Gunnels, Philip Heidelberger |
| 2009 | Zero-content augmented caches. | Julien Dusser, Thomas Piquet, Andr Seznec |
| 2009 | MPI-aware compiler optimizations for improving communication-computation overlap. | Anthony Danalis, Lori L. Pollock, D. Martin Swany, John Cavazos |