| 2007 | Automatic nonblocking communication for partitioned global address space programs. | Wei-Yu Chen, Dan Bonachea, Costin Iancu, Katherine A. Yelick |
| 2007 | Cooperative cache partitioning for chip multiprocessors. | Jichuan Chang, Gurindar S. Sohi |
| 2007 | Performance driven data cache prefetching in a dynamic software optimization system. | Jean Christophe Beyler, Philippe Clauss |
| 2007 | An operation stacking framework for large ensemble computations. | Mehmet Belgin, Calvin J. Ribbens, Godmar Back |
| 2007 | GridRod: a dynamic runtime scheduler for grid workflows. | Shahaan Ayyub, David Abramson |
| 2007 | Scheduling FFT computation on SMP and multicore systems. | Ayaz Ali, S. Lennart Johnsson, Jaspal Subhlok |
| 2007 | Tradeoff between data-, instruction-, and thread-level parallelism in stream processors. | Jung Ho Ahn, Mattan Erez, William J. Dally |
| 2007 | Compression in cache design. | Ali-Reza Adl-Tabatabai, Anwar M. Ghuloum, Shobhit O. Kanaujia |
| 2007 | Optimization of data prefetch helper threads with path-expression based statistical modeling. | Tor M. Aamodt, Paul Chow |
| 2006 | TMA: a trap-based memory architecture. | Hkan Zeffer, Zoran Radovic, Martin Karlsson, Erik Hagersten |
| 2006 | The exigency of benchmark and compiler drift: designing tomorrow's processors with yesterday's tools. | Joshua J. Yi, Hans Vandierendonck, Lieven Eeckhout, David J. Lilja |
| 2006 | Accelerating sparse matrix computations via data compression. | Jeremiah Willcock, Andrew Lumsdaine |
| 2006 | User-guided symbiotic space-sharing of real workloads. | Jonathan Weinberg, Allan Snavely |
| 2006 | Multigrid and Gauss-Seidel smoothers revisited: parallelization on chip multiprocessors. | Dan Wallin, Henrik Lf, Erik Hagersten, Sverker Holmgren |
| 2006 | A scalable low power issue queue for large instruction window processors. | Rajesh Vivekanandham, Bharadwaj S. Amrutur, R. Govindarajan |
| 2006 | Violated dependence analysis. | Nicolas Vasilache, Cdric Bastoul, Albert Cohen, Sylvain Girbal |
| 2006 | Scalable, fault tolerant membership for MPI tasks on HPC systems. | Jyothish Varma, Chao Wang, Frank Mueller, Christian Engelmann, Stephen L. Scott |
| 2006 | Sensitivity analysis of knapsack-based task scheduling on the grid. | Daniel C. Vanderster, Nikitas J. Dimopoulos |
| 2006 | A modern high-performance processor pipeline. | Marc Tremblay |
| 2006 | A scalable communication layer for multi-dimensional hyper crossbar network using multiple gigabit ethernet. | Shinji Sumimoto, Kazuichi Ooe, Kouichi Kumon, Taisuke Boku, Mitsuhisa Sato, Akira Ukawa |
| 2006 | Scientific applications vs. SPEC-FP: a comparison of program behavior. | Kyle Rupnow, Arun Rodrigues, Keith D. Underwood, Katherine Compton |
| 2006 | Probabilistic accuracy bounds for fault-tolerant computations that discard tasks. | Martin C. Rinard |
| 2006 | Selective predicate prediction for out-of-order processors. | Eduardo Quiones, Joan-Manuel Parcerisa, Antonio Gonzlez |
| 2006 | Profitable loop fusion and tiling using model-driven empirical search. | Apan Qasem, Ken Kennedy |
| 2006 | Quantum mechanical approaches to information processing. | Steven Prawer |