| 2007 | Using fine grain multithreading for energy efficient computing. | Alex Gontmakher, Avi Mendelson, Assaf Schuster |
| 2007 | Parallel programming environment: a key to translating tera-scale platforms into a big success. | Jesse Z. Fang |
| 2007 | Pervasive parallel computing: an historic opportunity for innovation in programming and architecture. | Andrew A. Chien |
| 2007 | Transactional collection classes. | Brian D. Carlstrom, Austen McDonald, Michael Carbin, Christos Kozyrakis, Kunle Olukotun |
| 2007 | Toward terabyte pattern mining: an architecture-conscious solution. | Gregory Buehrer, Srinivasan Parthasarathy, Shirish Tatikonda, Tahsin M. Kur, Joel H. Saltz |
| 2007 | Automatic mapping of nested loops to FPGAS. | Uday Bondhugula, J. Ramanujam, P. Sadayappan |
| 2007 | Reordering constraints for pthread-style locks. | Hans-Juergen Boehm |
| 2007 | Dynamic multigrain parallelization on the cell broadband engine. | Filip Blagojevic, Dimitrios S. Nikolopoulos, Alexandros Stamatakis, Christos D. Antonopoulos |
| 2007 | Promised messages: recovering from inconsistent global states. | Franoise Baude, Denis Caromel, Christian Delb, Ludovic Henrio |
| 2007 | Performance evaluation of the cray XT3 configured with dual core opteron processors. | Richard F. Barrett, Sadaf R. Alam, Jeffrey S. Vetter |
| 2007 | Adaptive work stealing with parallelism feedback. | Kunal Agrawal, Yuxiong He, Charles E. Leiserson |
| 2007 | May-happen-in-parallel analysis of X10 programs. | Shivali Agarwal, Rajkishore Barik, Vivek Sarkar, R. K. Shyamasundar |
| 2007 | Transactional programming in a multi-core environment. | Ali-Reza Adl-Tabatabai, Christos Kozyrakis, Bratin Saha |
| 2007 | Potential show-stoppers for transactional synchronization. | Ali-Reza Adl-Tabatabai, David Dice, Maurice Herlihy, Nir Shavit, Christos Kozyrakis, Christoph von Praun, Michael L. Scott |
| 2006 | Accurate and efficient runtime detection of atomicity errors in concurrent programs. | Liqiang Wang, Scott D. Stoller |
| 2006 | Proving correctness of highly-concurrent linearisable objects. | Viktor Vafeiadis, Maurice Herlihy, Tony Hoare, Marc Shapiro |
| 2006 | RDMA read based rendezvous protocol for MPI over InfiniBand: design alternatives and benefits. | Sayantan Sur, Hyun-Wook Jin, Lei Chai, Dhabaleswar K. Panda |
| 2006 | Parallel programming and code selection in fortress. | Guy L. Steele Jr. |
| 2006 | Parallel programming in modern web search engines. | Raymie Stata |
| 2006 | Minimizing execution time in MPI programs on an energy-constrained, power-scalable cluster. | Robert Springer, David K. Lowenthal, Barry Rountree, Vincent W. Freeh |
| 2006 | A case study in top-down performance estimation for a large-scale parallel application. | Ilya Sharapov, Robert Kroeger, Guy Delamarter, Razvan Cheveresan, Matthew Ramsay |
| 2006 | Scalable synchronous queues. | William N. Scherer III, Doug Lea, Michael L. Scott |
| 2006 | McRT-STM: a high performance software transactional memory system for a multi-core runtime. | Bratin Saha, Ali-Reza Adl-Tabatabai, Richard L. Hudson, Chi Cao Minh, Ben Hertzberg |
| 2006 | On-line automated performance diagnosis on thousands of processes. | Philip C. Roth, Barton P. Miller |
| 2006 | Hardware profile-guided automatic page placement for ccNUMA systems. | Jaydeep Marathe, Frank Mueller |