| 2008 | On the correctness of transactional memory. | Rachid Guerraoui, Michal Kapalka |
| 2008 | Design and implementation of a high-performance MPI for C# and the common language infrastructure. | Douglas P. Gregor, Andrew Lumsdaine |
| 2008 | FastForward for efficient pipeline parallelism: a cache-optimized concurrent lock-free queue. | John Giacomoni, Tipp Moseley, Manish Vachharajani |
| 2008 | Massive parallel LDPC decoding on GPU. | Gabriel Falco Paiva Fernandes, Leonel Sousa, Vtor Manuel Mendes da Silva |
| 2008 | Dynamic performance tuning of word-based software transactional memory. | Pascal Felber, Christof Fetzer, Torvald Riegel |
| 2008 | Matrix product on heterogeneous master-worker platforms. | Jack J. Dongarra, Jean-Francois Pineau, Yves Robert, Frdric Vivien |
| 2008 | All-window profiling of concurrent executions. | Chen Ding, Trishul M. Chilimbi |
| 2008 | High performance dense linear algebra on a spatially distributed processor. | Jeffrey R. Diamond, Behnam Robatmili, Stephen W. Keckler, Robert A. van de Geijn, Kazushige Goto, Doug Burger |
| 2008 | Scalable packet classification using interpreting: a cross-platform multi-core solution. | Haipeng Cheng, Zheng Chen, Bei Hua, Xinan Tang |
| 2008 | SuperMatrix: a multithreaded runtime scheduling system for algorithms-by-blocks. | Ernie Chan, Field G. Van Zee, Paolo Bientinesi, Enrique S. Quintana-Ort, Gregorio Quintana-Ort, Robert A. van de Geijn |
| 2008 | Type inference for locality analysis of distributed data structures. | Satish Chandra, Vijay A. Saraswat, Vivek Sarkar, Rastislav Bodk |
| 2008 | A case study in SIMD text processing with parallel bit streams: UTF-8 to UTF-16 transcoding. | Robert D. Cameron |
| 2008 | Compiler-enhanced incremental checkpointing for OpenMP applications. | Greg Bronevetsky, Daniel Marques, Keshav Pingali, Radu Rugina, Sally A. McKee |
| 2008 | Practical experiences with Java software transactional memory. | Evgueni Brevnov, Yuri Dolgov, Boris Kuznetsov, Dmitry Yershov, Vyacheslav Shakin, Dong-yuan Chen, Vijay Menon, Suresh Srinivas |
| 2008 | Software transactional memory for large scale clusters. | Robert L. Bocchino Jr., Vikram S. Adve, Bradford L. Chamberlain |
| 2008 | Automatic data movement and computation mapping for multi-level parallel architectures with explicitly managed memories. | Muthu Manikandan Baskaran, Uday Bondhugula, Sriram Krishnamoorthy, J. Ramanujam, Atanas Rountev, P. Sadayappan |
| 2008 | Semantics-based distributed I/O for mpiBLAST. | Pavan Balaji, Wu-chun Feng, Jeremy S. Archuleta, Heshan Lin, Rajkumar Kettimuthu, Rajeev Thakur, Xiaosong Ma |
| 2008 | Experiences using adaptive concurrency in transactional memory with Lee's routing algorithm. | Mohammad Ansari, Christos Kotselidis, Kim Jarvis, Mikel Lujn, Chris C. Kirkham, Ian Watson |
| 2008 | Compilers and parallel computing systems. | Frances E. Allen |
| 2008 | Safer open-nested transactions through ownership. | Kunal Agrawal, I-Ting Angelina Lee, Jim Sukha |
| 2008 | Nested parallelism in transactional memory. | Kunal Agrawal, Jeremy T. Fineman, Jim Sukha |
| 2007 | Supporting fault-tolerance in streaming grid applications. | Qian Zhu, Liang Chen, Gagan Agrawal |
| 2007 | Optimized lock assignment and allocation: a method for exploiting concurrency among critical sections. | Yuan Zhang, Vugranam C. Sreedhar, Weirong Zhu, Vivek Sarkar, Guang R. Gao |
| 2007 | Barrier matching for programs with textually unaligned barriers. | Yuan Zhang, Evelyn Duesterwald |
| 2007 | Self-adaptive applications on the grid. | Gosia Wrzesinska, Jason Maassen, Henri E. Bal |