| 2011 | SpiceC: scalable parallelism via implicit copying and explicit commit. | Min Feng, Rajiv Gupta, Yi Hu |
| 2011 | Auto-tuning of fast fourier transform on graphics processors. | Yuri Dotsenko, Sara S. Baghsorkhi, Brandon Lloyd, Naga K. Govindaraju |
| 2011 | SCRATCH: a tool for automatic analysis of dma races. | Alastair F. Donaldson, Daniel Kroening, Philipp Rmmer |
| 2011 | ULCC: a user-level facility for optimizing shared cache performance on multicores. | Xiaoning Ding, Kaibo Wang, Xiaodong Zhang |
| 2011 | Two examples of parallel programming without concurrency constructs (PP-CC). | Chen Ding |
| 2011 | Algorithm-based recovery for HPL. | Teresa Davies, Zizhong Chen, Christer Karlsson, Hui Liu |
| 2011 | A domain-specific approach to heterogeneous parallelism. | Hassan Chafi, Arvind K. Sujeeth, Kevin J. Brown, HyoukJoong Lee, Anand R. Atreya, Kunle Olukotun |
| 2011 | Copperhead: compiling an embedded data parallel language. | Bryan Catanzaro, Michael Garland, Kurt Keutzer |
| 2011 | Automatic safety proofs for asynchronous memory operations. | Matko Botincan, Mike Dodds, Alastair F. Donaldson, Matthew J. Parkinson |
| 2011 | Programming the memory hierarchy revisited: supporting irregular parallelism in sequoia. | Michael Bauer, John Clark, Eric Schkufza, Alex Aiken |
| 2010 | Debugging programs that use atomic blocks and transactional memory. | Ferad Zyulkyarov, Tim Harris, Osman S. Unsal, Adrin Cristal, Mateo Valero |
| 2010 | Does cache sharing on modern CMP matter to the performance of contemporary multithreaded programs? | Eddy Z. Zhang, Yunlian Jiang, Xipeng Shen |
| 2010 | Continuous speculative program parallelization in software. | Chao Zhang, Chen Ding, Xiaoming Gu, Kirk Kelsey, Tongxin Bai, Xiaobing Feng |
| 2010 | Fast tridiagonal solvers on the GPU. | Yao Zhang, Jonathan Cohen, John D. Owens |
| 2010 | PHANTOM: predicting performance of parallel applications on large-scale parallel machines using a single node. | Jidong Zhai, Wenguang Chen, Weimin Zheng |
| 2010 | An optimizing compiler for GPGPU programs with input-data sharing. | Yi Yang, Ping Xiang, Jingfei Kong, Huiyang Zhou |
| 2010 | Using data structure knowledge for efficient lock generation and strong atomicity. | Gautam Upadhyaya, Samuel P. Midkiff, Vijay S. Pai |
| 2010 | Lazy binary-splitting: a run-time adaptive work-stealing scheduler. | Alexandros Tzannes, George C. Caragea, Rajeev Barua, Uzi Vishkin |
| 2010 | Extreme scale computing: challenges and opportunities. | Josep Torrellas, Bill Gropp, Jaime H. Moreno, Kunle Olukotun, Vivek Sarkar |
| 2010 | Analyzing lock contention in multithreaded applications. | Nathan R. Tallent, John M. Mellor-Crummey, Allan Porterfield |
| 2010 | Composable thread coloring. | Dean F. Sutherland, William L. Scherlis |
| 2010 | CUDAlign: using GPU to accelerate the comparison of megabase genomic sequences. | Edans Flavius de Oliveira Sandes, Alba Cristina Magalhaes Alves de Melo |
| 2010 | Is transactional programming actually easier? | Christopher J. Rossbach, Owen S. Hofmann, Emmett Witchel |
| 2010 | The LOFAR correlator: implementation and performance analysis. | John W. Romein, P. Chris Broekema, Jan David Mol, Rob van Nieuwpoort |
| 2010 | Thread to strand binding of parallel network applications in massive multi-threaded systems. | Petar Radojkovic, Vladimir Cakarevic, Javier Verd, Alex Pajuelo, Francisco J. Cazorla, Mario Nemirovsky, Mateo Valero |