| 2012 | Algorithm-based fault tolerance for dense matrix factorizations. | Peng Du, Aurlien Bouteiller, George Bosilca, Thomas Hrault, Jack J. Dongarra |
| 2012 | Scalable parallel debugging with statistical assertions. | Minh Ngoc Dinh, David Abramson, Chao Jin, Andrew Gontarek, Bob Moench, Luiz De Rose |
| 2012 | Lock cohorting: a general technique for designing NUMA locks. | David Dice, Virendra J. Marathe, Nir Shavit |
| 2012 | A speculation-friendly binary search tree. | Tyler Crain, Vincent Gramoli, Michel Raynal |
| 2012 | Exploring parallelism in volume ray casting: understanding the programming issues of multithreaded accelerators. | Guilherme Cox, Cleomar Pereira da Silva, Leandro F. Cupertino, Cristiana Bentes, Ricardo C. Farias |
| 2012 | A case for secure and scalable hypervisor using safe language. | Haibo Chen, Binyu Zang |
| 2012 | PARRAY: a unifying array representation for heterogeneous parallelism. | Yifeng Chen, Xiang Cui, Hong Mei |
| 2012 | Performance analysis of parallel constraint-based local search. | Yves Caniou, Daniel Diaz, Florian Richoux, Philippe Codognet, Salvador Abreu |
| 2012 | NDetermin: inferring nondeterministic sequential specifications for parallelism correctness. | Jacob Burnim, Tayfun Elmas, George C. Necula, Koushik Sen |
| 2012 | Efficient deadlock avoidance for streaming computation with filtering. | Jeremy D. Buhler, Kunal Agrawal, Peng Li, Roger D. Chamberlain |
| 2012 | S: a scripting language for high-performance RESTful web services. | Daniele Bonetta, Achille Peternier, Cesare Pautasso, Walter Binder |
| 2012 | Internally deterministic parallel algorithms can be fast. | Guy E. Blelloch, Jeremy T. Fineman, Phillip B. Gibbons, Julian Shun |
| 2012 | Automatic communication optimizations through memory reuse strategies. | Muthu Manikandan Baskaran, Nicolas Vasilache, Benot Meister, Richard Lethin |
| 2012 | Communication avoiding successive band reduction. | Grey Ballard, James Demmel, Nicholas Knight |
| 2012 | Efficient performance evaluation of memory hierarchy for highly multithreaded graphics processors. | Sara S. Baghsorkhi, Isaac Gelado, Matthieu Delahaye, Wen-mei W. Hwu |
| 2012 | AGC: adaptive global clock in software transactional memory. | Ehsan Atoofian, Amir Ghanbari Bavarsad |
| 2012 | Programming parallel embedded and consumer applications in OpenMP superscalar. | Michael Andersch, Chi Ching Chi, Ben H. H. Juurlink |
| 2012 | Optimizing remote accesses for offloaded kernels: application to high-level synthesis for FPGA. | Christophe Alias, Alain Darte, Alexandru Plesco |
| 2012 | An hybrid model for very high level threads. | Jafar Al-Gharaibeh, Clinton L. Jeffery, Kostas N. Oikonomou |
| 2011 | GRace: a low-overhead mechanism for detecting data races in GPU programs. | Mai Zheng, Vignesh T. Ravi, Feng Qin, Gagan Agrawal |
| 2011 | Cooperative reasoning for preemptive execution. | Jaeheon Yi, Caitlin Sadowski, Cormac Flanagan |
| 2011 | All-window profiling and composable models of cache sharing. | Xiaoya Xiang, Bin Bao, Tongxin Bai, Chen Ding, Trishul M. Chilimbi |
| 2011 | ScalaExtrap: trace-based communication extrapolation for spmd programs. | Xing Wu, Frank Mueller |
| 2011 | Active pebbles: a programming model for highly parallel fine-grained data-driven computations. | Jeremiah Willcock, Torsten Hoefler, Nicholas Gerard Edmonds, Andrew Lumsdaine |
| 2011 | COREMU: a scalable and portable parallel full-system emulator. | Zhaoguo Wang, Ran Liu, Yufei Chen, Xi Wu, Haibo Chen, Weihua Zhang, Binyu Zang |