| 2010 | Modeling transactional memory workload performance. | Donald E. Porter, Emmett Witchel |
| 2010 | KRASH: reproducible CPU load generation on many cores machines. | Swann Perarnau, Guillaume Huard |
| 2010 | Intra-application shared cache partitioning for multithreaded applications. | Sai Prashanth Muralidhara, Mahmut T. Kandemir, Padma Raghavan |
| 2010 | Structure-driven optimizations for amorphous data-parallel programs. | Mario Mndez-Lojo, Donald Nguyen, Dimitrios Prountzos, Xin Sui, Muhammad Amber Hassaan, Milind Kulkarni, Martin Burtscher, Keshav Pingali |
| 2010 | Effective communication and computation overlap with hybrid MPI/SMPSs. | Vladimir Marjanovic, Jess Labarta, Eduard Ayguad, Mateo Valero |
| 2010 | Compiler aided selective lock assignment for improving the performance of software transactional memory. | Sandya Mannarswamy, Dhruva R. Chakrabarti, Kaushik Rajan, Sujoy Saraswati |
| 2010 | Scheduling support for transactional memory contention management. | Walther Maldonado, Patrick Marlier, Pascal Felber, Adi Suissa, Danny Hendler, Alexandra Fedorova, Julia L. Lawall, Gilles Muller |
| 2010 | Towards scalable and transparent parallelization of multiplayer games using transactional memory support. | Daniel Lupei, Bogdan Simion, Don Pinto, Matthew Misler, Mihai Burcea, William Krick, Cristiana Amza |
| 2010 | Improving parallelism and locality with asynchronous algorithms. | Lixia Liu, Zhiyuan Li |
| 2010 | A symbolic verifier for CUDA programs. | Guodong Li, Ganesh Gopalakrishnan, Robert M. Kirby, Daniel J. Quinlan |
| 2010 | Featherweight X10: a core calculus for async-finish parallelism. | Jonathan K. Lee, Jens Palsberg |
| 2010 | Data transformations enabling loop vectorization on multithreaded data parallel architectures. | Byunghyun Jang, Perhaad Mistry, Dana Schaa, Rodrigo Dominguez, David R. Kaeli |
| 2010 | Load balancing on speed. | Steven Hofmeyr, Costin Iancu, Filip Blagojevic |
| 2010 | Application heartbeats for software performance and health. | Henry Hoffmann, Jonathan Eastep, Marco D. Santambrogio, Jason E. Miller, Anant Agarwal |
| 2010 | Scalable communication protocols for dynamic sparse data exchange. | Torsten Hoefler, Christian Siebert, Andrew Lumsdaine |
| 2010 | SLAW: a scalable locality-aware adaptive work-stealing scheduler for multi-core systems. | Yi Guo, Yisheng Zhao, Vincent Cav, Vivek Sarkar |
| 2010 | Symbolic prefetching in transactional distributed shared memory. | Alokika Dash, Brian Demsky |
| 2010 | NOrec: streamlining STM by abolishing ownership records. | Luke Dalessandro, Michael F. Spear, Michael L. Scott |
| 2010 | GAMBIT: effective unit testing for concurrency libraries. | Katherine E. Coons, Sebastian Burckhardt, Madanlal Musuvathi |
| 2010 | Model-driven autotuning of sparse matrix-vector multiply on GPUs. | JeeWhan Choi, Amik Singh, Richard W. Vuduc |
| 2010 | Applying the concurrent collections programming model to asynchronous parallel dense linear algebra. | Aparna Chandramowlishwaran, Kathleen Knobe, Richard W. Vuduc |
| 2010 | New abstractions for effective performance analysis of STM programs. | Dhruva R. Chakrabarti |
| 2010 | Supporting lock-free composition of concurrent data objects. | Daniel Cederman, Philippas Tsigas |
| 2010 | Scaling LAPACK panel operations using parallel cache assignment. | Anthony M. Castaldo, R. Clint Whaley |
| 2010 | The pilot library for novice MPI programmers. | John D. Carter, William B. Gardner, Gary Grwal |