| 2012 | A work-stealing scheduler for X10's task parallelism with suspension. | Olivier Tardieu, Haichuan Wang, Haibo Lin |
| 2012 | Using GPU's to accelerate stencil-based computation kernels for the development of large scale scientific applications on heterogeneous systems. | Jian Tao, Marek Blazewicz, Steven R. Brandt |
| 2012 | PMA: Pixel-based multi-anchor algorithm for image recognition on multi-core systems. | Xiaoxin Tang, Long Zheng, Jun Ma, Yao Shen, Li Li, Minyi Guo |
| 2012 | Chestnut: a GPU programming language for non-experts. | Andrew Stromme, Ryan Carlson, Tia Newhall |
| 2012 | Establishing a Miniapp as a programmability proxy. | Andrew Stone, John M. Dennis, Michelle Strout |
| 2012 | Techniques for the parallelization of unstructured grid applications on multi-GPU systems. | Lizandro D. Solano-Quinde, Brett M. Bode, Arun K. Somani |
| 2012 | A performance analysis framework for identifying potential benefits in GPGPU applications. | Jaewoong Sim, Aniruddha Dasgupta, Hyesoon Kim, Richard W. Vuduc |
| 2012 | Faster topology-aware collective algorithms through non-minimal communication. | Paul Sack, William Gropp |
| 2012 | Efficient execution of time-step computations with pipelined parallelism and inter-thread data locality optimizaitions. | Apan Qasem |
| 2012 | Concurrent tries with efficient non-blocking snapshots. | Aleksandar Prokopec, Nathan Grasso Bronson, Phil Bagwell, Martin Odersky |
| 2012 | Concurrent breakpoints. | Chang-Seo Park, Koushik Sen |
| 2012 | Efficient memory management of a hierarchical and a hybrid main memory for MN-MATE platform. | Kyu Ho Park, Sung Kyu Park, Hyunchul Seok, Woomin Hwang, Dong-Jae Shin, Jong Hun Choi, Ki-Woong Park |
| 2012 | The boat hull model: adapting the roofline model to enable performance prediction for parallel computing. | Cedric Nugteren, Henk Corporaal |
| 2012 | An infrastructure for dynamic optimization of parallel programs. | Albert Noll, Thomas R. Gross |
| 2012 | Scalable parallel minimum spanning forest computation. | Sadegh Nobari, Thanh-Tung Cao, Panagiotis Karras, Stphane Bressan |
| 2012 | New strategy for coarse grid solvers in parallel multigrid methods using OpenMP/MPI hybrid programming models. | Kengo Nakajima |
| 2012 | Collective algorithms for sub-communicators. | Anshul Mittal, Nikhil Jain, Thomas George, Yogish Sabharwal, Sameer Kumar |
| 2012 | CPHASH: a cache-partitioned hash table. | Zviad Metreveli, Nickolai Zeldovich, M. Frans Kaashoek |
| 2012 | Scalable GPU graph traversal. | Duane Merrill, Michael Garland, Andrew S. Grimshaw |
| 2012 | A GPU implementation of inclusion-based points-to analysis. | Mario Mndez-Lojo, Martin Burtscher, Keshav Pingali |
| 2012 | Mechanizing the expert dense linear algebra developer. | Bryan Marker, Andy Terrel, Jack Poulson, Don S. Batory, Robert A. van de Geijn |
| 2012 | Verification of software barriers. | Alexander Malkis, Anindya Banerjee |
| 2012 | A lock-free, array-based priority queue. | Yujie Liu, Michael F. Spear |
| 2012 | FlexBFS: a parallelism-aware implementation of breadth-first search on GPU. | Gu Liu, Hong An, Wenting Han, Xiaoqiang Li, Tao Sun, Wei Zhou, Xuechao Wei, Xulong Tang |
| 2012 | GKLEE: concolic verification and test generation for GPUs. | Guodong Li, Peng Li, Geoffrey Sawaya, Ganesh Gopalakrishnan, Indradeep Ghosh, Sreeranga P. Rajan |