| 2012 | Extending a C-like language for portable SIMD programming. | Roland Leia, Sebastian Hack, Ingo Wald |
| 2012 | Deterministic parallel random-number generation for dynamic-multithreading platforms. | Charles E. Leiserson, Tao B. Schardl, Jim Sukha |
| 2012 | Revisiting shared virtual memory systems for non-coherent memory-coupled cores. | Stefan Lankes, Pablo Reble, Oliver Sinnen, Carsten Clauss |
| 2012 | A hybrid approach of OpenMP for clusters. | Okwan Kwon, Fahed Jubair, Rudolf Eigenmann, Samuel P. Midkiff |
| 2012 | Synchronization views for event-loop actors. | Joeri De Koster, Stefan Marr, Theo D'Hondt |
| 2012 | A methodology for creating fast wait-free data structures. | Alex Kogan, Erez Petrank |
| 2012 | Automatic datatype generation and optimization. | Fredrik Kjolstad, Torsten Hoefler, Marc Snir |
| 2012 | OpenCL as a unified programming model for heterogeneous CPU/GPU clusters. | Jungwon Kim, Sangmin Seo, Jun Lee, Jeongho Nah, Gangwon Jo, Jaejin Lee |
| 2012 | GHOST: GPGPU-offloaded high performance storage I/O deduplication for primary storage system. | Chulmin Kim, Ki-Woong Park, Kyu Ho Park |
| 2012 | Efficient SIMD code generation for irregular kernels. | Seonggun Kim, Hwansoo Han |
| 2012 | Portable parallel performance from sequential, productive, embedded domain-specific languages. | Shoaib Kamil, Derrick Coetzee, Scott Beamer, Henry Cook, Ekaterina Gonina, Jonathan Harper, Jeffrey Morlan, Armando Fox |
| 2012 | Networks beat pipelines: the design of FG 2.0. | Peter C. Johnson, Thomas H. Cormen |
| 2012 | Massively parallel breadth first search using a tree-structured memory model. | Tom St. John, Jack B. Dennis, Guang R. Gao |
| 2012 | Adapting the polyhedral model as a framework for efficient speculative parallelization. | Alexandra Jimborean, Philippe Clauss, Benot Pradelle, Luis Mastrangelo, Vincent Loechner |
| 2012 | OpenMP-style parallelism in data-centered multicore computing with R. | Lei Jiang, Pragneshkumar B. Patel, George Ostrouchov, Ferdinand Jamitzky |
| 2012 | Scalable framework for mapping streaming applications onto multi-GPU systems. | Huynh Phung Huynh, Andrei Hagiescu, Weng-Fai Wong, Rick Siow Mong Goh |
| 2012 | Communication-centric optimizations by dynamically detecting collective operations. | Torsten Hoefler, Timo Schneider |
| 2012 | Semi-sparse algorithm based on multi-layer optimization for recommendation system. | Hu Guan, Huakang Li, Minyi Guo |
| 2012 | An overview of CMPI: network performance aware MPI in the cloud. | Yifan Gong, Bingsheng He, Jianlong Zhong |
| 2012 | Speculative parallelization on GPGPUs. | Min Feng, Rajiv Gupta, Laxmi N. Bhuyan |
| 2012 | Revisiting the combining synchronization technique. | Panagiota Fatourou, Nikolaos D. Kallimanis |
| 2012 | Function flow: making synchronization easier in task parallelism. | Xuepeng Fan, Hai Jin, Liang Zhu, Xiaofei Liao, Chencheng Ye, Xuping Tu |
| 2012 | DOJ: dynamically parallelizing object-oriented programs. | Yong Hun Eom, Stephen Yang, James Christopher Jenista, Brian Demsky |
| 2012 | Better speedups using simpler parallel programming for graph connectivity and biconnectivity. | James Alexander Edwards, Uzi Vishkin |
| 2012 | Kokkos Array performance-portable manycore programming model. | H. Carter Edwards, Daniel Sunderland |