| 2016 | Runtime aware architectures. | Mateo Valero |
| 2016 | Effect of portable fine-grained locality on energy efficiency and performance in concurrent search trees. | Ibrahim Umar, Otto J. Anshus, Phuong Hoai Ha |
| 2016 | Coarse grain parallelization of deep neural networks. | Marc Gonzlez Tallada |
| 2016 | Adding approximate counters. | Guy L. Steele Jr., Jean-Baptiste Tristan |
| 2016 | Multi-GPU implementation of the Horizontal Diffusion method of the Weather Research and Forecast Model. | Lizandro D. Solano-Quinde, Ronald Gualan-Saavedra, Miguel Ziga-Prieto |
| 2016 | JParEnt: Parallel Entropy Decoding for JPEG Decompression on Heterogeneous Multicore Architectures. | Wasuwee Sodsong, Minyoung Jung, Jinwoo Park, Bernd Burgstaller |
| 2016 | Generic messages: capability-based shared memory parallelism for event-loop systems. | Luca Salucci, Daniele Bonetta, Stefan Marr, Walter Binder |
| 2016 | On ordering transaction commit. | Mohamed M. Saad, Roberto Palmieri, Binoy Ravindran |
| 2016 | Benchmarking weak memory models. | Carl G. Ritson, Scott Owens |
| 2016 | Working together to build the heterogeneous processing ecosystem. | Andrew Richards |
| 2016 | Performance portable GPU code generation for matrix multiplication. | Toomas Remmelg, Thibaut Lutz, Michel Steuwer, Christophe Dubach |
| 2016 | Verification of MPI Java programs using software model checking. | Waqas ur Rehman, Muhammad Sohaib Ayub, Junaid Haroon Siddiqui |
| 2016 | Effective resource management for enhancing performance of 2D and 3D stencils on GPUs. | Prashant Singh Rawat, Changwan Hong, Mahesh Ravishankar, Vinod Grover, Louis-Nol Pouchet, P. Sadayappan |
| 2016 | Tidex: a mutual exclusion lock. | Pedro Ramalhete, Andreia Correia |
| 2016 | Improving efficacy of internal binary search trees using local recovery. | Arunmoezhi Ramachandran, Neeraj Mittal |
| 2016 | Preemption-aware planning on big-data systems. | Marco Rabozzi, Matteo Mazzucchelli, Roberto Cordone, Giovanni Matteo Fumarola, Marco D. Santambrogio |
| 2016 | OPR: deterministic group replay for one-sided communication. | Xuehai Qian, Koushik Sen, Paul Hargrove, Costin Iancu |
| 2016 | Implementing directed acyclic graphs with the heterogeneous system architecture. | Sooraj Puthoor, Ashwin M. Aji, Shuai Che, Mayank Daga, Wei Wu, Bradford M. Beckmann, Gregory Rodgers |
| 2016 | CUDA acceleration for Xen virtual machines in infiniband clusters with rCUDA. | Javier Prades, Carlos Reao, Federico Silla |
| 2016 | An evaluation of current SIMD programming models for C++. | Angela Pohl, Biagio Cosenza, Mauricio Alvarez-Mesa, Chi Ching Chi, Ben H. H. Juurlink |
| 2016 | Causal consistency: beyond memory. | Matthieu Perrin, Achour Mostfaoui, Claude Jard |
| 2016 | Simplifying programming and load balancing of data parallel applications on heterogeneous systems. | Borja Prez, Jos Luis Bosque, Ramn Beivide |
| 2016 | Efficient distributed workstealing via matchmaking. | Hrushit Parikh, Vinit Deodhar, Ada Gavrilovska, Santosh Pande |
| 2016 | A scalable lock-free hash table with open addressing. | Jesper Puge Nielsen, Sven Karlsson |
| 2016 | Parallel type-checking with haskell using saturating LVars and stream generators. | Ryan R. Newton, mer S. Agacan, Peter P. Fogg, Sam Tobin-Hochstadt |