| 2016 | Lease/release: architectural support for scaling contended data structures. | Syed Kamran Haider, William Hasenplaugh, Dan Alistarh |
| 2016 | Optimistic concurrency with OPTIK. | Rachid Guerraoui, Vasileios Trigonakis |
| 2016 | Parallel Locality and Parallelization Quality. | Bernard Goossens, David Parello, Katarzyna Porada, Djallal Rahmoune |
| 2016 | An interval constrained memory allocator for the Givy GAS runtime. | Franois Gindraud, Fabrice Rastello, Albert Cohen, Franois Broquedis |
| 2016 | Support for data parallelism in the CAL actor language. | Essayas Gebrewahid, Mehmet Ali Arslan, Andreas Karlsson, Zain-ul-Abdin |
| 2016 | On Guided Installation of Basic Linear Algebra Routines in Nodes with Manycore Components. | Luis-Pedro Garca, Javier Cuenca, Francisco-Jos Herrera, Domingo Gimnez |
| 2016 | Affinity-aware work-stealing for integrated CPU-GPU processors. | Naila Farooqui, Rajkishore Barik, Brian T. Lewis, Tatiana Shpeisman, Karsten Schwan |
| 2016 | A systems perspective on GPU computing: a tribute to Karsten Schwan. | Naila Farooqui |
| 2016 | NUMA-aware scheduling and memory allocation for data-flow task-parallel applications. | Andi Drebes, Antoniu Pop, Karine Heydemann, Nathalie Drach, Albert Cohen |
| 2016 | Embedding Semantics of the Single-Producer/Single-Consumer Lock-Free Queue into a Race Detection Tool. | Manuel F. Dolz, David del Rio Astorga, Javier Fernndez, Jos Daniel Garca, Flix Garca Carballeira, Marco Danelutto, Massimo Torquati |
| 2016 | Accelerating Dynamic Data Race Detection Using Static Thread Interference Analysis. | Peng Di, Yulei Sui |
| 2016 | Refined transactional lock elision. | Dave Dice, Alex Kogan, Yossi Lev |
| 2016 | GPU centric extensions for parallel strongly connected components computation. | Shrinivas Devshatwar, Madhur Amilkanthwar, Rupesh Nasre |
| 2016 | Distributed Halide. | Tyler Denniston, Shoaib Kamil, Saman P. Amarasinghe |
| 2016 | Declarative coordination of graph-based parallel programs. | Flvio Cruz, Ricardo Rocha, Seth Copen Goldstein |
| 2016 | Software-managed Cache Coherence for fast One-Sided Communication. | Steffen Christgau, Bettina Schnor |
| 2016 | AUTOGEN: automatic discovery of cache-oblivious parallel recursive algorithms for solving dynamic programs. | Rezaul Alam Chowdhury, Pramod Ganapathi, Jesmin Jahan Tithi, Charles Bachmeier, Bradley C. Kuszmaul, Charles E. Leiserson, Armando Solar-Lezama, Yuan Tang |
| 2016 | Samsara parallel: a non-BSP parallel-in-time model. | Yifeng Chen, Kun Huang, Bei Wang, Guohui Li, Xiang Cui |
| 2016 | ESTIMA: extrapolating scalability of in-memory applications. | Georgios Chatzopoulos, Aleksandar Dragojevic, Rachid Guerraoui |
| 2016 | A programming system for future proofing performance critical libraries. | Li-Wen Chang, Izzat El Hajj, Hee-Seok Kim, Juan Gmez-Luna, Abdul Dakkak, Wen-mei W. Hwu |
| 2016 | Contention-conscious, locality-preserving locks. | Milind Chabbi, John M. Mellor-Crummey |
| 2016 | Drinking from both glasses: combining pessimistic and optimistic tracking of cross-thread dependences. | Man Cao, Minjia Zhang, Aritra Sengupta, Michael D. Bond |
| 2016 | Multi-core on-the-fly SCC decomposition. | Vincent Bloemen, Alfons Laarman, Jaco van de Pol |
| 2016 | Designing high performance communication runtime for GPU managed memory: early experiences. | Dip Sankar Banerjee, Khaled Hamidouche, Dhabaleswar K. Panda |
| 2016 | Discovering Pipeline Parallel Patterns in Sequential Legacy C++ Codes. | David del Rio Astorga, Manuel F. Dolz, Luis Miguel Snchez, Jos Daniel Garca |