| 2016 | The virtues of conflict: analysing modern concurrency. | Ganesh Narayanaswamy, Saurabh Joshi, Daniel Kroening |
| 2016 | Multitasking Real-time Embedded GPU Computing Tasks. | Pinar Muyan-zelik, John D. Owens |
| 2016 | Grain graphs: OpenMP performance analysis made easy. | Ananya Muddukrishna, Peter A. Jonsson, Artur Podobas, Mats Brorsson |
| 2016 | Auto-vectorizing a large-scale production unstructured-mesh CFD application. | Gihan R. Mudalige, Istvn Z. Reguly, Michael B. Giles |
| 2016 | On designing NUMA-aware concurrency control for scalable transactional memory. | Mohamed Mohamedin, Roberto Palmieri, Sebastiano Peluso, Binoy Ravindran |
| 2016 | Merge-based sparse matrix-vector multiplication (SpMV) using the CSR storage format. | Duane Merrill, Michael Garland |
| 2016 | Keep calm and react with foresight: strategies for low-latency and energy-efficient elastic data stream processing. | Tiziano De Matteis, Gabriele Mencagli |
| 2016 | Unifying fixed code and fixed data mapping of load-imbalanced pipelined loops. | Aristeidis Mastoras, Thomas R. Gross |
| 2016 | An Evaluation of Emerging Many-Core Parallel Programming Models. | Matt Martineau, Simon McIntosh-Smith, Michael Boulton, Wayne P. Gaudin |
| 2016 | DSMR: a shared and distributed memory algorithm for single-source shortest path problem. | Saeed Maleki, Donald Nguyen, Andrew Lenharth, Mara Jess Garzarn, David A. Padua, Keshav Pingali |
| 2016 | Concurrent hash tables: fast | Tobias Maier, Peter Sanders, Roman Dementiev |
| 2016 | Production-guided concurrency debugging. | Nuno Machado, Brandon Lucia, Lus E. T. Rodrigues |
| 2016 | Data-centric combinatorial optimization of parallel code. | Hao Luo, Guoyang Chen, Pengcheng Li, Chen Ding, Xipeng Shen |
| 2016 | Hybrid CPU-GPU scheduling and execution of tree traversals. | Jianqiao Liu, Nikhil Hegde, Milind Kulkarni |
| 2016 | Work stealing for interactive services to meet target latency. | Jing Li, Kunal Agrawal, Sameh Elnikety, Yuxiong He, I-Ting Angelina Lee, Chenyang Lu, Kathryn S. McKinley |
| 2016 | A new SIMD iterative connected component labeling algorithm. | Lionel Lacassagne, Laurent Cabaret, Daniel Etiemble, Farouk Hebache, Andrea Petreto |
| 2016 | User-assisted storage reuse determination for dynamic task graphs. | Mehmet Can Kurt, Bin Ren, Sriram Krishnamoorthy, Gagan Agrawal |
| 2016 | Code vectorization using Intel Array Notation. | Olaf Krzikalla, Georg Zitzlsberger |
| 2016 | Efficient Parallelization of Complex Automotive Systems. | Julian Kienberger, Christian Saad, Stefan Kuntz, Bernhard Bauer |
| 2016 | A high-performance parallel algorithm for nonnegative matrix factorization. | Ramakrishnan Kannan, Grey Ballard, Haesun Park |
| 2016 | DomLock: a new multi-granularity locking technique for hierarchies. | Saurabh Kalikar, Rupesh Nasre |
| 2016 | Enhancing Metaheuristic-based Virtual Screening Methods on Massively Parallel and Heterogeneous Systems. | Baldomero Imbernn, Jos M. Cecilia, Domingo Gimnez |
| 2016 | Flow Driven GPGPU Programming combining Textual and Graphical Programming. | Thomas Hoegg, Guenther Fiedler, Christian Koehler, Andreas Kolb |
| 2016 | SPIRIT: a runtime system for distributed irregular tree applications. | Nikhil Hegde, Jianqiao Liu, Milind Kulkarni |
| 2016 | Multi-stage programming for GPUs in C++ using PACXX. | Michael Haidl, Michel Steuwer, Tim Humernbrum, Sergei Gorlatch |