| 2016 | Insights into the Fallback Path of Best-Effort Hardware Transactional Memory Systems. | Ricardo Quislant, Eladio Gutirrez, Emilio L. Zapata, Oscar G. Plata |
| 2016 | Piecewise Holistic Autotuning of Compiler and Runtime Parameters. | Mihail Popov, Chadi Akel, William Jalby, Pablo de Oliveira Castro |
| 2016 | Are Low-Power SoCs Feasible for Heterogenous HPC Workloads? | Max Plauth, Andreas Polze |
| 2016 | Parallel String Matching. | Philip Pfaffe, Martin Peter Tillmann, Sarah Lutteropp, Bernhard Scheirle, Kevin Zerr |
| 2016 | An Autonomic Parallel Strategy for the Projection of Ecological Niche Models in Heterogeneous Computational Environments. | Fernanda G. O. Passos, Vinod E. F. Rebello |
| 2016 | HAP: A Heterogeneity-Conscious Runtime System for Adaptive Pipeline Parallelism. | Jinsu Park, Woongki Baek |
| 2016 | High Performance Parallel Summed-Area Table Kernels for Multi-core and Many-core Systems. | Angelos Papatriantafyllou, Dimitris Sacharidis |
| 2016 | Speed-Up Computational Finance Simulations with OpenCL on Intel Xeon Phi. | Michail Papadimitriou, Joris Cramwinckel, Ana Lucia Varbanescu |
| 2016 | High Performance Small RNA Detection with Pipelined Task Parallel Computation Model. | Linqiang Ouyang, Jin H. Park |
| 2016 | Heating as a Cloud-Service, A Position Paper (Industrial Presentation). | Yanik Ngoko |
| 2016 | In-Cache Streaming: Morphable Infrastructure for Many-Core Processing Systems. | Nuno Neves, Adrien Mussio, Fabien Gonalves, Pedro Toms, Nuno Roma |
| 2016 | Lattice Boltzmann Flow Simulation on Android Devices for Interactive Mobile-Based Learning. | Philipp Neumann, Michael Zellner |
| 2016 | Improving Bioinformatics Analysis of Large Sequence Datasets Parallelizing Tools for Population Genomics. | Javier Navarro, Gonzalo Vera, Sebastin Ramos-Onsins, Porfidio Hernndez |
| 2016 | A Cooperative Approach to Virtual Machine Based Fault Injection. | Thomas J. Naughton, Christian Engelmann, Geoffroy Valle, Ferrol Aderholdt, Stephen L. Scott |
| 2016 | Design and Verification of Distributed Phasers. | Karthik Murthy, Sri Raj Paul, Kuldeep S. Meel, Tiago Cogumbreiro, John M. Mellor-Crummey |
| 2016 | Distributed In-GPU Data Cache for Document-Oriented Data Store via PCIe over 10 Gbit Ethernet. | Shin Morishima, Hiroki Matsutani |
| 2016 | On the Inherent Resilience of Integer Operations. | Laura Monroe, William M. Jones, Scott R. Lavigne, Claude H. Davis IV, Qiang Guan, Nathan DeBardeleben |
| 2016 | Computation-Aware Dynamic Frequency Scaling: Parsimonious Evaluation of the Time-Energy Trade-Off Using Design of Experiments. | Luis Felipe Millani, Lucas Mello Schnorr |
| 2016 | Balancing Speedup and Accuracy in Smart City Parallel Applications. | Carlo Mastroianni, Eugenio Cesario, Andrea Giordano |
| 2016 | High-Performance Matrix-Matrix Multiplications of Very Small Matrices. | Ian Masliah, Ahmad Abdelfattah, Azzam Haidar, Stanimire Tomov, Marc Baboulin, Jol Falcou, Jack J. Dongarra |
| 2016 | Automatic Benchmark Profiling Through Advanced Trace Analysis. | Alexis Martin, Vania Marangozova-Martin |
| 2016 | Theano-MPI: A Theano-Based Distributed Training Framework. | He Ma, Fei Mao, Graham W. Taylor |
| 2016 | Toward a General I/O Arbitration Framework for netCDF Based Big Data Processing. | Jianwei Liao, Balazs Gerofi, Guo-Yuan Lien, Seiya Nishizawa, Takemasa Miyoshi, Hirofumi Tomita, Yutaka Ishikawa |
| 2016 | Optimized Execution Strategies for Sequence Aligners on NUMA Architectures. | Josefina Lenis, Miquel ngel Senar |
| 2016 | Performance Prediction and Ranking of SpMV Kernels on GPU Architectures. | Christoph Lehnert, Rudolf Berrendorf, Jan P. Ecker, Florian Mannuss |