| 2016 | Towards an Empirical Study Design for Concurrent Software Testing. | Silvana M. Melo, Paulo S. L. Souza, Simone R. S. Souza |
| 2016 | Asynchronous In Situ Connected-Components Analysis for Complex Fluid flows. | James E. McClure, Mark A. Berrill, Jan F. Prins, Cass T. Miller |
| 2016 | Left-Preconditioned Communication-Avoiding Conjugate Gradient Methods for Multiphase CFD Simulations on the K Computer. | Akie Mayumi, Yasuhiro Idomura, Takuya Ina, Susumu Yamada, Toshiyuki Imamura |
| 2016 | PGAS Communication Runtime for Extreme Large Data Computation. | Ryo Matsumiya, Toshio Endo |
| 2016 | Tr: blob storage meets built-in transactions. | Pierre Matri, Alexandru Costan, Gabriel Antoniu, Jess Montes, Mara S. Prez |
| 2016 | Simulation and performance analysis of the ECMWF tape library system. | Markus Msker, Lars Nagel, Tim S, Andr Brinkmann, Lennart Sorth |
| 2016 | Performance Analysis and Optimization of Clang's OpenMP 4.5 GPU Support. | Matt Martineau, Simon McIntosh-Smith, Carlo Bertolli, Arpith C. Jacob, Samuel F. Anto, Alexandre E. Eichenberger, Gheorghe-Teodor Bercea, Tong Chen, Tian Jin, Kevin O'Brien, Georgios Rokos, Hyojin Sung, Zehra Sura |
| 2016 | A PCIe congestion-aware performance model for densely populated accelerator servers. | Maxime Martinasso, Grzegorz Kwasniewski, Sadaf R. Alam, Thomas C. Schulthess, Torsten Hoefler |
| 2016 | HPC Benchmarking: Problem Size Matters. | Vladimir Marjanovic, Jos Gracia, Colin W. Glass |
| 2016 | Distributed-memory large deformation diffeomorphic 3D image registration. | Andreas Mang, Amir Gholami, George Biros |
| 2016 | Towards Serverless Execution of Scientific Workflows - HyperFlow Case Study. | Maciej Malawski |
| 2016 | Optimal execution of co-analysis for large-scale molecular dynamics simulations. | Preeti Malakar, Venkatram Vishwanath, Christopher Knight, Todd S. Munson, Michael E. Papka |
| 2016 | Automatic Mapping of Array Operations to Specific Architectures. | Simon Andreas Frimann Lund, Mads Ruben Burgdorff Kristensen, Brian Vinter |
| 2016 | Mrs: High Performance MapReduce for Iterative and Asynchronous Algorithms in Python. | Jeffrey Lund, Chace Ashcraft, Andrew W. McNabb, Kevin D. Seppi |
| 2016 | Towards Achieving Performance Portability Using Directives for Accelerators. | M. Graham Lopez, Vernica G. Vergara Larrea, Wayne Joubert, Oscar R. Hernandez, Azzam Haidar, Stanimire Tomov, Jack J. Dongarra |
| 2016 | DAOS and friends: a proposal for an exascale storage system. | Jay F. Lofstead, Ivo Jimenez, Carlos Maltzahn, Quincey Koziol, John Bent, Eric Barton |
| 2016 | Optimizing memory efficiency for deep convolutional neural networks on GPUs. | Chao Li, Yi Yang, Min Feng, Srimat T. Chakradhar, Huiyang Zhou |
| 2016 | Sanity Tool: Lightweight Diagnostics for Individual User Accounts on Supercomputer Systems. | Si Liu, Robert T. McLay, Doug James |
| 2016 | Compiler-directed lightweight checkpointing for fine-grained guaranteed soft error recovery. | Qingrui Liu, Changhee Jung, Dongyoon Lee, Devesh Tiwari |
| 2016 | An Efficient Parallel Trust-Based Recommendation Method on Multicores. | Huafeng Liu, Liping Jing, MiaoMiao Cheng |
| 2016 | Server-side log data analytics for I/O workload characterization and coordination on large shared storage systems. | Yang Liu, Raghul Gunasekaran, Xiaosong Ma, Sudharshan S. Vazhkudai |
| 2016 | 20 Years of Teaching Parallel Processing to Computer Science Seniors. | Jie Liu |
| 2016 | Improved Data-Aware Task Dispatching for Batch Queuing Systems. | Xieming Li, Osamu Tatebe |
| 2016 | Learning to Diagnose Stragglers in Distributed Computing. | Cong Li, Huanxing Shen, Tai Huang |
| 2016 | Boosting Python Performance on Intel Processors: A Case Study of Optimizing Music Recognition. | Yuanzhe Li, Loren Schwiebert |