| 2016 | Effective Dynamic Load Balance using Space-Filling Curves for Large-Scale SPH Simulations on GPU-rich Supercomputers. | Satori Tsuzuki, Takayuki Aoki |
| 2016 | Performance-Portable Autotuning of OpenCL Kernels for Convolutional Layers of Deep Neural Networks. | Yaohung M. Tsai, Piotr Luszczek, Jakub Kurzak, Jack J. Dongarra |
| 2016 | LLVM Framework and IR Extensions for Parallelization, SIMD Vectorization and Offloading. | Xinmin Tian, Hideki Saito, Ernesto Su, Abhinav Gaba, Matt Masten, ric Garcia, Ayal Zaks |
| 2016 | Topology-Aware Data Aggregation for Intensive I/O on Large-Scale Supercomputers. | Francois Tessier, Preeti Malakar, Venkatram Vishwanath, Emmanuel Jeannot, Florin Isaila |
| 2016 | In Situ Statistical Analysis for Parametric Studies. | Thophile Terraz, Bruno Raffin, Alejandro Ribs, Yvan Fournier |
| 2016 | Extreme scale plasma turbulence simulations on top supercomputers worldwide. | William M. Tang, Bei Wang, Stphane Ethier, Grzegorz Kwasniewski, Torsten Hoefler, Khaled Z. Ibrahim, Kamesh Madduri, Samuel Williams, Leonid Oliker, Carlos Rosales-Fernandez, Timothy J. Williams |
| 2016 | Elastic multi-resource fairness: balancing fairness and efficiency in coupled CPU-GPU architectures. | Shanjiang Tang, Bingsheng He, Shuhao Zhang, Zhaojie Niu |
| 2016 | Design of Fault Tolerant Pwrake Workflow System Supported by Gfarm File System. | Masahiro Tanaka, Osamu Tatebe |
| 2016 | In-Staging Data Placement for Asynchronous Coupling of Task-Based Scientific Workflows. | Qian Sun, Melissa Romanus, Tong Jin, Hongfeng Yu, Peer-Timo Bremer, Steve Petruzza, Scott Klasky, Manish Parashar |
| 2016 | An Overview of Performance Portability in the Uintah Runtime System through the Use of Kokkos. | Daniel Sunderland, Brad Peterson, John A. Schmidt, Alan Humphrey, Jeremy Thornock, Martin Berzins |
| 2016 | Keynote: The quantum step in parallel execution through dynamic adaptive runtime and programming strategies. | Thomas L. Sterling |
| 2016 | Efficient Parallelization of MATLAB Stencil Applications for Multi-core Clusters. | Johannes Spazier, Steffen Christgau, Bettina Schnor |
| 2016 | Online Input Data Reduction in Scientific Workflows. | Renan Souza, Vtor Silva, Alvaro L. G. A. Coutinho, Patrick Valduriez, Marta Mattoso |
| 2016 | Integrating Domain-Data Steering with Code-Profiling Tools to Debug Data-Intensive Workflows. | Vtor Silva, Leonardo Neves, Renan Souza, Alvaro L. G. A. Coutinho, Daniel de Oliveira, Marta Mattoso |
| 2016 | Modular HPC I/O Characterization with Darshan. | Shane Snyder, Philip H. Carns, Kevin Harms, Robert B. Ross, Glenn K. Lockwood, Nicholas J. Wright |
| 2016 | An exploration of optimization algorithms for high performance tensor completion. | Shaden Smith, Jongsoo Park, George Karypis |
| 2016 | Advantages, Disadvantages and Misunderstandings About Document Driven Design for Scientific Software. | Spencer Smith, Thulasi Jegatheesan, Diane Kelly |
| 2016 | Performance of MPI Codes Written in Python with NumPy and mpi4py. | Ross Smith |
| 2016 | Klimatic: A Virtual Data Lake for Harvesting and Distribution of Geospatial Data. | Tyler J. Skluzacek, Kyle Chard, Ian T. Foster |
| 2016 | SLA-aware Interactive Workflow Assistant for HPC Parameter Sweeping Experiments. | Bruno Silva, Marco Aurlio Stelmar Netto, Renato Luiz de Freitas Cunha |
| 2016 | Using Simple PID Controllers to Prevent and Mitigate Faults in Scientific Workflows. | Rafael Ferreira da Silva, Rosa Filgueira, Ewa Deelman, Erola Pairo-Castineira, Ian Michael Overton, Malcolm P. Atkinson |
| 2016 | Energy-efficient Mapping of Big Data Workflows under Deadline Constraints. | Tong Shu, Chase Qishi Wu |
| 2016 | Multi-scale CAFE Framework for Simulating Fracture in Heterogeneous Materials Implemented in Fortran Co-arrays and MPI. | Anton Shterenlikht, Lee Margetts, Jose D. Arregui-Mena, Luis Cebamanos |
| 2016 | Single-Sided Statistic Multiplexed High Performance Computing. | Justin Y. Shi, Yasin Celik |
| 2016 | Transient guarantees: maximizing the value of idle cloud capacity. | Supreeth Shastri, Amr Rizk, David Irwin |