| 2016 | Computational Efficiency vs. Maintainability and Portability. Experiences with the Sparse Grid Code SG++. | Dirk Pflger, David Pfander |
| 2016 | The Scalability-Efficiency/Maintainability-Portability Trade-Off in Simulation Software Engineering: Examples and a Preliminary Systematic Literature Review. | Dirk Pflger, Miriam Mehl, Julian Valentin, Florian Lindner, David Pfander, Stefan Wagner, Daniel Graziotin, Yang Wang |
| 2016 | A Multi-tenant Fair Share Approach to Full-text Search Engine. | Zong Peng, Beth Plale |
| 2016 | Static Cost Estimation for Data Layout Selection on GPUs. | Yuhan Peng, Max Grossman, Vivek Sarkar |
| 2016 | Enabling Work Migration in CoMD to Study Dynamic Load Imbalance Solutions. | Olga Pearce, Hadia Ahmed, Rasmus W. Larsen, David F. Richards |
| 2016 | Economic Viability of Hardware Overprovisioning in Power-Constrained High Performance Computing. | Tapasya Patki, David K. Lowenthal, Barry L. Rountree, Martin Schulz, Bronis R. de Supinski |
| 2016 | Application of PGAS Programming to Power Grid Simulation. | Bruce Palmer |
| 2016 | High level abstractions and automatic optimization techniques for the programming of irregular algorithms. | David A. Padua |
| 2016 | Scientific Workflows at DataWarp-Speed: Accelerated Data-Intensive Science Using NERSC's Burst Buffer. | Andrey Ovsyannikov, Melissa Romanus, Brian van Straalen, Gunther H. Weber, David Trebotich |
| 2016 | Towards Fast Scalable Solvers for Charge Equilibration in Molecular Dynamics Applications. | Kurt A. O'Hearn, Hasan Metin Aktulga |
| 2016 | Distributed Multithreaded Breadth-First Search on Large Graphs Using DXGraph. | Stefan Nothaas, Kevin Beineke, Michael Schttner |
| 2016 | FlipBack: automatic targeted protection against silent data corruption. | Xiang Ni, Laxmikant V. Kal |
| 2016 | VIPACT: A Visualization Interface for Analyzing Calling Context Trees. | Huu Tan Nguyen, Lai Wei, Abhinav Bhatele, Todd Gamblin, David Bhme, Martin Schulz, Kwan-Liu Ma, Peer-Timo Bremer |
| 2016 | Perilla: metadata-based optimizations of an asynchronous runtime for adaptive mesh refinement. | Tan Nguyen, Didem Unat, Weiqun Zhang, Ann S. Almgren, Muhammed Nufail Farooqi, John Shalf |
| 2016 | A Case Study: Test-Driven Development in a Microscopy Image-Processing Project. | Aziz Nanthaamornphong |
| 2016 | Simulations of below-ground dynamics of fungi: 1.184 pflops attained by automated generation and autotuning of temporal blocking codes. | Takayuki Muranushi, Hideyuki Hotta, Junichiro Makino, Seiya Nishizawa, Hirofumi Tomita, Keigo Nitadori, Masaki Iwasawa, Natsuki Hosono, Yutaka Maruyama, Hikaru Inoue, Hisashi Yashiro, Yoshifumi Nakamura |
| 2016 | Code Complexity versus Performance for GPU-accelerated Scientific Applications. | W. K. Umayanganie Munipala, Shirley V. Moore |
| 2016 | The ARES High-Level Intermediate Representation. | Nick Moss, Kei Davis, Patrick S. McCormick |
| 2016 | Performance Scaling Variability and Energy Analysis for a Resilient ULFM-based PDE Solver. | Karla Morris, Francesco Rizzi, Brendan Cook, Paul Mycek, Olivier P. Le Matre, Omar M. Knio, Khachik Sargsyan, Kathryn Dahlgren, Bert J. Debusschere |
| 2016 | Efficient delaunay tessellation through K-D tree decomposition. | Dmitriy Morozov, Tom Peterka |
| 2016 | Applications of the FACE-IT Data Science Portal and Workflow Engine for Operational Food Quality Prediction and Assessment: Mussel Farm Monitoring in the Bay of Napoli, Italy. | Raffaele Montella, Alison Brizius, Diana Di Luccio, Cheryl H. Porter, Joshua Elliott, Ravi K. Madduri, David Kelly, Angelo Riccio, Ian T. Foster |
| 2016 | Clustering Based on Task Dependency for Data-Intensive Workflow Scheduling Optimization. | Ei Ei Mon, Myint Myint Thein, May Thu Aung |
| 2016 | Understanding performance interference in next-generation HPC systems. | Oscar H. Mondragon, Patrick G. Bridges, Scott Levy, Kurt B. Ferreira, Patrick M. Widener |
| 2016 | An Extension of OpenACC Directives for Out-of-Core Stencil Computation with Temporal Blocking. | Nobuhiro Miki, Fumihiko Ino, Kenichi Hagihara |
| 2016 | Merge-based parallel sparse matrix-vector multiplication. | Duane Merrill, Michael Garland |