| 2016 | HISC/R: An Efficient Hypersparse-Matrix Storage Format for Scalable Graph Processing. | Robert Kirchgessner, Giovanni De La Torre, Alan D. George, Vitaliy Gleyzer |
| 2016 | Translating OpenMP device constructs to OpenCL using unnecessary data transfer elimination. | Junghyun Kim, Yong-Jun Lee, Jung-Ho Park, Jaejin Lee |
| 2016 | A Massively Parallel Distributed N-body Application Implemented with HPX. | Zahra Khatami, Hartmut Kaiser, Patricia Grubel, Adrian Serio, J. Ramanujam |
| 2016 | Designing scalable | Arif M. Khan, Alex Pothen, Md. Mostofa Ali Patwary, Mahantesh Halappanavar, Nadathur Rajagopalan Satish, Narayanan Sundaram, Pradeep Dubey |
| 2016 | Towards Automatic HBM Allocation Using LLVM: A Case Study with Knights Landing. | Dounia Khaldi, Barbara M. Chapman |
| 2016 | Distributed Training of Deep Neural Networks: Theoretical and Practical Limits of Parallel Scalability. | Janis Keuper, Franz-Josef Pfreundt |
| 2016 | PFEAST: a high performance sparse eigenvalue solver using distributed-memory linear solvers. | James Kestyn, Vasileios Kalantzis, Eric Polizzi, Yousef Saad |
| 2016 | In-Situ Visual Exploration of Multivariate Volume Data Based on Particle Based Volume Rendering. | Takuma Kawamura, Tomoyuki Noda, Yasuhiro Idomura |
| 2016 | Measuring and understanding throughput of network topologies. | Sangeetha Abdu Jyothi, Ankit Singla, Brighten Godfrey, Alexandra Kolla |
| 2016 | Block iterative methods and recycling for improved scalability of linear solvers. | Pierre Jolivet, Pierre-Henri Tournier |
| 2016 | Topology and Affinity Aware Hierarchical and Distributed Load-Balancing in Charm++. | Emmanuel Jeannot, Guillaume Mercier, Francois Tessier |
| 2016 | Evaluating HPC networks via simulation of parallel workloads. | Nikhil Jain, Abhinav Bhatele, Sam White, Todd Gamblin, Laxmikant V. Kal |
| 2016 | Randomized Sketching for Large-Scale Sparse Ridge Regression Problems. | Chander Iyer, Christopher D. Carothers, Petros Drineas |
| 2016 | A machine learning framework for performance coverage analysis of proxy applications. | Tanzima Z. Islam, Jayaraman J. Thiagarajan, Abhinav Bhatele, Martin Schulz, Todd Gamblin |
| 2016 | Power-Efficient Breadth-First Search with DRAM Row Buffer Locality-Aware Address Mapping. | Satoshi Imamura, Yuichiro Yasui, Koji Inoue, Takatsugu Ono, Hiroshi Sasaki, Katsuki Fujisawa |
| 2016 | Strassen's algorithm reloaded. | Jianyu Huang, Tyler M. Smith, Greg M. Henry, Robert A. van de Geijn |
| 2016 | DCA: a DRAM-cache-aware DRAM controller. | Cheng-Chieh Huang, Vijay Nagarajan, Arpit Joshi |
| 2016 | Towards Automatic and Flexible Unit Test Generation for Legacy HPC Code. | Christian Hovy, Julian M. Kunkel |
| 2016 | Extremely Scalable Algorithm for 108-atom Quantum Material Simulation on the Full System of the K Computer. | Takeo Hoshi, Hiroto Imachi, Kiyoshi Kumahata, Masaaki Terai, Kengo Miyamoto, Kazuo Minami, Fumiyoshi Shoji |
| 2016 | Metaprogramming-Enabled Parallel Execution of Apparently Sequential C++ Code. | David S. Hollman, Janine C. Bennett, Hemanth Kolla, Jonathan Lifflander, Nicole Slattengren, Jeremiah J. Wilke |
| 2016 | The vectorization of the tersoff multi-body potential: an exercise in performance portability. | Markus Hhnerbach, Ahmed E. Ismail, Paolo Bientinesi |
| 2016 | An ECM-based Energy-Efficiency Optimization Approach for Bandwidth-Limited Streaming Kernels on Recent Intel Xeon Processors. | Johannes Hofmann, Dietmar Fey |
| 2016 | The AllScale Runtime Interface - Theoretical Foundation and Concept. | Arne Hendricks, Thomas Heller, Herbert Jordan, Peter Thoman, Thomas Fahringer, Dietmar Fey |
| 2016 | Dynamic Load Balancing for High-Performance Graph Processing on Hybrid CPU-GPU Platforms. | Stijn Heldens, Ana Lucia Varbanescu, Alexandru Iosup |
| 2016 | MetaMorph: a library framework for interoperable kernels on multi- and many-core clusters. | Ahmed E. Helal, Paul Sathre, Wu-chun Feng |