| 2016 | Understanding error propagation in GPGPU applications. | Guanpeng Li, Karthik Pattabiraman, Chen-Yong Cher, Pradip Bose |
| 2016 | Enabling efficient preemption for SIMT architectures with lightweight context switching. | Zhen Lin, Lars Nyland, Huiyang Zhou |
| 2016 | Optimizing Sparse Tensor Times Matrix on Multi-core and Many-Core Architectures. | Jiajia Li, Yuchen Ma, Chenggang Yan, Richard W. Vuduc |
| 2016 | Designing MPI library with on-demand paging (ODP) of infiniband: challenges and benefits. | Mingzhe Li, Khaled Hamidouche, Xiaoyi Lu, Hari Subramoni, Jie Zhang, Dhabaleswar K. Panda |
| 2016 | Optimizing PLASMA Eigensolver on Large Shared Memory Systems. | Cheng Liao |
| 2016 | Improving application resilience to memory errors with lightweight compression. | Scott Levy, Kurt B. Ferreira, Patrick G. Bridges |
| 2016 | Characterizing parallel scientific applications on commodity clusters: an empirical study of a tapered fat-tree. | Edgar A. Len, Ian Karlin, Abhinav Bhatele, Steven H. Langer, Chris Chambreau, Louis H. Howell, Trent D'Hooge, Matthew L. Leininger |
| 2016 | Enhancing infiniband with openflow-style SDN capability. | Jason Lee, Zhou Tong, Karthik Achalkar, Xin Yuan, Michael Lang |
| 2016 | Extended task queuing: active messages for heterogeneous systems. | Michael LeBeane, Brandon Potter, Abhisek Pan, Alexandru Dutu, Vinay Agarwala, Wonchan Lee, Deepak Majeti, Bibek Ghimire, Eric Van Tassell, Samuel Wasmundt, Brad Benton, Maurcio Breternitz, Michael L. Chu, Mithuna Thottethodi, Lizy K. John, Steven K. Reinhardt |
| 2016 | Runtime Power Limiting of Parallel Applications on Intel Xeon Phi Processors. | Gary Lawson, Vaibhav Sundriyal, Masha Sosonkina, Yuzhong Shen |
| 2016 | An OpenCL Framework for Distributed Apps on a Multidimensional Network of FPGAs. | Abhijeet Lawande, Alan D. George, Herman Lam |
| 2016 | High-Performance Python-C++ Bindings with PyPy and Cling. | Wim T. L. P. Lavrijsen, Aditi Dutta |
| 2016 | Get out of the Way! Applying Compression to Internal Data Structures. | Robert Latham, Matthieu Dorier, Robert B. Ross |
| 2016 | OpenACC Cache Directive: Opportunities and Optimizations. | Ahmad Lashgar, Amirali Baniasadi |
| 2016 | Performance modeling of in situ rendering. | Matthew Larsen, Cyrus Harrison, James Kress, David Pugmire, Jeremy S. Meredith, Hank Childs |
| 2016 | Extending a Message Passing Runtime to Support Partitioned, Global Logical Address Spaces. | D. Brian Larkins, James Dinan |
| 2016 | A HYDRA UQ Workflow for NIF Ignition Experiments. | Steven H. Langer, Brian K. Spears, J. Luc Peterson, John Everett Field, Ryan Nora, Scott Brandon |
| 2016 | Devito: Towards a Generic Finite Difference DSL Using Symbolic Python. | Michael Lange, Navjot Kukreja, Mathias Louboutin, Fabio Luporini, Felippe Vieira, Vincenzo Pandolfo, Paulius Velesko, Paulius Kazakas, Gerard Gorman |
| 2016 | Floating-Point Shadow Value Analysis. | Michael O. Lam, Barry L. Rountree |
| 2016 | Pinpointing scale-dependent integer overflow bugs in large-scale parallel applications. | Ignacio Laguna, Martin Schulz |
| 2016 | G-store: high-performance graph store for trillion-edge processing. | Pradeep Kumar, H. Howie Huang |
| 2016 | Devito: Automated Fast Finite Difference Computation. | Navjot Kukreja, Mathias Louboutin, Felippe Vieira, Fabio Luporini, Michael Lange, Gerard Gorman |
| 2016 | Visualization and Analysis Requirements for In Situ Processing for a Large-Scale Fusion Simulation Code. | James Kress, David Pugmire, Scott Klasky, Hank Childs |
| 2016 | Towards Energy Efficient Data Management in HPC: The Open Ethernet Drive Approach. | Anthony Kougkas, Anthony Fleck, Xian-He Sun |
| 2016 | PIPES: a language and compiler for task-based programming on distributed-memory clusters. | Martin Kong, Louis-Nol Pouchet, P. Sadayappan, Vivek Sarkar |