| 2016 | Efficient Processing of Large Graphs via Input Reduction. | Amlan Kusum, Keval Vora, Rajiv Gupta, Iulian Neamtiu |
| 2016 | Network-Managed Virtual Global Address Space for Message-driven Runtimes. | Abhishek Kulkarni, Luke Dalessandro, Ezra Kissel, Andrew Lumsdaine, Thomas L. Sterling, D. Martin Swany |
| 2016 | IMPACC: A Tightly Integrated MPI+OpenACC Framework Exploiting Shared Memory Parallelism. | Jungwon Kim, Seyong Lee, Jeffrey S. Vetter |
| 2016 | LUT Optimization In Implementation Of Combinational Karatsuba Ofman On Virtex-6 FPGA. | Deepak Kapoor, Rahul Yamasani, Saket Saurav, Abhishek Bajpai |
| 2016 | Distributed Incremental Pattern Matching on Streaming Graphs. | Jyun-Sheng Kao, Jerry Chou |
| 2016 | Self-configuring Software-defined Overlay Bypass for Seamless Inter- and Intra-cloud Virtual Networking. | Kyuho Jeong, Renato J. O. Figueiredo |
| 2016 | Improving GPU Performance Through Resource Sharing. | Vishwesh Jatala, Jayvant Anantpur, Amey Karkare |
| 2016 | Challenges in Transition. | Kazuaki Ishizaki |
| 2016 | Evaluation of Pattern Matching Workloads in Graph Analysis Systems. | Seokyong Hong, Sangkeun Lee, Seung-Hwan Lim, Sreenivas R. Sukumar, Ranga Raju Vatsavai |
| 2016 | Scalable I/O-Aware Job Scheduling for Burst Buffer Enabled HPC Clusters. | Stephen Herbein, Dong H. Ahn, Don Lipari, Thomas R. W. Scogland, Marc Stearman, Mark Grondona, Jim Garlick, Becky Springmeyer, Michela Taufer |
| 2016 | A Scalable Runtime for the ECOSCALE Heterogeneous Exascale Hardware Platform. | Paul Harvey, Konstantin Bakanov, Ivor T. A. Spence, Dimitrios S. Nikolopoulos |
| 2016 | Automatic Hybridization of Runtime Systems. | Kyle C. Hale, Conor Hetland, Peter A. Dinda |
| 2016 | SWAT: A Programmable, In-Memory, Distributed, High-Performance Computing Platform. | Max Grossman, Vivek Sarkar |
| 2016 | A Multi-Kernel Survey for High-Performance Computing. | Balazs Gerofi, Yutaka Ishikawa, Rolf Riesen, Robert W. Wisniewski, Yoonho Park, Bryan S. Rosenburg |
| 2016 | The READEX Project for Dynamic Energy Efficiency Tuning. | Michael Gerndt |
| 2016 | Implications of Heterogeneous Memories in Next Generation Server Systems. | Ada Gavrilovska |
| 2016 | A Cross-Enclave Composition Mechanism for Exascale System Software. | Noah Evans, Kevin T. Pedretti, Brian Kocoloski, John R. Lange, Michael Lang, Patrick G. Bridges |
| 2016 | SDS-Sort: Scalable Dynamic Skew-aware Parallel Sorting. | Bin Dong, Surendra Byna, Kesheng Wu |
| 2016 | With Extreme Scale Computing the Rules Have Changed. | Jack J. Dongarra |
| 2016 | Routing on the Dependency Graph: A New Approach to Deadlock-Free High-Performance Routing. | Jens Domke, Torsten Hoefler, Satoshi Matsuoka |
| 2016 | NVL-C: Static Analysis Techniques for Efficient, Correct Programming of Non-Volatile Main Memory Systems. | Joel E. Denny, Seyong Lee, Jeffrey S. Vetter |
| 2016 | GPUrdma: GPU-side library for high performance networking from GPU kernels. | Feras Daoud, Amir Wated, Mark Silberstein |
| 2016 | Betweenness Centrality in an HSA-enabled System. | Shuai Che, Marc S. Orr, Gregory Rodgers, Jonathan Gallmeier |
| 2016 | DD-Graph: A Highly Cost-Effective Distributed Disk-based Graph-Processing Framework. | Yongli Cheng, Fang Wang, Hong Jiang, Yu Hua, Dan Feng, XiuNeng Wang |
| 2016 | Scaling Spark on HPC Systems. | Nicholas Chaimov, Allen D. Malony, Shane Canon, Costin Iancu, Khaled Z. Ibrahim, Jay Srinivasan |