| 2017 | Exploiting Common Neighborhoods to Optimize MPI Neighborhood Collectives. | Seyed Hessam Mirsadeghi, Jesper Larsson Trff, Pavan Balaji, Ahmad Afsahi |
| 2017 | Last Level Collective Hardware Prefetching For Data-Parallel Applications. | George Michelogiannakis, John Shalf |
| 2017 | Exact and Parallel Triangle Counting in Dynamic Graphs. | Devavret Makkar, David A. Bader, Oded Green |
| 2017 | Designing Registration Caching Free High-Performance MPI Library with Implicit On-Demand Paging (ODP) of InfiniBand. | Mingzhe Li, Xiaoyi Lu, Hari Subramoni, Dhabaleswar K. Panda |
| 2017 | Applying Graph Analytics to Understand Compute Core Usage and Publication Trends in a Petascale Supercomputing Facility. | Sangkeun Lee, Sudharshan S. Vazhkudai, Raghul Gunasekaran |
| 2017 | Parallel Deep Convolutional Neural Network Training by Exploiting the Overlapping of Computation and Communication. | Sunwoo Lee, Dipendra Jha, Ankit Agrawal, Alok N. Choudhary, Wei-keng Liao |
| 2017 | ReCALL: Reordered Cache Aware Locality Based Graph Processing. | Kartik Lakhotia, Shreyas G. Singapura, Rajgopal Kannan, Viktor K. Prasanna |
| 2017 | Characterization of Data Movement Requirements for Sparse Matrix Computations on GPUs. | Sreyya Emre Kurt, Vineeth Thumma, Changwan Hong, Aravind Sukumaran-Rajam, P. Sadayappan |
| 2017 | Reducing Network Congestion and Synchronization Overhead During Aggregation of Hierarchical Data. | Sidharth Kumar, Duong Hoang, Steve Petruzza, John Edwards, Valerio Pascucci |
| 2017 | Context-Aware Memory Profiling for Speculative Parallelism. | Changsu Kim, Juhyun Kim, Juwon Kang, Jae W. Lee, Hanjun Kim |
| 2017 | Numerical Simulation of Aerospace Applications Using Overset Mesh. | Alok Khaware, Vinay Kumar Gupta, Abhilash Rajan |
| 2017 | Distributed Algorithm for High-Utility Subgraph Pattern Mining Over Big Data Platforms. | Alind Khare, Vikram Goyal, Srikanth Baride, Sushil K. Prasad, Michael McDermott, Dhara Shah |
| 2017 | CFD Invited Speakers [3 abstracts]. | Amit P. Kesarkar, Dipak Maiti, Lourens Post |
| 2017 | Scalable Exact Parent Sets Identification in Bayesian Networks Learning with Apache Spark. | Subhadeep Karan, Jaroslaw Zola |
| 2017 | Parallel Asynchronous Distributed-Memory Maximal Independent Set Algorithm with Work Ordering. | Thejaka Amila Kanewala, Marcin Zalewski, Andrew Lumsdaine |
| 2017 | ARM Wrestling with Big Data: A Study of Commodity ARM64 Server for Big Data Workloads. | Jayanth Kalyanasundaram, Yogesh Simmhan |
| 2017 | Software Troubleshooting Using Machine Learning. | Neha Mukund Kalibhat, Shreya Varshini, Chid Kollengode, Dinkar Sitaram, Subramaniam Kalambur |
| 2017 | Shared-Memory Graph Truss Decomposition. | Humayun Kabir, Kamesh Madduri |
| 2017 | Integrating External Resources with a Task-Based Programming Model. | Zhihao Jia, Sean Treichler, Galen M. Shipman, Michael Bauer, Noah Watkins, Carlos Maltzahn, Patrick S. McCormick, Alex Aiken |
| 2017 | Efficient Fork-Join on GPUs Through Warp Specialization. | Arpith Chacko Jacob, Alexandre E. Eichenberger, Hyojin Sung, Samuel F. Anto, Gheorghe-Teodor Bercea, Carlo Bertolli, Alexey Bataev, Tian Jin, Tong Chen, Zehra Sura, Georgios Rokos, Kevin O'Brien |
| 2017 | Kernel-Assisted Communication Engine for MPI on Emerging Manycore Processors. | Jahanzeb Maqbool Hashmi, Khaled Hamidouche, Hari Subramoni, Dhabaleswar K. Panda |
| 2017 | MPI-LiFE: Designing High-Performance Linear Fascicle Evaluation of Brain Connectome with MPI. | Shashank Gugnani, Xiaoyi Lu, Franco Pestilli, Cesar F. Caiafa, Dhabaleswar K. Panda |
| 2017 | CFD Introduction. | Vinay R. Gopala, Aditya Singh, Sunil Appanaboyina |
| 2017 | Thrust++: Extending Thrust Framework for Better Abstraction and Performance. | Ajai V. George, Sankar Manoj, Sanket R. Gupte, Sayantan Mitra, Santonu Sarkar |
| 2017 | Computing Just What You Need: Online Data Analysis and Reduction at Extreme Scales. | Ian T. Foster |