| 2021 | Large-Message Nonblocking MPI_Iallgather and MPI Ibcast Offload via BlueField-2 DPU. | Nick Sarkauskas, Mohammadreza Bayatpour, Tu Tran, Bharath Ramesh, Hari Subramoni, Dhabaleswar K. Panda |
| 2021 | RSP-Hist: Approximate Histograms for Big Data Exploration on Hadoop Clusters. | Salman Salloum, Joshua Zhexue Huang |
| 2021 | A Model of Graph Transactional Coverage Patterns with Applications to Drug Discovery. | A. Srinivas Reddy, P. Krishna Reddy, Anirban Mondal, U. Deva Priyakumar |
| 2021 | PILOT: a Runtime System to Manage Multi-tenant GPU Unified Memory Footprint. | John Ravi, Tri Nguyen, Huiyang Zhou, Michela Becchi |
| 2021 | A computational technique for parallel solution of diagonally dominant banded linear systems. | S. Chandra Sekhara Rao, Rabia Kamra |
| 2021 | SYMBIOMON: A High-Performance, Composable Monitoring Service. | Srinivasan Ramesh, Robert B. Ross, Matthieu Dorier, Allen D. Malony, Philip H. Carns, Kevin A. Huck |
| 2021 | Predictive Analysis of Large-Scale Coupled CFD Simulations with the CPX Mini-App. | Archie Powell, K. Choudry, Arun Prabhakar, Istvn Z. Reguly, Dario Amirante, Stephen A. Jarvis, Gihan R. Mudalige |
| 2021 | CUDA-DClust+: Revisiting Early GPU-Accelerated DBSCAN Clustering Designs. | Madhav Poudel, Michael Gowanlock |
| 2021 | Optimizing k-path selection for randomized interconnection networks. | Md Nahid Newaz, Md Atiqul Mollah |
| 2021 | How to Avoid Zero-Spacing in Fractionally-Strided Convolution? A Hardware-Algorithm Co-Design Methodology. | Yuan Meng, Sanmukh R. Kuppannagari, Rajgopal Kannan, Viktor K. Prasanna |
| 2021 | JACC: An OpenACC Runtime Framework with Kernel-Level and Multi-GPU Parallelization. | Kazuaki Matsumura, Simon Garcia de Gonzalo, Antonio J. Pea |
| 2021 | A Programming API Implementation for Secure Data Analytics Applications with Homomorphic Encryption on GPUs. | Shuangsheng Lou, Gagan Agrawal |
| 2021 | Dynamic Voltage and Frequency Scaling to Improve Energy-Efficiency of Hardware Accelerators. | Siqin Liu, Avinash Karanth |
| 2021 | Parallel Algorithms for Efficient Computation of High-Order Line Graphs of Hypergraphs. | Xu T. Liu, Jesun Firoz, Andrew Lumsdaine, Cliff A. Joslyn, Sinan G. Aksoy, Brenda Praggastis, Assefaw H. Gebremedhin |
| 2021 | Optimizing Multi-Range based Error-Bounded Lossy Compression for Scientific Datasets. | Yuanjian Liu, Sheng Di, Kai Zhao, Sian Jin, Cheng Wang, Kyle Chard, Dingwen Tao, Ian T. Foster, Franck Cappello |
| 2021 | Shrinking Sample Search Algorithm for Automatic Tuning of GPU Kernels. | Xiang Li, Gagan Agrawal |
| 2021 | Asynchronous I/O Strategy for Large-Scale Deep Learning Applications. | Sunwoo Lee, Qiao Kang, Kewei Wang, Jan Balewski, Alex Sim, Ankit Agrawal, Alok N. Choudhary, Peter Nugent, Kesheng Wu, Wei-keng Liao |
| 2021 | Shared-memory implementation of the Karp-Sipser kernelization process. | Johannes Langguth, Ioannis Panagiotas, Bora Uar |
| 2021 | HiPC 2021 Workshop on Parallel Programming in the Exascale Era (PPEE 2021). | Vivek Kumar, Swarnendu Biswas, Vishwesh Jatala |
| 2021 | Teaching High Productivity and High Performance in an Introductory Parallel Programming Course. | Vivek Kumar |
| 2021 | Multi-Stage Memory Efficient Strassen's Matrix Multiplication on GPU. | Arjun Gopala Krishnan, Dhrubajyoti Goswami |
| 2021 | Deciding Non-Compressible Blocks in Sparse Direct Solvers using Incomplete Factorization. | Esragul Korkmaz, Mathieu Faverge, Pierre Ramet, Grgoire Pichon |
| 2021 | MulConn: User-Transparent I/O Subsystem for High-Performance Parallel File Systems. | Hwajung Kim, Jiwoo Bang, Dong Kyu Sung, Hyeonsang Eom, Heon Y. Yeom, Hanul Sung |
| 2021 | Towards Zero-Waste Recovery and Zero-Overhead Checkpointing in Ensemble Data Assimilation. | Kai Keller, Adrin Cristal Kestelman, Leonardo Bautista-Gomez |
| 2021 | HPC@SCALE: A Hands-on Approach for Training Next-Gen HPC Software Architects. | Tanzima Z. Islam, Chase Phelps |