| 2021 | Meeting the real-time challenges of ground-based telescopes using low-rank matrix computations. | Hatem Ltaief, Jesse Cranney, Damien Gratadour, Yuxi Hong, Laurent Gatineau, David E. Keyes |
| 2021 | Reducing redundancy in data organization and arithmetic calculation for stencil computations. | Kun Li, Liang Yuan, Yunquan Zhang, Yue Yue |
| 2021 | Closing the "quantum supremacy" gap: achieving real-time simulation of a random quantum circuit using a new Sunway supercomputer. | Yong (Alexander) Liu, Xin (Lucy) Liu, Fang (Nancy) Li, Haohuan Fu, Yuling Yang, Jiawei Song, Pengpeng Zhao, Zhen Wang, Dajia Peng, Huarong Chen, Chu Guo, Heliang Huang, Wenzhao Wu, Dexun Chen |
| 2021 | HatRPC: hint-accelerated thrift RPC over RDMA. | Tianxi Li, Haiyang Shi, Xiaoyi Lu |
| 2021 | RIBBON: cost-effective and qos-aware deep learning model inference using a diverse pool of cloud computing instances. | Baolin Li, Rohan Basu Roy, Tirthak Patel, Vijay Gadepally, Karen Gettings, Devesh Tiwari |
| 2021 | STM-multifrontal QR: streaming task mapping multifrontal QR factorization empowered by GCN. | Shengle Lin, Wangdong Yang, Haotian Wang, Qinyun Tsai, Kenli Li |
| 2021 | Understanding Effectiveness of Multi-Error-Bounded Lossy Compression for Preserving Ranges of Interest in Scientific Analysis. | Yuanjian Lin, Sheng Di, Kai Zhao, Sian Jin, Cheng Wang, Kyle Chard, Dingwen Tao, Ian T. Foster, Franck Cappello |
| 2021 | SW_Qsim: a minimize-memory quantum simulator with high-performance on a new Sunway supercomputer. | Fang Li, Xin Liu, Yong Liu, Pengpeng Zhao, Yuling Yang, Honghui Shang, Weizhe Sun, Zhen Wang, Enming Dong, Dexun Chen |
| 2021 | SV-sim: scalable PGAS-based state vector simulation of quantum circuits. | Ang Li, Bo Fang, Christopher E. Granade, Guen Prawiroatmodjo, Bettina Heim, Martin Roetteler, Sriram Krishnamoorthy |
| 2021 | Resilient error-bounded lossy compressor for data transfer. | Sihuan Li, Sheng Di, Kai Zhao, Xin Liang, Zizhong Chen, Franck Cappello |
| 2021 | Automated In Situ Computational Steering Using Ascent's Capable Yes-No Machine. | Margaret Lawson, Cyrus Harrison, Eric Brugger, Aaron Skinner, Matthew Larsen |
| 2021 | No Coherence? No Problem! Virtual Shared Memory for MPSoCs. | Tobias Langer, Jonas Rabenstein, Timo Hnig, Wolfgang Schrder-Preikschat |
| 2021 | Low Overhead Security Isolation using Lightweight Kernels and TEEs. | John R. Lange, Nicholas Gordon, Brian Gaines |
| 2021 | On the parallel I/O optimality of linear algebra kernels: near-optimal matrix factorizations. | Grzegorz Kwasniewski, Marko Kabic, Tal Ben-Nun, Alexandros Nikolaos Ziogas, Jens Eirik Saethre, Andr Gaillard, Timo Schneider, Maciej Besta, Anton Kozhevnikov, Joost VandeVondele, Torsten Hoefler |
| 2021 | CAKE: matrix multiplication using constant-bandwidth blocks. | H. T. Kung, Vikas Natesh, Andrew Sabot |
| 2021 | Cuttlefish: library for achieving energy efficiency in multicore parallel programs. | Sunil Kumar, Akshat Gupta, Vivek Kumar, Sridutt Bhalachandra |
| 2021 | 3D acoustic-elastic coupling with gravity: the dynamics of the 2018 palu, sulawesi earthquake and tsunami. | Lukas Krenz, Carsten Uphoff, Thomas Ulrich, Alice-Agnes Gabriel, Lauren S. Abrahams, Eric M. Dunham, Michael Bader |
| 2021 | Exploring Lossy Compressibility through Statistical Correlations of Scientific Datasets. | David Krasowska, Julie Bessac, Robert Underwood, Jon C. Calhoun, Sheng Di, Franck Cappello |
| 2021 | Arithmetic-intensity-guided fault tolerance for neural network inference on GPUs. | Jack Kosaian, K. V. Rashmi |
| 2021 | ndzip-gpu: efficient lossless compression of scientific floating-point data on GPUs. | Fabian Knorr, Peter Thoman, Thomas Fahringer |
| 2021 | OSCAR Parallelizing and Power Reducing Compiler and API for Heterogeneous Multicores : (Invited Paper). | Hironori Kasahara, Keiji Kimura, Toshiaki Kitamura, Hiroki Mikami, Kazutaka Morita, Kazuki Fujita, Kazuki Yamamoto, Tohma Kawasumi |
| 2021 | Optimization of Asynchronous Communication Operations through Eager Notifications. | Amir Kamil, Dan Bonachea |
| 2021 | A Python-based High-Level Programming Flow for CPU-FPGA Heterogeneous Systems : (Invited Paper). | Sitao Huang, Kun Wu, Sai Rahul Chalamalasetti, Izzat El Hajj, Cong Xu, Paolo Faraboschi, Deming Chen |
| 2021 | TFProf: Profiling Large Taskflow Programs with Modern D3 and C++. | Tsung-Wei Huang |
| 2021 | Characterization and prediction of deep learning workloads in large-scale GPU datacenters. | Qinghao Hu, Peng Sun, Shengen Yan, Yonggang Wen, Tianwei Zhang |