| 2019 | HYPHA: a framework based on separation of parallelisms to accelerate persistent homology matrix reduction. | Simon Zhang, Mengbai Xiao, Chengxin Guo, Liang Geng, Hao Wang, Xiaodong Zhang |
| 2019 | Laius: Towards latency awareness and improved utilization of spatial multitasking accelerators in datacenters. | Wei Zhang, Weihao Cui, Kaihua Fu, Quan Chen, Daniel Edward Mawhirter, Bo Wu, Chao Li, Minyi Guo |
| 2019 | GreenMM: energy efficient GPU matrix multiplication through undervolting. | Hadi Zamani, Yuanlai Liu, Devashree Tripathy, Laxmi N. Bhuyan, Zizhong Chen |
| 2019 | Can we trust profiling results?: understanding and fixing the inaccuracy in modern profilers. | Hao Xu, Qingsen Wang, Shuang Song, Lizy Kurian John, Xu Liu |
| 2019 | GPUGuard: mitigating contention based side and covert channel attacks on GPUs. | Qiumin Xu, Hoda Naghibijouybari, Shibo Wang, Nael B. Abu-Ghazaleh, Murali Annavaram |
| 2019 | IA-SpGEMM: an input-aware auto-tuning framework for parallel sparse matrix-matrix multiplication. | Zhen Xie, Guangming Tan, Weifeng Liu, Ninghui Sun |
| 2019 | Parallelizing cryo-EM 3D reconstruction on GPU cluster with a partitioned and streamed model. | Kunpeng Wang, Shizhen Xu, Haohuan Fu, Hongkun Yu, Wenlai Zhao, Guangwen Yang |
| 2019 | Address-stride assisted approximate load value prediction in GPUs. | Haonan Wang, Mohamed Assem Ibrahim, Sparsh Mittal, Adwait Jog |
| 2019 | Multi-criteria partitioning of multi-block structured grids. | Hengjie Wang, Aparna Chandramowlishwaran |
| 2019 | WCCV: improving the vectorization of IF-statements with warp-coherent conditions. | Huihui Sun, Florian Fey, Jie Zhao, Sergei Gorlatch |
| 2019 | A communication-avoiding 3D sparse triangular solver. | Piyush Sao, Ramakrishnan Kannan, Xiaoye Sherry Li, Richard W. Vuduc |
| 2019 | Efficient thread/page/parallelism autotuning for NUMA systems. | Mihail Popov, Alexandra Jimborean, David Black-Schaffer |
| 2019 | Efficient hierarchical online-autotuning: a case study on polyhedral accelerator mapping. | Philip Pfaffe, Tobias Grosser, Martin Peter Tillmann |
| 2019 | Performance optimization of reactive molecular dynamics simulations with dynamic charge distribution models on distributed memory platforms. | Kurt A. O'Hearn, Abdullah Alperen, Hasan Metin Aktulga |
| 2019 | Deep reuse: streamline CNN inference on the fly via coarse-grained computation reuse. | Lin Ning, Xipeng Shen |
| 2019 | SDC: a software defined cache for efficient data indexing. | Fan Ni, Song Jiang, Hong Jiang, Jian Huang, Xingbo Wu |
| 2019 | Full-stack optimization for accelerating CNNs using powers-of-two weights with FPGA validation. | Bradley McDanel, Sai Qian Zhang, H. T. Kung, Xin Dong |
| 2019 | Efficient GPU tree walks for effective distributed n-body simulations. | Jianqiao Liu, Michael P. Robson, Thomas Quinn, Milind Kulkarni |
| 2019 | Efficient and effective sparse tensor reordering. | Jiajia Li, Bora Uar, mit V. atalyrek, Jimeng Sun, Kevin J. Barker, Richard W. Vuduc |
| 2019 | DeepHiR: improving high-radix router throughput with deep hybrid memory buffer microarchitecture. | Cunlu Li, Dezun Dong, Xiangke Liao, John Kim, Changhyun Kim |
| 2019 | GPU snapshot: checkpoint offloading for GPU-dense systems. | Kyushick Lee, Michael B. Sullivan, Siva Kumar Sastry Hari, Timothy Tsai, Stephen W. Keckler, Mattan Erez |
| 2019 | Least squares solvers for distributed-memory machines with GPU accelerators. | Jakub Kurzak, Mark Gates, Ali Charara, Asim YarKhan, Jack J. Dongarra |
| 2019 | AMPT-GA: automatic mixed precision floating point tuning for GPU applications. | Pradeep V. Kotipalli, Ranvijay Singh, Paul Wood, Ignacio Laguna, Saurabh Bagchi |
| 2019 | GPU road network graph contraction and SSSP query. | Roozbeh Karimi, David M. Koppelman, Chris J. Michael |
| 2019 | Henosis: workload-driven small array consolidation and placement for HDF5 applications on heterogeneous data stores. | Donghe Kang, Vedang Patel, Ashwati Nair, Spyros Blanas, Yang Wang, Srinivasan Parthasarathy |