| 2020 | No barrier in the road: a comprehensive study and optimization of ARM barriers. | Nian Liu, Binyu Zang, Haibo Chen |
| 2020 | MatRox: modular approach for improving data locality in hierarchical (Mat)rix App(Rox)imation. | Bangtian Liu, Kazem Cheshmi, Saeed Soori, Michelle Mills Strout, Maryam Mehri Dehnavi |
| 2020 | PLUM: static parallel program locality analysis under uniform multiplexing. | Fangzhou Liu, Dong Chen, Wesley Smith, Chen Ding |
| 2020 | Generating energy-efficient code for parallel applications specified by streaming task graphs with dynamic elements. | Sebastian Litzinger, Jrg Keller |
| 2020 | A parallel sparse tensor benchmark suite on CPUs and GPUs. | Jiajia Li, Mahesh Lakshminarasimhan, Xiaolong Wu, Ang Li, Catherine Olschanowsky, Kevin J. Barker |
| 2020 | Taming unbalanced training workloads in deep learning with partial collective operations. | Shigang Li, Tal Ben-Nun, Salvatore Di Girolamo, Dan Alistarh, Torsten Hoefler |
| 2020 | How to speed Connected Component Labeling up with SIMD RLE algorithms. | Florian Lemaitre, Arthur M. Hennequin, Lionel Lacassagne |
| 2020 | Lock-free transactional vector. | Kenneth Lamar, Christina L. Peterson, Damian Dechev |
| 2020 | Testing concurrency on the JVM with lincheck. | Nikita Koval, Maria Sokolova, Alexander Fedorov, Dan Alistarh, Dmitry Tsitelov |
| 2020 | Restricted memory-friendly lock-free bounded queues. | Nikita Koval, Vitaly Aksenov |
| 2020 | Overlapping host-to-device copy and computation using hidden unified memory. | Jaehoon Jung, Daeyoung Park, Youngdong Do, Jungho Park, Jaejin Lee |
| 2020 | Identifying scalability bottlenecks for large-scale parallel programs with graph analysis. | Yuyang Jin, Haojie Wang, Xiongchao Tang, Torsten Hoefler, Xu Liu, Jidong Zhai |
| 2020 | A novel data transformation and execution strategy for accelerating sparse matrix multiplication on GPUs. | Peng Jiang, Changwan Hong, Gagan Agrawal |
| 2020 | Understand the overheads of storage data structures on persistent memory. | Abdullah Al Raqibul Islam, Dong Dai |
| 2020 | Parallel and distributed bounded model checking of multi-threaded programs. | Omar Inverso, Catia Trubiani |
| 2020 | <u>G</u>PU <u>i</u>nitiated <u>O</u>penSHMEM: correct and efficient intra-kernel networking for dGPUs. | Khaled Hamidouche, Michael LeBeane |
| 2020 | Custom code generation for a graph DSL. | Bikash Gogoi, Unnikrishnan Cheramangalath, Rupesh Nasre |
| 2020 | The Minos Computing Library: efficient parallel programming for extremely heterogeneous systems. | Roberto Gioiosa, Burcu Ozcelik Mutlu, Seyong Lee, Jeffrey S. Vetter, Giulio Picierro, Marco Cesati |
| 2020 | Towards a portable hierarchical view of distributed shared memory systems: challenges and solutions. | Millad Ghane, Sunita Chandrasekaran, Margaret S. Cheung |
| 2020 | Kite: efficient and available release consistency for the datacenter. | Vasilis Gavrielatos, Antonios Katsarakis, Vijay Nagarajan, Boris Grot, Arpit Joshi |
| 2020 | Self-adjusting task granularity for Global load balancer library on clusters of many-core processors. | Patrick Finnerty, Tomio Kamada, Chikara Ohta |
| 2020 | Neighbor-list-free molecular dynamics on sunway TaihuLight supercomputer. | Xiaohui Duan, Ping Gao, Meng Zhang, Tingjian Zhang, Hongsong Meng, Yuxuan Li, Bertil Schmidt, Haohuan Fu, Lin Gan, Wei Xue, Guangwen Yang, Weiguo Liu |
| 2020 | Detecting and reproducing error-code propagation bugs in MPI implementations. | Daniel DeFreez, Antara Bhowmick, Ignacio Laguna, Cindy Rubio-Gonzlez |
| 2020 | A wait-free universal construction for large objects. | Andreia Correia, Pedro Ramalhete, Pascal Felber |
| 2020 | TardisTM: incremental repair for transactional memory. | Daming D. Chen, Phillip B. Gibbons, Todd C. Mowry |