| 2023 | High-Performance and Scalable Agent-Based Simulation with BioDynaMo. | Lukas Breitwieser, Ahmad Hesam, Fons Rademakers, Juan Gmez-Luna, Onur Mutlu |
| 2023 | Visibility Algorithms for Dynamic Dependence Analysis and Distributed Coherence. | Michael Bauer, Elliott Slaughter, Sean Treichler, Wonchan Lee, Michael Garland, Alex Aiken |
| 2023 | TL4x: Buffered Durable Transactions on Disk as Fast as in Memory. | Gal Assa, Andreia Correia, Pedro Ramalhete, Valerio Schiavoni, Pascal Felber |
| 2023 | Unexpected Scaling in Path Copying Trees. | Vitaly Aksenov, Trevor Brown, Alexander Fedorov, Ilya Kokorin |
| 2023 | Harnessing Extreme Heterogeneity for Ocean Modeling with Tensors. | Li Tang, Philip Jones, Scott Pakin |
| 2023 | Transactional Composition of Nonblocking Data Structures. | Wentao Cai, Haosen Wen, Michael L. Scott |
| 2022 | Vapro: performance variance detection and diagnosis for production-run parallel applications. | Liyan Zheng, Jidong Zhai, Xiongchao Tang, Haojie Wang, Teng Yu, Yuyang Jin, Shuaiwen Leon Song, Wenguang Chen |
| 2022 | High performance GPU concurrent B+tree. | Weihua Zhang, Chuanlei Zhao, Lu Peng, Yuzhe Lin, Fengzhe Zhang, Jinhu Jiang |
| 2022 | An LLVM-based open-source compiler for NVIDIA GPUs. | Da Yan, Wei Wang, Xiaowen Chu |
| 2022 | LB-HM: load balance-aware data placement on heterogeneous memory for task-parallel HPC applications. | Zhen Xie, Jie Liu, Sam Ma, Jiajia Li, Dong Li |
| 2022 | A W-cycle algorithm for efficient batched SVD on GPUs. | Junmin Xiao, Qing Xue, Hui Ma, Xiaoyang Zhang, Guangming Tan |
| 2022 | Parallel block-delayed sequences. | Sam Westrick, Mike Rainey, Daniel Anderson, Guy E. Blelloch |
| 2022 | FliT: a library for simple and efficient persistent algorithms. | Yuanhao Wei, Naama Ben-David, Michal Friedman, Guy E. Blelloch, Erez Petrank |
| 2022 | ParGeo: a library for parallel computational geometry. | Yiqiu Wang, Shangdi Yu, Laxman Dhulipala, Yan Gu, Julian Shun |
| 2022 | QGTC: accelerating quantized graph neural networks via GPU tensor core. | Yuke Wang, Boyuan Feng, Yufei Ding |
| 2022 | Beyond worst-case analysis: observed low depth for a P-complete problem. | Uzi Vishkin |
| 2022 | Understanding wafer-scale GPU performance using an architectural simulator. | Chris Thames, Hang Yan, Yifan Sun |
| 2022 | Accelerating data transfer between host and device using idle GPU. | Yuya Tatsugi, Akira Nukada |
| 2022 | Elimination (a, b)-trees with fast, durable updates. | Anubhav Srivastava, Trevor Brown |
| 2022 | Rethinking graph data placement for graph neural network training on multiple GPUs. | Shihui Song, Peng Jiang |
| 2022 | Systematically extending a high-level code generator with support for tensor cores. | Lukas Siefke, Bastian Kpcke, Sergei Gorlatch, Michel Steuwer |
| 2022 | Mashup: making serverless computing useful for HPC workflows via hybrid execution. | Rohan Basu Roy, Tirthak Patel, Vijay Gadepally, Devesh Tiwari |
| 2022 | Understanding and detecting deep memory persistency bugs in NVM programs with DeepMC. | Benjamin Reidys, Jian Huang |
| 2022 | Multi-queues can be state-of-the-art priority schedulers. | Anastasiia Postnikova, Nikita Koval, Giorgi Nadiradze, Dan Alistarh |
| 2022 | Compiler-assisted scheduling for multi-instance GPUs. | Chris Porter, Chao Chen, Santosh Pande |