| 2026 | ICS | SHIRO: Near-Optimal Communication Strategies for Distributed Sparse Matrix Multiplication. | Chen Zhuang, Lingqi Zhang, Benjamin Brock, Du Wu, Peng Chen, Toshio Endo, Satoshi Matsuoka, Mohamed Wahib |
| 2025 | SC | Reproducibility Report for SC25 Paper MLP-Offload: Multi-Level, Multi-Path Offloading for LLM Pre-training to Break the GPU Memory Wall. | Benjamin Brock |
| 2025 | SC | Slicing Is All You Need: Towards A Universal One-Sided Algorithm for Distributed Matrix Multiplication. | Benjamin Brock, Renato Golin |
| 2024 | ICS | RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs. | Benjamin Brock, Aydin Bulu, Katherine A. Yelick |
| 2024 | ICS | Distributed Ranges: A Model for Distributed Data Structures, Algorithms, and Views. | Benjamin Brock, Robert Cohn, Suyash Bakshi, Tuomas Karna, Jeongnim Kim, Mateusz Nowak, Lukasz Slusarczyk, Kacper Stefanski, Timothy G. Mattson |
| 2022 | ICPP | Atos: A Task-Parallel GPU Scheduler for Graph Analytics. | Yuxin Chen, Benjamin Brock, Serban D. Porumbescu, Aydin Bulu, Katherine A. Yelick, John D. Owens |
| 2022 | SC | Scalable Irregular Parallelism with GPUs: Getting CPUs Out of the Way. | Yuxin Chen, Benjamin Brock, Serban D. Porumbescu, Aydin Bulu, Katherine A. Yelick, John D. Owens |
| 2021 | ICS | Distributed-memory parallel algorithms for sparse times tall-skinny-dense matrix multiplication. | Oguz Selvitopi, Benjamin Brock, Israt Nisa, Alok Tripathy, Katherine A. Yelick, Aydin Bulu |
| 2019 | ICCAD | Centrifuge: Evaluating full-system HLS-generated heterogenous-accelerator SoCs using FPGA-Acceleration. | Qijing Huang, Christopher Yarp, Sagar Karandikar, Nathan Pemberton, Benjamin Brock, Liang Ma, Guohao Dai, Robert Quitt, Krste Asanovic, John Wawrzynek |
| 2019 | ICPP | BCL: A Cross-Platform Distributed Data Structures Library. | Benjamin Brock, Aydin Bulu, Katherine A. Yelick |