| 2025 | GraCFL: A Holistically Designed Vertex-Centric Graph System for CFL Reachability. | Sakib Fuad, Amir Hossein Nodehi Sabet, Umar Farooq, Zhijia Zhao |
| 2025 | Generating Microservice Graphs with Production Characteristics for Efficient Resource Scaling. | Fanrong Du, Jiuchen Shi, Quan Chen, Pu Pang, Li Li, Minyi Guo |
| 2025 | Loop Fusion in Matrix Multiplications with Sparse Dependence. | Mohammad Mahdi Salehi Dezfuli, Kazem Cheshmi |
| 2025 | IA-Chol: Input-Aware Cholesky Decomposition on CPU and GPU. | Jixiao Deng, Qinglin Wang, Lin Chen, Tun Li, Bo Yang, Xinhai Chen, Jie Liu |
| 2025 | Taking GPU Programming Models to Task for Performance Portability. | Joshua Hoke Davis, Pranav Sivaraman, Joy Kitson, Konstantinos Parasyris, Harshitha Menon, Isaac Minn, Giorgis Georgakoudis, Abhinav Bhatele |
| 2025 | HARNESS: Holistic Resource Management for Diversely Scaled Edge Cloud Systems. | Ismet Dagli, Justin Davis, Mehmet Esat Belviranli |
| 2025 | DeCOS: Data-Efficient Reinforcement Learning for Compiler Optimization Selection Ignited by LLM. | Tianming Cui, Pen-Chung Yew, Stephen McCamant, Antonia Zhai |
| 2025 | CB-SpMV: A Data Aggregating and Balance Algorithm for for Cache-Friendly Block-Based SpMV on GPUs. | Xing Cong, FuKai Sun, YiFan Chen, Chenhao Xie, Yi Liu, Depei Qian |
| 2025 | SYprox: Combining Host and Device Perforation with Mixed Precision Approximation on Heterogeneous Architectures. | Lorenzo Carpentieri, Biagio Cosenza |
| 2025 | PortFC: Designing High-performance Deadlock-free BCube Networks. | Peirui Cao, Rui Ning, Hongwei Yang, Zhaochen Zhang, Chang Liu, Rui Li, Yongqi Yang, Yunzhuo Liu, Chengyuan Huang, Tao Sun, Xiaodong Duan, Guihai Chen, Chen Tian |
| 2025 | UJOpt: Heuristic Approach for Applying Unroll-and-Jam Optimization and Loop Order Selection. | Shilpa Babalad, Shirish K. Shevade, Matthew Jacob Thazhuthaveetil, R. Govindarajan |
| 2025 | A Multi-GPU Algorithm for Computing Maximal Independent Sets in Large Graphs. | Anju Mongandampulath Akathoott, Benila Virgin Jerald Xavier, Martin Burtscher |
| 2025 | Efficient Server Consolidation through a balanced mix of Transformer-based and Conventional Applications. | Pablo Abad, Pablo Prieto, Valentin Puente, Jos-ngel Gregorio |
| 2025 | StructILU: Dependency-Preserving Incomplete LU with Hierarchical Parallelism for Structured Grid PDEs on GPUs. | Hao Luo, Qianchao Zhu, Xiaochen Hao, Chunxi Lei, Chengdi Ma, Chenchen Zhang, Yun Liang, Chao Yang |
| 2025 | NeurLZ: An Online Neural Learning-based Method to Enhance Scientific Lossy Compression. | Wenqi Jia, Zhewen Hu, Youyuan Liu, Boyuan Zhang, Jinzhen Wang, Jinyang Liu, Wei Niu, Stavros Kalafatis, Junzhou Huang, Sian Jin, Daoce Wang, Jiannan Tian, Miao Yin |
| 2025 | Pushing the Limits of GPU Lossy Compression: A Hierarchical Delta Approach. | Boyuan Zhang, Yafan Huang, Sheng Di, Fengguang Song, Guanpeng Li, Franck Cappello |
| 2025 | BMQSim: Overcoming Memory Constraints in Quantum Circuit Simulation with a High-Fidelity Compression Framework. | Boyuan Zhang, Bo Fang, Fanjiang Ye, Luanzheng Guo, Fengguang Song, Nathan R. Tallent, Dingwen Tao |
| 2025 | ghZCCL: Advancing GPU-aware Collective Communications with Homomorphic Compression. | Jiajun Huang, Sheng Di, Yafan Huang, Zizhong Chen, Franck Cappello, Yanfei Guo, Rajeev Thakur |
| 2024 | Stencil Computation with Vector Outer Product. | Wenxuan Zhao, Liang Yuan, Baicheng Yan, Penghao Ma, Yunquan Zhang, Long Wang, Zhe Wang |
| 2024 | CLAY: CXL-based Scalable NDP Architecture Accelerating Embedding Layers. | Sungmin Yun, Hwayong Nam, Kwanhee Kyung, Jaehyun Park, Byeongho Kim, Yongsuk Kwon, Eojin Lee, Jung Ho Ahn |
| 2024 | Real-time High-resolution X-Ray Computed Tomography. | Du Wu, Peng Chen, Xiao Wang, Isaac Lyngaas, Takaaki Miyajima, Toshio Endo, Satoshi Matsuoka, Mohamed Wahib |
| 2024 | sys-sage: A Unified Representation of Dynamic Topologies & Attributes on HPC Systems. | Stepan Vanecek, Martin Schulz |
| 2024 | Differentiating Set Intersections in Maximal Clique Enumeration by Function and Subproblem Size. | Hans Vandierendonck |
| 2024 | SLIDEX: A Novel Architecture for Sliding Window Processing. | Ral Taranco, Jos-Mara Arnau, Antonio Gonzlez |
| 2024 | DeepHYDRA: A Hybrid Deep Learning and DBSCAN-Based Approach to Time-Series Anomaly Detection in Dynamically-Configured Systems. | Franz Kevin Stehle, Wainer Vandelli, Felix Zahn, Giuseppe Avolio, Holger Frning |