| 2025 | EquilibrIO: Taming the I/O Tides in High-Performance Computing. | Taylan zden, Ahmad Tarraf, Felix Wolf |
| 2025 | Efficient Multi-GPU Programming in Python: Reducing Synchronization and Access Overheads. | Lena Oden, Klaus Nlp |
| 2025 | Incremental Sparse Tensor Format for Maximizing Efficiency in Tensor-Vector Multiplications. | Xiaohe Niu, Georg Meyer, Dimosthenis Pasadakis, Albert-Jan Yzelman, Olaf Schenk |
| 2025 | Toward LLM-Compatible Log Representation Learning: A Hierarchical Semantic-Structural Framework for HPC Anomaly Detection. | Dumo Ngwenya, Dhouha Kbaier, Patrick Wong |
| 2025 | Parallel Selected Inversion of Block-Tridiagonal with Arrowhead Matrices. | Vincent Maillou, Lisa Gaedke-Merzhuser, Alexandros Nikolaos Ziogas, Olaf Schenk, Mathieu Luisier |
| 2025 | SoCL: Scalable and Latency-Optimized Microservices in Serverless Edge Computing. | Shuaibing Lu, Bojin Xiang, Jie Wu, Ziyu You, Wentong Cai |
| 2025 | Optimizing I/O for an Exascale Implicit Kinetic Plasma Simulation using the Rabbit Storage System. | Ian Lumsden, Hariharan Devarajan, Izzet Yildirim, Stefano Markidis, Andong Hu, Ivy Peng, Luca Pennati, Dewi Yokelson, Stephanie Brink, Olga Pearce, Tom Scogland, Bronis R. de Supinski, Gian Luca Delzanno, Anthony Kougkas, Xian-He Sun, Michela Taufer |
| 2025 | SYCL QPU: an LLVM-based QPU simulation framework built using DPC++. | Miguel Leal, Francisco Javier Cardama, Toms F. Pena |
| 2025 | FastEM: an efficient EM algorithm for learning Gaussian mixture models on compute clusters. | Wojciech Kwedlo |
| 2025 | BBView: A View-Aware Burst-Buffer Mechanism for MPI-IO. | Sohei Koyama, Osamu Tatebe |
| 2025 | PRT: An Efficient Pipeline Reuse Technology for Large Models Training. | Zeyu Ji, Banghao Zhai, Zhonghao Zhang, Qi Chu, Bin Liu |
| 2025 | FIFO-MEP: An Efficient Multi-Eviction-Point FIFO Cache with Stable Demotion for Burst-Oriented Access Mitigation. | Ranhao Jia, Yunfei Gu, Chentao Wu, Jie Li, Minyi Guo, Liqiang Zhang, Zaigui Zhang, Haijun Zhang |
| 2025 | Lessons from Profiling and Optimizing Placement in AMR Codes. | Ankush Jain, Charles D. Cranor, Qing Zheng, Dominic Manno, George Amvrosiadis, Gary A. Grider |
| 2025 | Performance Evaluation of HPC Benchmarks on a RISC-V Based Cluster. | Rushikesh Jadhav, Surendra Billa, Yogeshwar Sonawane, Sanjay Wandhekar |
| 2025 | A similarity-aware MOE-based method for optimizing tensor programs across diverse GPUs. | Haowen Hou, Zihan Wang, Yining Song, Lei Gong, Chao Wang, Xi Li, Xuehai Zhou |
| 2025 | Accelerating Key-Value Data Structures Using AVX-512 SIMD Extensions. | MohammadReza HoseinyFarahabady, Javid Taheri, Albert Y. Zomaya |
| 2025 | PALLAS: HPC trace analysis at scale. | Catherine Guelque, Valentin Honor, Philippe Swartvagher, Franois Trahay |
| 2025 | Communication Notification Through User-Level Interrupts for the BXI Network. | Charles Goedefroit, Alexandre Denis, Mathieu Barbe, Brice Goglin, Grgoire Pichon |
| 2025 | GPU-CPU Shared Memory Performance Analysis on NVIDIA GH200. | Norihisa Fujita, Taisuke Boku, Tomo Yoshida, Takuto Shirai, Miwako Tsuji |
| 2025 | Deadline-Aware Resource Allocation and Scheduling of Serverless Workloads on Heterogeneous Clusters. | Matthias Fritz, Siegfried Benkner, Enes Bajrovic |
| 2025 | Closing the HPC-Cloud Convergence Gap: Multi-Tenant Slingshot RDMA for Kubernetes. | Philipp A. Friese, Ahmed Eleliemy, Utz-Uwe Haus, Martin Schulz |
| 2025 | Scaling Deep Learning Molecular Dynamics to 500M Atoms on 4096-Node ARMv8 Clusters. | Qi Du, Feng Wang, Chengkun Wu, Han Wang, Yongpeng Liu, Zhaoyin Zhou, Kenli Li |
| 2025 | TRACE: A Targeted Recommender for VM Assignment in Cloud Environment. | Hongji Dong, Yunlong Cheng, Tin Ping Chan, Xiaofeng Gao, Guihai Chen |
| 2025 | CFseq: A Framework for Constructing Compression-Friendly Field Sequences for Network Logs. | Yunwei Dai, Tao Huang, Shuo Wang, Yong Wang |
| 2025 | NSYS2PRV: Detailed and Quantitative Analysis of Large-Scale GPU Execution Traces with Paraver. | Marc Clasc, Jess Labarta, Marta Garcia-Gasulla |