| 2024 | GMM: An Efficient GPU Memory Management-based Model Serving System for Multiple DNN Inference Models. | XinYu Piao, Jong-Kook Kim |
| 2024 | Sparse Gradient Communication with AlltoAll for Accelerating Distributed Deep Learning. | Jing Peng, Zihan Li, Shaohuai Shi, Bo Li |
| 2024 | BandSlim: A Novel Bandwidth and Space-Efficient KV-SSD with an Escape-from-Block Approach. | Junhyeok Park, Chang-Gyu Lee, Soon Hwang, Soonyeal Yang, Jungki Noh, Woosuk Chung, Junghee Lee, Youngjae Kim |
| 2024 | Improving efficiency of Monte Carlo method via code intrinsic framework. | Qifeng Pan, Ralf Schneider |
| 2024 | Pluto and Charon: A Time and Memory Efficient Collaborative Edge AI Framework for Personal LLMs Fine-tuning. | Bei Ouyang, Shengyuan Ye, Liekang Zeng, Tianyi Qian, Jingyi Li, Xu Chen |
| 2024 | Selective Memory Compression for GPU Memory Oversubscription Management. | Abdun Nihaal, Madhu Mutyam |
| 2024 | Significantly Improving Fixed-Ratio Compression Framework for Resource-limited Applications. | Tri Nguyen, Md Hasanur Rahman, Sheng Di, Michela Becchi |
| 2024 | Sparsity-Aware Communication for Distributed Graph Neural Network Training. | Ujjaini Mukhopadhyay, Alok Tripathy, Oguz Selvitopi, Katherine A. Yelick, Aydin Bulu |
| 2024 | Detailed Analysis and Optimization of Irregular-Shaped Matrix Multiplication on Multi-Core DSPs. | Haotian Mo, Qinglin Wang, Linyu Liao, Biao Li, Lihua Chi, Jie Liu |
| 2024 | Parallelization of the Banded Needleman & Wunsch Algorithm on UPMEM PiM Architecture for Long DNA Sequence Alignment. | Meven Mognol, Dominique Lavenier, Julien Legriel |
| 2024 | Cache Line Pinning for Mitigating Row Hammer Attack. | Praseetha M, Madhu Mutyam, Venkata Kalyan Tavva |
| 2024 | zQoS: Unleashing full performance capabilities of NVMe SSDs while enforcing SLOs in distributed storage systems. | Liuying Ma, Zhenqing Liu, Jin Xiong, Yue Wu, Renhai Chen, Xi Peng, Ying Zhang, Gong Zhang, Dejun Jiang |
| 2024 | A Hybrid Machine Learning Method for Cross-Platform Performance Prediction of Parallel Applications. | Kaveh Mahdavi |
| 2024 | FedCA: Efficient Federated Learning with Client Autonomy. | Na Lv, Zhi Shen, Chen Chen, Zhifeng Jiang, Jiayi Zhang, Quan Chen, Minyi Guo |
| 2024 | Federated Edge Learning with Blurred or Pseudo Data Sharing. | Yinlong Li, Hao Zhang, Siyao Cheng, Jie Liu |
| 2024 | In-Situ Binary Segmentation of 3D time-dependent Flows into Laminar and Turbulent Regions. | Jiahui Liu, Tobias Edwards, Kristina Durovic, Philipp Schlatter, Tino Weinkauf |
| 2024 | Hi-ZNS: High Space Efficiency and Zero-Copy LSM-Tree Based Stores on ZNS SSDs. | Renping Liu, Junhua Chen, Peng Chen, Linbo Long, Anping Xiong, Duo Liu |
| 2024 | Murmuration: On-the-fly DNN Adaptation for SLO-Aware Distributed Inference in Dynamic Edge Environments. | Jieyu Lin, Minghao Li, Sai Qian Zhang, Alberto Leon-Garcia |
| 2024 | Designing Non-uniform Locally Repairable Codes for Wide Stripes under Skewed File Accesses. | Guantian Lin, Si Wu, Cheng Li, Yinlong Xu |
| 2024 | RIA: Return on Investment Auto-scaler for Serverless Edge Functions. | Huadong Li, Hui Liu, Aoqi Chen, Xirui Ma, Qiaoqiao Liu, Junzhao Du |
| 2024 | Thawbringer: An Orchestrator to Mitigate Cascading Cold Starts of Serverless Function Chains. | Huadong Li, Hui Liu, Aoqi Chen, Xirui Ma, Junzhao Du |
| 2024 | High-Performance Sorting-Based K-mer Counting in Distributed Memory with Flexible Hybrid Parallelism. | Yifan Li, Giulia Guidi |
| 2024 | High-Performance 3D convolution on the Latest Generation Sunway Processor. | Jialin Li, Zhichen Feng, Yaqian Gao, Shaobo Tian, Haoyuan Zhang, Huang Ye, Jian Zhang |
| 2024 | Rethinking Low-Carbon Edge Computing System Design with Renewable Energy Sharing. | Hanlong Liao, Guoming Tang, Deke Guo, Yi Wang, Ruide Cao |
| 2024 | Exploring Scalability in C++ Parallel STL Implementations. | Ruben Laso, Diego Krupitza, Sascha Hunold |