| 2026 | HPCA | eGPU: Production-Scale Elastic Sharing Over 10,000 GPUs. | Xiaochuan Tang, Hao Qi, Jianbo Dong, Yinghao Yu, Zhennan Xue, Zhengyu Zhang, Daocheng Ying, Zheng Cao, Xiaoyi Lu |
| 2026 | HPDC | When RDMA Goes Long-Haul: Characterization, Modeling, and Verbs-Level Emulation with Implications for Federated Learning. | Yuke Li, Zhonghao Chen, Xiaoyi Lu |
| 2026 | ICS | PACER: A Userspace Network Rate Controller in MPI with Adaptive Compression for Parallel Applications. | Yuke Li, Darren Ng, Arjun Kashyap, Sheng Di, Guanpeng Li, Xiaoyi Lu |
| 2025 | EWSN | PULSE: Power Usage Monitoring Leveraging Sensing of Electromagnetic Field. | Rahul Sidramappa Hoskeri, Xiaoyi Lu, Hua Huang |
| 2025 | HPDC | DPU-KV: On the Benefits of DPU Offloading for In-Memory Key-Value Stores at the Edge. | Arjun Kashyap, Yuke Li, Xiaoyi Lu |
| 2025 | ICS | Understanding the Idiosyncrasies of Emerging BlueField DPUs. | Arjun Kashyap, Yuke Li, Darren Ng, Xiaoyi Lu |
| 2025 | PPoPP | SBMGT: Scaling Bayesian Multinomial Group Testing. | Weicong Chen, Hao Qi, Curtis Tatsuoka, Xiaoyi Lu |
| 2025 | SC | Compression Error Sensitivity Analysis for Different Experts in MoE Model Inference. | Songkai Ma, Zhaorui Zhang, Sheng Di, Benben Liu, Xiaodong Yu, Xiaoyi Lu, Dan Wang |
| 2025 | SC | DPAR: High-Performance, Secure, and Scalable Differential Privacy-based AllReduce. | Hao Qi, Weicong Chen, Chenghong Wang, Xiaoyi Lu |
| 2025 | SC | HPC-R1: Characterizing R1-like Large Reasoning Models on HPC. | Adam Weingram, Duo Zhang, Zhonghao Chen, Hao Qi, Xiaoyi Lu |
| 2024 | ICNP | Kspeed: Beating I/O Bottlenecks of Data Provisioning for RDMA Training Clusters. | Jianbo Dong, Hao Qi, Tianjing Xu, Xiaoli Liu, Chen Wei, Rongyao Wang, Xiaoyi Lu, Zheng Cao, Binzhang Fu |
| 2024 | ICS | gZCCL: Compression-Accelerated Collective Communication Framework for GPU Clusters. | Jiajun Huang, Sheng Di, Xiaodong Yu, Yujia Zhai, Jinyang Liu, Yafan Huang, Ken Raffenetti, Hui Zhou, Kai Zhao, Xiaoyi Lu, Zizhong Chen, Franck Cappello, Yanfei Guo, Rajeev Thakur |
| 2024 | SC | hZCCL: Accelerating Collective Communication with Co-Designed Homomorphic Compression. | Jiajun Huang, Sheng Di, Xiaodong Yu, Yujia Zhai, Jinyang Liu, Zizhe Jian, Xin Liang, Kai Zhao, Xiaoyi Lu, Zizhong Chen, Franck Cappello, Yanfei Guo, Rajeev Thakur |
| 2024 | SC | Versatile Datapath Soft Error Detection on the Cheap for HPC Applications. | Yafan Huang, Sheng Di, Zhaorui Zhang, Xiaoyi Lu, Guanpeng Li |
| 2023 | HOTI | Characterizing Lossy and Lossless Compression on Emerging BlueField DPU Architectures. | Yuke Li, Arjun Kashyap, Yanfei Guo, Xiaoyi Lu |
| 2023 | HOTI | Performance Characterization of Large Language Models on High-Speed Interconnects. | Hao Qi, Liuyao Dai, Weicong Chen, Zhen Jia, Xiaoyi Lu |
| 2022 | HiPC | HiBGT: High-Performance Bayesian Group Testing for COVID-19. | Weicong Chen, Curtis Tatsuoka, Xiaoyi Lu |
| 2022 | HPDC | NVMe-oAF: Towards Adaptive NVMe-oF for IO-Intensive Workloads on HPC Cloud. | Arjun Kashyap, Xiaoyi Lu |
| 2021 | CIS | Global Adaptive Optimization Parameters For Robust Pupil Location. | Yang Wang, Xiaoyi Lu, Wenjun Zhou |
| 2021 | HPDC | DStore: A Fast, Tailless, and Quiescent-Free Object Store for PMEM. | Shashank Gugnani, Xiaoyi Lu |
| 2021 | SC | HatRPC: hint-accelerated thrift RPC over RDMA. | Tianxi Li, Haiyang Shi, Xiaoyi Lu |
| 2020 | SC | RDMP-KV: designing remote direct memory persistence based key-value stores with PMEM. | Tianxi Li, Dipti Shankar, Shashank Gugnani, Xiaoyi Lu |
| 2020 | SC | INEC: fast and coherent in-network erasure coding. | Haiyang Shi, Xiaoyi Lu |
| 2019 | HiPC | SCOR-KV: SIMD-Aware Client-Centric and Optimistic RDMA-Based Key-Value Store for Emerging CPU Architectures. | Dipti Shankar, Xiaoyi Lu, Dhabaleswar K. Panda |
| 2019 | HPDC | UMR-EC: A Unified and Multi-Rail Erasure Coding Library for High-Performance Distributed Storage Systems. | Haiyang Shi, Xiaoyi Lu, Dipti Shankar, Dhabaleswar K. Panda |
| 2019 | SC | TriEC: tripartite graph based erasure coding NIC offload. | Haiyang Shi, Xiaoyi Lu |
| 2018 | CLOUD | High-Performance Multi-Rail Erasure Coding Library over Modern Data Center Architectures: Early Experiences. | Haiyang Shi, Xiaoyi Lu, Dipti Shankar, Dhabaleswar K. Panda |
| 2018 | CLUSTER | Cutting the Tail: Designing High Performance Message Brokers to Reduce Tail Latencies in Stream Processing. | M. Haseeb Javed, Xiaoyi Lu, Dhabaleswar K. Panda |
| 2018 | HiPC | OC-DNN: Exploiting Advanced Unified Memory Capabilities in CUDA 9 and Volta GPUs for Out-of-Core DNN Training. | Ammar Ahmad Awan, Ching-Hsiang Chu, Hari Subramoni, Xiaoyi Lu, Dhabaleswar K. Panda |
| 2018 | HiPC | Accelerating TensorFlow with Adaptive RDMA-Based gRPC. | Rajarshi Biswas, Xiaoyi Lu, Dhabaleswar K. Panda |
| 2018 | UCC | Analyzing, Modeling, and Provisioning QoS for NVMe SSDs. | Shashank Gugnani, Xiaoyi Lu, Dhabaleswar K. Panda |
| 2017 | CCGRID | Swift-X: Accelerating OpenStack Swift with RDMA for Building an Efficient HPC Cloud. | Shashank Gugnani, Xiaoyi Lu, Dhabaleswar K. Panda |
| 2017 | CLUSTER | A Scalable Network-Based Performance Analysis Tool for MPI on Large-Scale HPC Systems. | Hari Subramoni, Xiaoyi Lu, Dhabaleswar K. Panda |
| 2017 | HiPC | MPI-LiFE: Designing High-Performance Linear Fascicle Evaluation of Brain Connectome with MPI. | Shashank Gugnani, Xiaoyi Lu, Franco Pestilli, Cesar F. Caiafa, Dhabaleswar K. Panda |
| 2017 | HiPC | Designing Registration Caching Free High-Performance MPI Library with Implicit On-Demand Paging (ODP) of InfiniBand. | Mingzhe Li, Xiaoyi Lu, Hari Subramoni, Dhabaleswar K. Panda |
| 2017 | HOTI | Characterizing Deep Learning over Big Data (DLoBD) Stacks on RDMA-Capable Networks. | Xiaoyi Lu, Haiyang Shi, M. Haseeb Javed, Rajarshi Biswas, Dhabaleswar K. Panda |
| 2017 | ICDCS | High-Performance and Resilient Key-Value Store with Online Erasure Coding for Big Data Workloads. | Dipti Shankar, Xiaoyi Lu, Dhabaleswar K. Panda |
| 2017 | ICPP | Efficient and Scalable Multi-Source Streaming Broadcast on GPU Clusters for Deep Learning. | Ching-Hsiang Chu, Xiaoyi Lu, Ammar Ahmad Awan, Hari Subramoni, Jahanzeb Maqbool Hashmi, Bracy Elton, Dhabaleswar K. Panda |
| 2017 | SC | Scalable reduction collectives with data partitioning-based multi-leader design. | Mohammadreza Bayatpour, Sourav Chakraborty, Hari Subramoni, Xiaoyi Lu, Dhabaleswar K. Panda |
| 2017 | UCC | HPC Meets Cloud: Building Efficient Clouds for HPC, Big Data, and Deep Learning Middleware and Applications. | Dhabaleswar K. Panda, Xiaoyi Lu |
| 2017 | UCC | Is Singularity-based Container Technology Ready for Running MPI Applications on HPC Clouds? | Jie Zhang, Xiaoyi Lu, Dhabaleswar K. Panda |
| 2017 | VEE | Designing Locality and NUMA Aware MPI Runtime for Nested Virtualization based HPC Cloud with SR-IOV Enabled InfiniBand. | Jie Zhang, Xiaoyi Lu, Dhabaleswar K. Panda |
| 2016 | CloudCom | Designing Virtualization-Aware and Automatic Topology Detection Schemes for Accelerating Hadoop on SR-IOV-Enabled Clouds. | Shashank Gugnani, Xiaoyi Lu, Dhabaleswar K. Panda |
| 2016 | CloudCom | Impact of HPC Cloud Networking Technologies on Accelerating Hadoop RPC and HBase. | Xiaoyi Lu, Dipti Shankar, Shashank Gugnani, Hari Subramoni, Dhabaleswar K. Panda |
| 2016 | EuroPar | Slurm-V: Extending Slurm for Building Efficient HPC Cloud with SR-IOV and IVShmem. | Jie Zhang, Xiaoyi Lu, Sourav Chakraborty, Dhabaleswar K. Panda |
| 2016 | HiPC | Mizan-RMA: Accelerating Mizan Graph Processing Framework with MPI RMA. | Mingzhe Li, Xiaoyi Lu, Khaled Hamidouche, Jie Zhang, Dhabaleswar K. Panda |
| 2016 | ICPP | High Performance MPI Library for Container-Based HPC Cloud on InfiniBand Clusters. | Jie Zhang, Xiaoyi Lu, Dhabaleswar K. Panda |
| 2016 | ICS | High Performance Design for HDFS with Byte-Addressability of NVM and RDMA. | Nusrat Sharmin Islam, Md. Wasi-ur-Rahman, Xiaoyi Lu, Dhabaleswar K. Panda |
| 2016 | SC | Designing MPI library with on-demand paging (ODP) of infiniband: challenges and benefits. | Mingzhe Li, Khaled Hamidouche, Xiaoyi Lu, Hari Subramoni, Jie Zhang, Dhabaleswar K. Panda |
| 2016 | SC | Can Non-volatile Memory Benefit MapReduce Applications on HPC Clusters? | Md. Wasi-ur-Rahman, Nusrat Sharmin Islam, Xiaoyi Lu, Dhabaleswar K. Panda |
| 2016 | SBAC-PAD | MR-Advisor: A Comprehensive Tuning Tool for Advising HPC Users to Accelerate MapReduce Applications on Supercomputers. | Md. Wasi-ur-Rahman, Nusrat Sharmin Islam, Xiaoyi Lu, Dipti Shankar, Dhabaleswar K. Panda |
| 2015 | CCGRID | Triple-H: A Hybrid Approach to Accelerate HDFS on HPC Clusters with Heterogeneous Storage Architecture. | Nusrat Sharmin Islam, Xiaoyi Lu, Md. Wasi-ur-Rahman, Dipti Shankar, Dhabaleswar K. Panda |
| 2015 | CCGRID | MVAPICH2 over OpenStack with SR-IOV: An Efficient Approach to Build HPC Clouds. | Jie Zhang, Xiaoyi Lu, Mark Daniel Arnold, Dhabaleswar K. Panda |
| 2015 | CLUSTER | High Performance MPI Datatype Support with User-Mode Memory Registration: Challenges, Designs, and Benefits. | Mingzhe Li, Hari Subramoni, Khaled Hamidouche, Xiaoyi Lu, Dhabaleswar K. Panda |
| 2015 | EuroPar | High-Performance and Scalable Design of MPI-3 RMA on Xeon Phi Clusters. | Mingzhe Li, Khaled Hamidouche, Xiaoyi Lu, Jian Lin, Dhabaleswar K. Panda |
| 2015 | HiPC | High Performance OpenSHMEM Strided Communication Support with InfiniBand UMR. | Mingzhe Li, Khaled Hamidouche, Xiaoyi Lu, Jie Zhang, Jian Lin, Dhabaleswar K. Panda |
| 2015 | ICDCS | Accelerating Apache Hive with MPI for Data Warehouse Systems. | Lu Chao, Chundian Li, Fan Liang, Xiaoyi Lu, Zhiwei Xu |
| 2015 | ICPP | Accelerating I/O Performance of Big Data Analytics on HPC Clusters through RDMA-Based Key-Value Store. | Nusrat Sharmin Islam, Dipti Shankar, Xiaoyi Lu, Md. Wasi-ur-Rahman, Dhabaleswar K. Panda |
| 2015 | ISPASS | Can RDMA benefit online data processing workloads on memcached and MySQL? | Dipti Shankar, Xiaoyi Lu, Jithin Jose, Md. Wasi-ur-Rahman, Nusrat S. Islam, Dhabaleswar K. Panda |
| 2014 | CLUSTER | High performance OpenSHMEM for Xeon Phi clusters: Extensions, runtime designs and application co-design. | Jithin Jose, Khaled Hamidouche, Xiaoyi Lu, Sreeram Potluri, Jie Zhang, Karen Tomko, Dhabaleswar K. Panda |
| 2014 | CLUSTER | Scalable Graph500 design with MPI-3 RMA. | Mingzhe Li, Xiaoyi Lu, Sreeram Potluri, Khaled Hamidouche, Jithin Jose, Karen Tomko, Dhabaleswar K. Panda |
| 2014 | EuroPar | MapReduce over Lustre: Can RDMA-Based Approach Benefit? | Md. Wasi-ur-Rahman, Xiaoyi Lu, Nusrat Sharmin Islam, Raghunath Rajachandrasekar, Dhabaleswar K. Panda |
| 2014 | EuroPar | Can Inter-VM Shmem Benefit MPI Applications on SR-IOV Based Virtualized Infiniband Clusters? | Jie Zhang, Xiaoyi Lu, Jithin Jose, Rong Shi, Dhabaleswar K. Panda |
| 2014 | HiPC | High performance MPI library over SR-IOV enabled infiniband clusters. | Jie Zhang, Xiaoyi Lu, Jithin Jose, Mingzhe Li, Rong Shi, Dhabaleswar K. Panda |
| 2014 | HOTI | Accelerating Spark with RDMA for Big Data Processing: Early Experiences. | Xiaoyi Lu, Md. Wasi-ur-Rahman, Nusrat S. Islam, Dipti Shankar, Dhabaleswar K. Panda |
| 2014 | HPDC | SOR-HDFS: a SEDA-based approach to maximize overlapping in RDMA-enhanced HDFS. | Nusrat S. Islam, Xiaoyi Lu, Md. Wasi-ur-Rahman, Dhabaleswar K. Panda |
| 2014 | ICPP | HAND: A Hybrid Approach to Accelerate Non-contiguous Data Movement Using MPI Datatypes on GPU Clusters. | Rong Shi, Xiaoyi Lu, Sreeram Potluri, Khaled Hamidouche, Jie Zhang, Dhabaleswar K. Panda |
| 2014 | ICPP | Performance Modeling for RDMA-Enhanced Hadoop MapReduce. | Md. Wasi-ur-Rahman, Xiaoyi Lu, Nusrat Sharmin Islam, Dhabaleswar K. Panda |
| 2014 | ICS | HOMR: a hybrid approach to exploit maximum overlapping in MapReduce over high performance interconnects. | Md. Wasi-ur-Rahman, Xiaoyi Lu, Nusrat Sharmin Islam, Dhabaleswar K. Panda |
| 2014 | PPoPP | Initial study of multi-endpoint runtime for MPI+OpenMP hybrid programming model on multi-core systems. | Miao Luo, Xiaoyi Lu, Khaled Hamidouche, Krishna Chaitanya Kandalla, Dhabaleswar K. Panda |
| 2014 | VLDB | On Big Data Benchmarking. | Rui Han, Xiaoyi Lu, Jiangtao Xu |
| 2014 | VLDB | Performance Benefits of DataMPI: A Case Study with BigDataBench. | Fan Liang, Chen Feng, Xiaoyi Lu, Zhiwei Xu |
| 2014 | VLDB | A Micro-benchmark Suite for Evaluating Hadoop MapReduce on High-Performance Networks. | Dipti Shankar, Xiaoyi Lu, Md. Wasi-ur-Rahman, Nusrat S. Islam, Dhabaleswar K. Panda |
| 2013 | CCGRID | SR-IOV Support for Virtualization on InfiniBand Clusters: Early Experience. | Jithin Jose, Mingzhe Li, Xiaoyi Lu, Krishna Chaitanya Kandalla, Mark Daniel Arnold, Dhabaleswar K. Panda |
| 2013 | CLOUD | Does RDMA-based enhanced Hadoop MapReduce need a new performance model? | Md. Wasi-ur-Rahman, Xiaoyi Lu, Nusrat S. Islam, Dhabaleswar K. Panda |
| 2013 | CLUSTER | A scalable and portable approach to accelerate hybrid HPL on heterogeneous CPU-GPU clusters. | Rong Shi, Sreeram Potluri, Khaled Hamidouche, Xiaoyi Lu, Karen Tomko, Dhabaleswar K. Panda |
| 2013 | HOTI | Tutorials. | Dhabaleswar K. Panda, Xiaoyi Lu |
| 2013 | HOTI | Can Parallel Replication Benefit Hadoop Distributed File System for High Performance Interconnects? | Nusrat S. Islam, Xiaoyi Lu, Md. Wasi-ur-Rahman, Dhabaleswar K. Panda |
| 2013 | ICPP | High-Performance Design of Hadoop RPC with RDMA over InfiniBand. | Xiaoyi Lu, Nusrat S. Islam, Md. Wasi-ur-Rahman, Jithin Jose, Hari Subramoni, Hao Wang, Dhabaleswar K. Panda |
| 2011 | ISPA | Vega LingCloud: A Resource Single Leasing Point System to Support Heterogeneous Application Modes on Shared Infrastructure. | Xiaoyi Lu, Jian Lin, Li Zha, Zhiwei Xu |
| 2010 | NPC | JAMILA: A Usable Batch Job Management System to Coordinate Heterogeneous Clusters and Diverse Applications over Grid or Cloud Infrastructure. | Juan Peng, Xiaoyi Lu, Boqun Cheng, Li Zha |
| 2010 | SERVICES | Investigating, Modeling, and Ranking Interface Complexity of Web Services on the World Wide Web. | Xiaoyi Lu, Jian Lin, Yongqiang Zou, Juan Peng, Xingwu Liu, Li Zha |
| 2009 | PDCAT | ICOMC: Invocation Complexity Of Multi-Language Clients for Classified Web Services and its Impact on Large Scale SOA Applications. | Xiaoyi Lu, Yongqiang Zou, Fei Xiong, Jian Lin, Li Zha |
| 2009 | SERVICES | A Model of Message-Based Debugging Facilities for Web or Grid Services. | Qiang Yue, Xiaoyi Lu, Zhiguang Shan, Zhiwei Xu, Haiyan Yu, Li Zha |
| 2008 | PDCAT | An Experimental Analysis for Memory Usage of GOS Core. | Xiaoyi Lu, Qiang Yue, Yongqiang Zou, Xiaoning Wang |