| 2026 | HPDC | HAT-MPI: Hierarchical Auto Tuning of MPI Inter-Node Communication on InfiniBand Clusters. | Sungjae Lee, Sam Tilford, Dhabaleswar K. Panda |
| 2026 | WACV | Supporting Ultra-High-Resolution Digital Agriculture Tasks with Fully Synthetic Curriculum Learning. | Jacob Hatef, Quentin Anthony, Nawras Alnaasan, Dhabaleswar K. Panda |
| 2025 | CLUSTER | Towards Dynamic Message Passing Protocols for Stencil-Based Communication Patterns. | Kaushik Kandadi Suresh, Bharath Ramesh, Goutham Kalikrishna Reddy Kuncham, Hari Subramoni, Dhabaleswar K. Panda |
| 2025 | HiPC | Performance Characterization of Data Transfer and Allocation Strategies on AMD MI300A APUs: Early Experiences. | Goutham Kalikrishna Reddy Kuncham, Siyuan Zhang, Bharath Ramesh, Kaushik Kandadi Suresh, Dhabaleswar K. Panda |
| 2025 | HiPC | Enhanced MPI Intra-Node Communication Framework: A Hybrid Approach with Cooperative DMA Channel-Based Data Transfer. | Shulei Xu, Tu Tran, Dhabaleswar K. Panda |
| 2025 | HOTI | Characterizing Communication Patterns in Distributed Large Language Model Inference. | Lang Xu, Kaushik Kandadi Suresh, Quentin Anthony, Nawras Alnaasan, Dhabaleswar K. Panda |
| 2025 | ICPP | Design and Optimization of GPU-Aware MPI Allreduce Using Direct Sendrecv Communication. | Chen-Chun Chen, Jinghan Yao, Hari Subramoni, Dhabaleswar K. Panda |
| 2025 | SC | OpenSHMEM MLIR: A Dialect for Compile-Time Optimization of One-Sided Communications. | Michael Beebe, Benjamin Michalowicz, Andrew McNamara, Yash Kumar, Dhabaleswar K. Panda, Yong Chen, Wendy Poole, Steve Poole |
| 2025 | SC | A Streaming Collectives Interface Targeting Dataflow Acceleration and HPC Workloads. | Nicholas Contini, Jake Queiser, Bharath Ramesh, Hari Subramoni, Dhabaleswar K. Panda |
| 2025 | SC | MPI Communication Performance on AMD MI300A: Microbenchmarks and Applications. | Goutham Kalikrishna Reddy Kuncham, Siyuan Zhang, Shoaib Mohammad, Chen-Chun Chen, Dhabaleswar K. Panda |
| 2024 | HiPC | HyperSack: Distributed Hyperparameter Optimization for Deep Learning using Resource-Aware Scheduling on Heterogeneous GPU Systems. | Nawras Alnaasan, Bharath Ramesh, Jinghan Yao, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda |
| 2024 | HiPC | Design and Implementation of Kernel-based MPI Reduction Operations for Intel GPU s. | Chen-Chun Chen, Goutham Kalikrishna Reddy Kuncham, Hari Subramoni, Dhabaleswar K. Panda |
| 2024 | HiPC | Effective and Efficient Offloading Designs for One-Sided Communication to SmartNICs. | Benjamin Michalowicz, Kaushik Kandadi Suresh, Hari Subramoni, Mustafa Abduljabbar, Dhabaleswar K. Panda, Steve Poole |
| 2024 | HiPC | Using BlueField-3 SmartNICs to Offload Vector Operations in Krylov Subspace Methods. | Kaushik Kandadi Suresh, Benjamin Michalowicz, Nick Contini, Bharath Ramesh, Mustafa Abduljabbar, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda |
| 2024 | HiPC | Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning. | Lang Xu, Quentin Anthony, Jacob Hatef, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda |
| 2024 | HOTI | Characterizing Communication in Distributed Parameter-Efficient Fine-Tuning for Large Language Models. | Nawras Alnaasan, Horng-Ruey Huang, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda |
| 2024 | HOTI | Demystifying the Communication Characteristics for Distributed Transformer Models. | Quentin Anthony, Benjamin Michalowicz, Jacob Hatef, Lang Xu, Mustafa Abdul Jabbar, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda |
| 2024 | HOTI | OHIO: Improving RDMA Network Scalability in MPI_Alltoall Through Optimized Hierarchical and Intra/Inter-Node Communication Overlap Design. | Tu Tran, Goutham Kalikrishna Reddy Kuncham, Bharath Ramesh, Shulei Xu, Hari Subramoni, Mustafa Abduljabbar, Dhabaleswar K. Panda |
| 2024 | ICPP | The Case for Co-Designing Model Architectures with Hardware. | Quentin Anthony, Jacob Hatef, Deepak Narayanan, Stella Biderman, Stas Bekman, Junqi Yin, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda |
| 2023 | CCGRID | ScaMP: Scalable Meta-Parallelism for Deep Learning Search. | Quentin Anthony, Lang Xu, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda |
| 2023 | CCGRID | ScaMP: Scalable Meta-Parallelism for Deep Learning Search. | Quentin Anthony, Lang Xu, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda |
| 2023 | CCGRID | Implementing and Optimizing a GPU-aware MPI Library for Intel GPUs: Early Experiences. | Chen-Chun Chen, Kawthar Shafie Khorassani, Goutham Kalikrishna Reddy Kuncham, Rahul Vaidya, Mustafa Abduljabbar, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda |
| 2023 | ICFEC | Performance Characterization of Using Quantization for DNN Inference on Edge Devices. | Hyunho Ahn, Tian Chen, Nawras Alnaasan, Aamir Shafi, Mustafa Abduljabbar, Hari Subramoni, Dhabaleswar K. Panda |
| 2023 | ICS | Enabling Reconfigurable HPC through MPI-based Inter-FPGA Communication. | Nicholas Contini, Bharath Ramesh, Kaushik Kandadi Suresh, Tu Tran, Benjamin Michalowicz, Mustafa Abduljabbar, Hari Subramoni, Dhabaleswar K. Panda |
| 2023 | SC | MPI-xCCL: A Portable MPI Library over Collective Communication Libraries for Various Accelerators. | Chen-Chun Chen, Kawthar Shafie Khorassani, Pouya Kousha, Qinghua Zhou, Jinghan Yao, Hari Subramoni, Dhabaleswar K. Panda |
| 2023 | SC | Democratizing HPC Access and Use with Knowledge Graphs. | Pouya Kousha, Vivekananda Sathu, Matthew Lieber, Hari Subramoni, Dhabaleswar K. Panda |
| 2022 | CLUSTER | Spark Meets MPI: Towards High-Performance Communication Framework for Spark using MPI. | Kinan Al-Attar, Aamir Shafi, Mustafa Abduljabbar, Hari Subramoni, Dhabaleswar K. Panda |
| 2022 | HiPC | AccDP: Accelerated Data-Parallel Distributed DNN Training for Modern GPU-Based HPC Clusters. | Nawras Alnaasan, Arpan Jain, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda |
| 2022 | HiPC | Designing Efficient Pipelined Communication Schemes using Compression in MPI Libraries. | Bharath Ramesh, Qinghua Zhou, Aamir Shafi, Mustafa Abduljabbar, Hari Subramoni, Dhabaleswar K. Panda |
| 2022 | HiPC | Efficient Personalized and Non-Personalized Alltoall Communication for Modern Multi-HCA GPU-Based Clusters. | Kaushik Kandadi Suresh, Akshay Paniraja Guptha, Benjamin Michalowicz, Bharath Ramesh, Mustafa Abduljabbar, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda |
| 2022 | HiPC | Accelerating Broadcast Communication with GPU Compression for Deep Learning Workloads. | Qinghua Zhou, Quentin Anthony, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda |
| 2022 | HOTI | Network Assisted Non-Contiguous Transfers for GPU-Aware MPI Libraries. | Kaushik Kandadi Suresh, Kawthar Shafie Khorassani, Chen-Chun Chen, Bharath Ramesh, Mustafa Abduljabbar, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda |
| 2021 | CCGRID | Adaptive and Hierarchical Large Message All-to-all Communication Algorithms for Large-scale Dense GPU Systems. | Kawthar Shafie Khorassani, Ching-Hsiang Chu, Quentin G. Anthony, Hari Subramoni, Dhabaleswar K. Panda |
| 2021 | HiPC | Towards Architecture-aware Hierarchical Communication Trees on Modern HPC Systems. | Bharath Ramesh, Jahanzeb Maqbool Hashmi, Shulei Xu, Aamir Shafi, Seyedeh Mahdieh Ghazimirsaeed, Mohammadreza Bayatpour, Hari Subramoni, Dhabaleswar K. Panda |
| 2021 | HiPC | DistMILE: A Distributed Multi-Level Framework for Scalable Graph Embedding. | Yuntian He, Saket Gurukar, Pouya Kousha, Hari Subramoni, Dhabaleswar K. Panda, Srinivasan Parthasarathy |
| 2021 | HiPC | Large-Message Nonblocking MPI_Iallgather and MPI Ibcast Offload via BlueField-2 DPU. | Nick Sarkauskas, Mohammadreza Bayatpour, Tu Tran, Bharath Ramesh, Hari Subramoni, Dhabaleswar K. Panda |
| 2021 | HiPC | Layout-aware Hardware-assisted Designs for Derived Data Types in MPI. | Kaushik Kandadi Suresh, Bharath Ramesh, Chen-Chun Chen, Seyedeh Mahdieh Ghazimirsaeed, Mohammadreza Bayatpour, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda |
| 2021 | HOTI | Accelerating CPU-based Distributed DNN Training on Modern HPC Clusters using BlueField-2 DPUs. | Arpan Jain, Nawras Alnaasan, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda |
| 2020 | CCGRID | Design and Characterization of InfiniBand Hardware Tag Matching in MPI. | Mohammadreza Bayatpour, Seyedeh Mahdieh Ghazimirsaeed, Shulei Xu, Hari Subramoni, Dhabaleswar K. Panda |
| 2020 | CLUSTER | Dynamic Kernel Fusion for Bulk Non-contiguous Data Transfer on GPU Clusters. | Ching-Hsiang Chu, Kawthar Shafie Khorassani, Qinghua Zhou, Hari Subramoni, Dhabaleswar K. Panda |
| 2020 | HiPC | Blink: Towards Efficient RDMA-based Communication Coroutines for Parallel Python Applications. | Aamir Shafi, Jahanzeb Maqbool Hashmi, Hari Subramoni, Dhabaleswar K. Panda |
| 2020 | SC | Scalable MPI Collectives using SHARP: Large Scale Performance Evaluation on the TACC Frontera System. | Bharath Ramesh, Kaushik Kandadi Suresh, Nick Sarkauskas, Mohammadreza Bayatpour, Jahanzeb Maqbool Hashmi, Hari Subramoni, Dhabaleswar K. Panda |
| 2020 | SC | GEMS: GPU-enabled memory-aware model-parallelism system for distributed DNN training. | Arpan Jain, Ammar Ahmad Awan, Asmaa M. Aljuhani, Jahanzeb Maqbool Hashmi, Quentin G. Anthony, Hari Subramoni, Dhabaleswar K. Panda, Raghu Machiraju, Anil Parwani |
| 2020 | SC | Exploring Hybrid MPI+Kokkos Tasks Programming Model. | Samuel Khuvis, Karen Tomko, Jahanzeb Maqbool Hashmi, Dhabaleswar K. Panda |
| 2019 | ASPLOS | Characterizing CUDA Unified Memory (UM)-Aware MPI Designs on Modern GPU Architectures. | Karthik Vadambacheri Manian, A. A. Ammar, Amit Ruhela, Ching-Hsiang Chu, Hari Subramoni, Dhabaleswar K. Panda |
| 2019 | CCGRID | Scalable Distributed DNN Training using TensorFlow and CUDA-Aware MPI: Characterization, Designs, and Performance Evaluation. | Ammar Ahmad Awan, Jeroen Bdorf, Ching-Hsiang Chu, Hari Subramoni, Dhabaleswar K. Panda |
| 2019 | CCGRID | Design and Characterization of Shared Address Space MPI Collectives on Modern Architectures. | Jahanzeb Maqbool Hashmi, Sourav Chakraborty, Mohammadreza Bayatpour, Hari Subramoni, Dhabaleswar K. Panda |
| 2019 | CLUSTER | Performance Characterization of DNN Training using TensorFlow and PyTorch on Modern Clusters. | Arpan Jain, Ammar Ahmad Awan, Quentin Anthony, Hari Subramoni, Dhabaleswar K. Panda |
| 2019 | HiPC | High-Performance Adaptive MPI Derived Datatype Communication for Modern Multi-GPU Systems. | Ching-Hsiang Chu, Jahanzeb Maqbool Hashmi, Kawthar Shafie Khorassani, Hari Subramoni, Dhabaleswar K. Panda |
| 2019 | HiPC | Designing a Profiling and Visualization Tool for Scalable and In-depth Analysis of High-Performance GPU Clusters. | Pouya Kousha, Bharath Ramesh, Kaushik Kandadi Suresh, Ching-Hsiang Chu, Arpan Jain, Nick Sarkauskas, Hari Subramoni, Dhabaleswar K. Panda |
| 2019 | HiPC | SCOR-KV: SIMD-Aware Client-Centric and Optimistic RDMA-Based Key-Value Store for Emerging CPU Architectures. | Dipti Shankar, Xiaoyi Lu, Dhabaleswar K. Panda |
| 2019 | HOTI | Designing Scalable and High-Performance MPI Libraries on Amazon Elastic Fabric Adapter. | Sourav Chakraborty, Shulei Xu, Hari Subramoni, Dhabaleswar K. Panda |
| 2019 | HOTI | Communication Profiling and Characterization of Deep Learning Workloads on Clusters with High-Performance Interconnects. | Ammar Ahmad Awan, Arpan Jain, Ching-Hsiang Chu, Hari Subramoni, Dhabaleswar K. Panda |
| 2019 | HPDC | UMR-EC: A Unified and Multi-Rail Erasure Coding Library for High-Performance Distributed Storage Systems. | Haiyang Shi, Xiaoyi Lu, Dipti Shankar, Dhabaleswar K. Panda |
| 2019 | PPoPP | High performance distributed deep learning: a beginner's guide. | Dhabaleswar K. Panda, Ammar Ahmad Awan, Hari Subramoni |
| 2019 | SC | Scaling TensorFlow, PyTorch, and MXNet using MVAPICH2 for High-Performance Deep Learning on Frontera. | Arpan Jain, Ammar Ahmad Awan, Hari Subramoni, Dhabaleswar K. Panda |
| 2019 | SC | Leveraging Network-level parallelism with Multiple Process-Endpoints for MPI Broadcast. | Amit Ruhela, Bharath Ramesh, Sourav Chakraborty, Hari Subramoni, Jahanzeb Maqbool Hashmi, Dhabaleswar K. Panda |
| 2018 | CLOUD | High-Performance Multi-Rail Erasure Coding Library over Modern Data Center Architectures: Early Experiences. | Haiyang Shi, Xiaoyi Lu, Dipti Shankar, Dhabaleswar K. Panda |
| 2018 | CLUSTER | SALaR: Scalable and Adaptive Designs for Large Message Reduction Collectives. | Mohammadreza Bayatpour, Jahanzeb Maqbool Hashmi, Sourav Chakraborty, Hari Subramoni, Pouya Kousha, Dhabaleswar K. Panda |
| 2018 | CLUSTER | Cutting the Tail: Designing High Performance Message Brokers to Reduce Tail Latencies in Stream Processing. | M. Haseeb Javed, Xiaoyi Lu, Dhabaleswar K. Panda |
| 2018 | HiPC | OC-DNN: Exploiting Advanced Unified Memory Capabilities in CUDA 9 and Volta GPUs for Out-of-Core DNN Training. | Ammar Ahmad Awan, Ching-Hsiang Chu, Hari Subramoni, Xiaoyi Lu, Dhabaleswar K. Panda |
| 2018 | HiPC | Accelerating TensorFlow with Adaptive RDMA-Based gRPC. | Rajarshi Biswas, Xiaoyi Lu, Dhabaleswar K. Panda |
| 2018 | SC | Cooperative rendezvous protocols for improved performance and overlap. | Sourav Chakraborty, Mohammadreza Bayatpour, Jahanzeb Maqbool Hashmi, Hari Subramoni, Dhabaleswar K. Panda |
| 2017 | CCGRID | Swift-X: Accelerating OpenStack Swift with RDMA for Building an Efficient HPC Cloud. | Shashank Gugnani, Xiaoyi Lu, Dhabaleswar K. Panda |
| 2017 | CLUSTER | Contention-Aware Kernel-Assisted MPI Collectives for Multi-/Many-Core Systems. | Sourav Chakraborty, Hari Subramoni, Dhabaleswar K. Panda |
| 2017 | CLUSTER | A Scalable Network-Based Performance Analysis Tool for MPI on Large-Scale HPC Systems. | Hari Subramoni, Xiaoyi Lu, Dhabaleswar K. Panda |
| 2017 | HiPC | MPI-LiFE: Designing High-Performance Linear Fascicle Evaluation of Brain Connectome with MPI. | Shashank Gugnani, Xiaoyi Lu, Franco Pestilli, Cesar F. Caiafa, Dhabaleswar K. Panda |
| 2017 | HiPC | Kernel-Assisted Communication Engine for MPI on Emerging Manycore Processors. | Jahanzeb Maqbool Hashmi, Khaled Hamidouche, Hari Subramoni, Dhabaleswar K. Panda |
| 2017 | HiPC | Designing Registration Caching Free High-Performance MPI Library with Implicit On-Demand Paging (ODP) of InfiniBand. | Mingzhe Li, Xiaoyi Lu, Hari Subramoni, Dhabaleswar K. Panda |
| 2017 | HOTI | Characterizing Deep Learning over Big Data (DLoBD) Stacks on RDMA-Capable Networks. | Xiaoyi Lu, Haiyang Shi, M. Haseeb Javed, Rajarshi Biswas, Dhabaleswar K. Panda |
| 2017 | ICDCS | High-Performance and Resilient Key-Value Store with Online Erasure Coding for Big Data Workloads. | Dipti Shankar, Xiaoyi Lu, Dhabaleswar K. Panda |
| 2017 | ICPP | Efficient and Scalable Multi-Source Streaming Broadcast on GPU Clusters for Deep Learning. | Ching-Hsiang Chu, Xiaoyi Lu, Ammar Ahmad Awan, Hari Subramoni, Jahanzeb Maqbool Hashmi, Bracy Elton, Dhabaleswar K. Panda |
| 2017 | ICPP | MPI-GDS: High Performance MPI Designs with GPUDirect-aSync for CPU-GPU Control Flow Decoupling. | Akshay Venkatesh, Khaled Hamidouche, Sreeram Potluri, Davide Rossetti, Ching-Hsiang Chu, Dhabaleswar K. Panda |
| 2017 | PPoPP | S-Caffe: Co-designing MPI Runtimes and Caffe for Scalable Deep Learning on Modern GPU Clusters. | Ammar Ahmad Awan, Khaled Hamidouche, Jahanzeb Maqbool Hashmi, Dhabaleswar K. Panda |
| 2017 | SC | An In-depth Performance Characterization of CPU- and GPU-based DNN Training on Modern Architectures. | Ammar Ahmad Awan, Hari Subramoni, Dhabaleswar K. Panda |
| 2017 | SC | Scalable reduction collectives with data partitioning-based multi-leader design. | Mohammadreza Bayatpour, Sourav Chakraborty, Hari Subramoni, Xiaoyi Lu, Dhabaleswar K. Panda |
| 2016 | CCGRID | SHMEMPMI - Shared Memory Based PMI for Improved Performance and Scalability. | Sourav Chakraborty, Hari Subramoni, Jonathan L. Perkins, Dhabaleswar K. Panda |
| 2016 | CCGRID | CUDA Kernel Based Collective Reduction Operations on Large-scale GPU Clusters. | Ching-Hsiang Chu, Khaled Hamidouche, Akshay Venkatesh, Ammar Ahmad Awan, Dhabaleswar K. Panda |
| 2016 | CloudCom | Re-Designing CNTK Deep Learning Framework on Modern GPU Enabled Clusters. | Dip Sankar Banerjee, Khaled Hamidouche, Dhabaleswar K. Panda |
| 2016 | CloudCom | Designing Virtualization-Aware and Automatic Topology Detection Schemes for Accelerating Hadoop on SR-IOV-Enabled Clouds. | Shashank Gugnani, Xiaoyi Lu, Dhabaleswar K. Panda |
| 2016 | CloudCom | Impact of HPC Cloud Networking Technologies on Accelerating Hadoop RPC and HBase. | Xiaoyi Lu, Dipti Shankar, Shashank Gugnani, Hari Subramoni, Dhabaleswar K. Panda |
| 2016 | CLUSTER | Adaptive and Dynamic Design for MPI Tag Matching. | Mohammadreza Bayatpour, Hari Subramoni, Sourav Chakraborty, Dhabaleswar K. Panda |
| 2016 | EuroPar | Slurm-V: Extending Slurm for Building Efficient HPC Cloud with SR-IOV and IVShmem. | Jie Zhang, Xiaoyi Lu, Sourav Chakraborty, Dhabaleswar K. Panda |
| 2016 | HiPC | CUDA M3: Designing Efficient CUDA Managed Memory-Aware MPI by Exploiting GDR and IPC. | Khaled Hamidouche, Ammar Ahmad Awan, Akshay Venkatesh, Dhabaleswar K. Panda |
| 2016 | HiPC | Mizan-RMA: Accelerating Mizan Graph Processing Framework with MPI RMA. | Mingzhe Li, Xiaoyi Lu, Khaled Hamidouche, Jie Zhang, Dhabaleswar K. Panda |
| 2016 | HPCC | Enabling Performance Efficient Runtime Support for Hybrid MPI+UPC++ Programming Models. | Jahanzeb Maqbool Hashmi, Khaled Hamidouche, Dhabaleswar K. Panda |
| 2016 | ICPADS | System-Level Scalable Checkpoint-Restart for Petascale Computing. | Jiajun Cao, Kapil Arya, Rohan Garg, L. Shawn Matott, Dhabaleswar K. Panda, Hari Subramoni, Jrme Vienne, Gene Cooperman |
| 2016 | ICPP | High Performance MPI Library for Container-Based HPC Cloud on InfiniBand Clusters. | Jie Zhang, Xiaoyi Lu, Dhabaleswar K. Panda |
| 2016 | ICS | High Performance Design for HDFS with Byte-Addressability of NVM and RDMA. | Nusrat Sharmin Islam, Md. Wasi-ur-Rahman, Xiaoyi Lu, Dhabaleswar K. Panda |
| 2016 | PPoPP | Designing high performance communication runtime for GPU managed memory: early experiences. | Dip Sankar Banerjee, Khaled Hamidouche, Dhabaleswar K. Panda |
| 2016 | SC | Efficient Reliability Support for Hardware Multicast-Based Broadcast in GPU-enabled Streaming Applications. | Ching-Hsiang Chu, Khaled Hamidouche, Hari Subramoni, Akshay Venkatesh, Bracy Elton, Dhabaleswar K. Panda |
| 2016 | SC | OpenSHMEM Non-blocking Data Movement Operations with MVAPICH2-X: Early Experiences. | Khaled Hamidouche, Jie Zhang, Dhabaleswar K. Panda, Karen Tomko |
| 2016 | SC | Designing MPI library with on-demand paging (ODP) of infiniband: challenges and benefits. | Mingzhe Li, Khaled Hamidouche, Xiaoyi Lu, Hari Subramoni, Jie Zhang, Dhabaleswar K. Panda |
| 2015 | CCGRID | Non-Blocking PMI Extensions for Fast MPI Startup. | Sourav Chakraborty, Hari Subramoni, Adam Moody, Akshay Venkatesh, Jonathan L. Perkins, Dhabaleswar K. Panda |
| 2015 | CCGRID | Power-Check: An Energy-Efficient Checkpointing Framework for HPC Clusters. | Raghunath Raja Chandrasekar, Akshay Venkatesh, Khaled Hamidouche, Dhabaleswar K. Panda |
| 2015 | CCGRID | Triple-H: A Hybrid Approach to Accelerate HDFS on HPC Clusters with Heterogeneous Storage Architecture. | Nusrat Sharmin Islam, Xiaoyi Lu, Md. Wasi-ur-Rahman, Dipti Shankar, Dhabaleswar K. Panda |
| 2015 | CCGRID | MVAPICH2 over OpenStack with SR-IOV: An Efficient Approach to Build HPC Clouds. | Jie Zhang, Xiaoyi Lu, Mark Daniel Arnold, Dhabaleswar K. Panda |
| 2015 | CLUSTER | Exploiting GPUDirect RDMA in Designing High Performance OpenSHMEM for NVIDIA GPU Clusters. | Khaled Hamidouche, Akshay Venkatesh, Ammar Ahmad Awan, Hari Subramoni, Ching-Hsiang Chu, Dhabaleswar K. Panda |
| 2015 | CLUSTER | High Performance MPI Datatype Support with User-Mode Memory Registration: Challenges, Designs, and Benefits. | Mingzhe Li, Hari Subramoni, Khaled Hamidouche, Xiaoyi Lu, Dhabaleswar K. Panda |
| 2015 | EuroPar | High-Performance and Scalable Design of MPI-3 RMA on Xeon Phi Clusters. | Mingzhe Li, Khaled Hamidouche, Xiaoyi Lu, Jian Lin, Dhabaleswar K. Panda |
| 2015 | HiPC | High Performance OpenSHMEM Strided Communication Support with InfiniBand UMR. | Mingzhe Li, Khaled Hamidouche, Xiaoyi Lu, Jie Zhang, Jian Lin, Dhabaleswar K. Panda |
| 2015 | HiPC | Offloaded GPU Collectives Using CORE-Direct and CUDA Capabilities on InfiniBand Clusters. | Akshay Venkatesh, Khaled Hamidouche, Hari Subramoni, Dhabaleswar K. Panda |
| 2015 | HOTI | Impact of InfiniBand DC Transport Protocol on Energy Consumption of All-to-All Collective Algorithms. | Hari Subramoni, Akshay Venkatesh, Khaled Hamidouche, Karen Tomko, Dhabaleswar K. Panda |
| 2015 | ICPP | Accelerating I/O Performance of Big Data Analytics on HPC Clusters through RDMA-Based Key-Value Store. | Nusrat Sharmin Islam, Dipti Shankar, Xiaoyi Lu, Md. Wasi-ur-Rahman, Dhabaleswar K. Panda |
| 2015 | ISPASS | Can RDMA benefit online data processing workloads on memcached and MySQL? | Dipti Shankar, Xiaoyi Lu, Jithin Jose, Md. Wasi-ur-Rahman, Nusrat S. Islam, Dhabaleswar K. Panda |
| 2014 | CLUSTER | High performance OpenSHMEM for Xeon Phi clusters: Extensions, runtime designs and application co-design. | Jithin Jose, Khaled Hamidouche, Xiaoyi Lu, Sreeram Potluri, Jie Zhang, Karen Tomko, Dhabaleswar K. Panda |
| 2014 | CLUSTER | Scalable Graph500 design with MPI-3 RMA. | Mingzhe Li, Xiaoyi Lu, Sreeram Potluri, Khaled Hamidouche, Jithin Jose, Karen Tomko, Dhabaleswar K. Panda |
| 2014 | EuroPar | MapReduce over Lustre: Can RDMA-Based Approach Benefit? | Md. Wasi-ur-Rahman, Xiaoyi Lu, Nusrat Sharmin Islam, Raghunath Rajachandrasekar, Dhabaleswar K. Panda |
| 2014 | EuroPar | Can Inter-VM Shmem Benefit MPI Applications on SR-IOV Based Virtualized Infiniband Clusters? | Jie Zhang, Xiaoyi Lu, Jithin Jose, Rong Shi, Dhabaleswar K. Panda |
| 2014 | HiPC | Designing efficient small message transfer mechanism for inter-node MPI communication on InfiniBand GPU clusters. | Rong Shi, Sreeram Potluri, Khaled Hamidouche, Jonathan L. Perkins, Mingzhe Li, Davide Rossetti, Dhabaleswar K. Panda |
| 2014 | HiPC | A high performance broadcast design with hardware multicast and GPUDirect RDMA for streaming applications on Infiniband clusters. | Akshay Venkatesh, Hari Subramoni, Khaled Hamidouche, Dhabaleswar K. Panda |
| 2014 | HiPC | High performance MPI library over SR-IOV enabled infiniband clusters. | Jie Zhang, Xiaoyi Lu, Jithin Jose, Mingzhe Li, Rong Shi, Dhabaleswar K. Panda |
| 2014 | HOTI | Accelerating Spark with RDMA for Big Data Processing: Early Experiences. | Xiaoyi Lu, Md. Wasi-ur-Rahman, Nusrat S. Islam, Dipti Shankar, Dhabaleswar K. Panda |
| 2014 | HPDC | SOR-HDFS: a SEDA-based approach to maximize overlapping in RDMA-enhanced HDFS. | Nusrat S. Islam, Xiaoyi Lu, Md. Wasi-ur-Rahman, Dhabaleswar K. Panda |
| 2014 | HPDC | MIC-Check: a distributed check pointing framework for the intel many integrated cores architecture. | Raghunath Rajachandrasekar, Sreeram Potluri, Akshay Venkatesh, Khaled Hamidouche, Md. Wasi-ur-Rahman, Dhabaleswar K. Panda |
| 2014 | ICPADS | Message from the general co-chairs IEEE ICPADS 2014. | Dhabaleswar K. Panda, Jang-Ping Sheu |
| 2014 | ICPP | HAND: A Hybrid Approach to Accelerate Non-contiguous Data Movement Using MPI Datatypes on GPU Clusters. | Rong Shi, Xiaoyi Lu, Sreeram Potluri, Khaled Hamidouche, Jie Zhang, Dhabaleswar K. Panda |
| 2014 | ICPP | Designing Topology-Aware Communication Schedules for Alltoall Operations in Large InfiniBand Clusters. | Hari Subramoni, Krishna Chaitanya Kandalla, Jithin Jose, Karen Tomko, Karl W. Schulz, Dmitry Pekurovsky, Dhabaleswar K. Panda |
| 2014 | ICPP | Performance Modeling for RDMA-Enhanced Hadoop MapReduce. | Md. Wasi-ur-Rahman, Xiaoyi Lu, Nusrat Sharmin Islam, Dhabaleswar K. Panda |
| 2014 | ICS | HOMR: a hybrid approach to exploit maximum overlapping in MapReduce over high performance interconnects. | Md. Wasi-ur-Rahman, Xiaoyi Lu, Nusrat Sharmin Islam, Dhabaleswar K. Panda |
| 2014 | PPoPP | Initial study of multi-endpoint runtime for MPI+OpenMP hybrid programming model on multi-core systems. | Miao Luo, Xiaoyi Lu, Khaled Hamidouche, Krishna Chaitanya Kandalla, Dhabaleswar K. Panda |
| 2013 | CCGRID | SR-IOV Support for Virtualization on InfiniBand Clusters: Early Experience. | Jithin Jose, Mingzhe Li, Xiaoyi Lu, Krishna Chaitanya Kandalla, Mark Daniel Arnold, Dhabaleswar K. Panda |
| 2013 | CCGRID | Efficient Intra-node Communication on Intel-MIC Clusters. | Sreeram Potluri, Akshay Venkatesh, Devendar Bureddy, Krishna Chaitanya Kandalla, Dhabaleswar K. Panda |
| 2013 | CLOUD | Does RDMA-based enhanced Hadoop MapReduce need a new performance model? | Md. Wasi-ur-Rahman, Xiaoyi Lu, Nusrat S. Islam, Dhabaleswar K. Panda |
| 2013 | CLUSTER | A scalable and portable approach to accelerate hybrid HPL on heterogeneous CPU-GPU clusters. | Rong Shi, Sreeram Potluri, Khaled Hamidouche, Xiaoyi Lu, Karen Tomko, Dhabaleswar K. Panda |
| 2013 | CLUSTER | Design of network topology aware scheduling services for large InfiniBand clusters. | Hari Subramoni, Devendar Bureddy, Krishna Chaitanya Kandalla, Karl W. Schulz, Bill Barth, Jonathan L. Perkins, Mark Daniel Arnold, Dhabaleswar K. Panda |
| 2013 | HOTI | Tutorials. | Dhabaleswar K. Panda, Xiaoyi Lu |
| 2013 | HOTI | Can Parallel Replication Benefit Hadoop Distributed File System for High Performance Interconnects? | Nusrat S. Islam, Xiaoyi Lu, Md. Wasi-ur-Rahman, Dhabaleswar K. Panda |
| 2013 | HOTI | Designing Optimized MPI Broadcast and Allreduce for Many Integrated Core (MIC) InfiniBand Clusters. | Krishna Chaitanya Kandalla, Akshay Venkatesh, Khaled Hamidouche, Sreeram Potluri, Devendar Bureddy, Dhabaleswar K. Panda |
| 2013 | HPDC | A 1 PB/s file system to checkpoint three million MPI tasks. | Raghunath Rajachandrasekar, Adam Moody, Kathryn Mohror, Dhabaleswar K. Panda |
| 2013 | ICPP | A Novel Functional Partitioning Approach to Design High-Performance MPI-3 Non-blocking Alltoallv Collective on Multi-core Systems. | Krishna Chaitanya Kandalla, Hari Subramoni, Karen Tomko, Dmitry Pekurovsky, Dhabaleswar K. Panda |
| 2013 | ICPP | High-Performance Design of Hadoop RPC with RDMA over InfiniBand. | Xiaoyi Lu, Nusrat S. Islam, Md. Wasi-ur-Rahman, Jithin Jose, Hari Subramoni, Hao Wang, Dhabaleswar K. Panda |
| 2013 | ICPP | Efficient Inter-node MPI Communication Using GPUDirect RDMA for InfiniBand Clusters with NVIDIA GPUs. | Sreeram Potluri, Khaled Hamidouche, Akshay Venkatesh, Devendar Bureddy, Dhabaleswar K. Panda |
| 2013 | ICS | MIC-RO: enabling efficient remote offload on heterogeneous many integrated core (MIC) clusters with InfiniBand. | Khaled Hamidouche, Sreeram Potluri, Hari Subramoni, Krishna Chaitanya Kandalla, Dhabaleswar K. Panda |
| 2013 | SC | MVAPICH-PRISM: a proxy-based communication framework using InfiniBand and SCIF for intel MIC clusters. | Sreeram Potluri, Devendar Bureddy, Khaled Hamidouche, Akshay Venkatesh, Krishna Chaitanya Kandalla, Hari Subramoni, Dhabaleswar K. Panda |
| 2012 | CCGRID | Scalable Memcached Design for InfiniBand Clusters Using Hybrid Transports. | Jithin Jose, Hari Subramoni, Krishna Chaitanya Kandalla, Md. Wasi-ur-Rahman, Hao Wang, Sundeep Narravula, Dhabaleswar K. Panda |
| 2012 | CLUSTER | Can Network-Offload Based Non-blocking Neighborhood MPI Collectives Improve Communication Overheads of Irregular Graph Algorithms? | Krishna Chaitanya Kandalla, Aydin Bulu, Hari Subramoni, Karen Tomko, Jrme Vienne, Leonid Oliker, Dhabaleswar K. Panda |
| 2012 | CLUSTER | Minimizing Network Contention in InfiniBand Clusters with a QoS-Aware Data-Staging Framework. | Raghunath Rajachandrasekar, Jai Jaswani, Hari Subramoni, Dhabaleswar K. Panda |
| 2012 | EuroPar | A Scalable InfiniBand Network Topology-Aware Performance Analysis Tool for MPI. | Hari Subramoni, Jrme Vienne, Dhabaleswar K. Panda |
| 2012 | HOTI | Performance Analysis and Evaluation of InfiniBand FDR and 40GigE RoCE on HPC and Cloud Computing Systems. | Jrme Vienne, Jitong Chen, Md. Wasi-ur-Rahman, Nusrat S. Islam, Hari Subramoni, Dhabaleswar K. Panda |
| 2012 | ICPP | Supporting Hybrid MPI and OpenSHMEM over InfiniBand: Design and Performance Evaluation. | Jithin Jose, Krishna Chaitanya Kandalla, Miao Luo, Dhabaleswar K. Panda |
| 2012 | ICPP | SSD-Assisted Hybrid Memory to Accelerate Memcached over High Performance Networks. | Xiangyong Ouyang, Nusrat S. Islam, Raghunath Rajachandrasekar, Jithin Jose, Miao Luo, Hao Wang, Dhabaleswar K. Panda |
| 2012 | ICS | Congestion avoidance on manycore high performance computing systems. | Miao Luo, Dhabaleswar K. Panda, Khaled Z. Ibrahim, Costin Iancu |
| 2012 | ISPASS | Understanding the communication characteristics in HBase: What are the fundamental bottlenecks? | Md. Wasi-ur-Rahman, Jian Huang, Jithin Jose, Xiangyong Ouyang, Hao Wang, Nusrat S. Islam, Hari Subramoni, Chet Murthy, Dhabaleswar K. Panda |
| 2012 | SC | High performance RDMA-based design of HDFS over InfiniBand. | Nusrat S. Islam, Md. Wasi-ur-Rahman, Jithin Jose, Raghunath Rajachandrasekar, Hao Wang, Hari Subramoni, Chet Murthy, Dhabaleswar K. Panda |
| 2012 | SC | Design of a scalable InfiniBand topology service to enable network-topology-aware placement of processes. | Hari Subramoni, Sreeram Potluri, Krishna Chaitanya Kandalla, Bill Barth, Jrme Vienne, Jeff Keasler, Karen A. Tomko, Karl W. Schulz, Adam Moody, Dhabaleswar K. Panda |
| 2011 | CCGRID | High Performance Pipelined Process Migration with RDMA. | Xiangyong Ouyang, Raghunath Rajachandrasekar, Xavier Besseron, Dhabaleswar K. Panda |
| 2011 | CLUSTER | Can a Decentralized Metadata Service Layer Benefit Parallel Filesystems? | Vilobh Meshram, Xavier Besseron, Xiangyong Ouyang, Raghunath Rajachandrasekar, Ravi Prakash, Dhabaleswar K. Panda |
| 2011 | CLUSTER | MPI Alltoall Personalized Exchange on GPGPU Clusters: Design Alternatives and Benefit. | Ashish Kumar Singh, Sreeram Potluri, Hao Wang, Krishna Chaitanya Kandalla, Sayantan Sur, Dhabaleswar K. Panda |
| 2011 | CLUSTER | Design and Evaluation of Network Topology-/Speed- Aware Broadcast Algorithms for InfiniBand Clusters. | Hari Subramoni, Krishna Chaitanya Kandalla, Jrme Vienne, Sayantan Sur, Bill Barth, Karen A. Tomko, Robert T. McLay, Karl W. Schulz, Dhabaleswar K. Panda |
| 2011 | CLUSTER | Optimized Non-contiguous MPI Datatype Communication for GPU Clusters: Design, Implementation and Evaluation with MVAPICH2. | Hao Wang, Sreeram Potluri, Miao Luo, Ashish Kumar Singh, Xiangyong Ouyang, Sayantan Sur, Dhabaleswar K. Panda |
| 2011 | EuroPar | INAM - A Scalable InfiniBand Network Analysis and Monitoring Tool. | N. Dandapanthula, Hari Subramoni, Jrme Vienne, Krishna Chaitanya Kandalla, Sayantan Sur, Dhabaleswar K. Panda, Ron Brightwell |
| 2011 | EuroPar | Can Checkpoint/Restart Mechanisms Benefit from Hierarchical Data Staging? | Raghunath Rajachandrasekar, Xiangyong Ouyang, Xavier Besseron, Vilobh Meshram, Dhabaleswar K. Panda |
| 2011 | HiPC | Multi-threaded UPC runtime with network endpoints: Design alternatives and evaluation on multi-core architectures. | Miao Luo, Jithin Jose, Sayantan Sur, Dhabaleswar K. Panda |
| 2011 | HOTI | Designing Non-blocking Broadcast with Collective Offload on InfiniBand Clusters: A Case Study with HPL. | Krishna Chaitanya Kandalla, Hari Subramoni, Jrme Vienne, S. Pai Raikar, Karen Tomko, Sayantan Sur, Dhabaleswar K. Panda |
| 2011 | HPCA | Beyond block I/O: Rethinking traditional storage primitives. | Xiangyong Ouyang, David W. Nellans, Robert Wipfel, David Flynn, Dhabaleswar K. Panda |
| 2011 | ICPP | Memcached Design on High Performance RDMA Capable Interconnects. | Jithin Jose, Hari Subramoni, Miao Luo, Minjia Zhang, Jian Huang, Md. Wasi-ur-Rahman, Nusrat S. Islam, Xiangyong Ouyang, Hao Wang, Sayantan Sur, Dhabaleswar K. Panda |
| 2011 | ICPP | CRFS: A Lightweight User-Level Filesystem for Generic Checkpoint/Restart. | Xiangyong Ouyang, Raghunath Rajachandrasekar, Xavier Besseron, Hao Wang, Jian Huang, Dhabaleswar K. Panda |
| 2010 | CCGRID | An MPI-Stream Hybrid Programming Model for Computational Clusters. | Emilio Pasquale Mancini, Gregory Marsh, Dhabaleswar K. Panda |
| 2010 | CCGRID | High Performance Data Transfer in Grid Environment Using GridFTP over InfiniBand. | Hari Subramoni, Ping Lai, Rajkumar Kettimuthu, Dhabaleswar K. Panda |
| 2010 | CLUSTER | RDMA-Based Job Migration Framework for MPI over InfiniBand. | Xiangyong Ouyang, Sonya Marcarelli, Raghunath Rajachandrasekar, Dhabaleswar K. Panda |
| 2010 | HOTI | Designing High-End Computing Systems with InfiniBand and High-Speed Ethernet. | Dhabaleswar K. Panda, Sayantan Sur, Pavan Balaji |
| 2010 | HOTI | Design and Evaluation of Generalized Collective Communication Primitives with Overlap Using ConnectX-2 Offload Engine. | Hari Subramoni, Krishna Chaitanya Kandalla, Sayantan Sur, Dhabaleswar K. Panda |
| 2010 | ICPP | Designing Power-Aware Collective Communication Algorithms for InfiniBand Clusters. | Krishna Chaitanya Kandalla, Emilio Pasquale Mancini, Sayantan Sur, Dhabaleswar K. Panda |
| 2010 | ICPP | Improving Application Performance and Predictability Using Multiple Virtual Lanes in Modern Multi-core InfiniBand Clusters. | Hari Subramoni, Ping Lai, Sayantan Sur, Dhabaleswar K. Panda |
| 2010 | ICS | Quantifying performance benefits of overlap using MPI-2 in a seismic modeling application. | Sreeram Potluri, Ping Lai, Karen A. Tomko, Sayantan Sur, Yifeng Cui, Mahidhar Tatineni, Karl W. Schulz, William L. Barth, Amitava Majumdar, Dhabaleswar K. Panda |
| 2010 | SC | Scalable Earthquake Simulation on Petascale Supercomputers. | Yifeng Cui, Kim B. Olsen, Thomas H. Jordan, Kwangyoon Lee, Jun Zhou, Patrick Small, Daniel Roten, Geoffrey Ely, Dhabaleswar K. Panda, Amit Chourasia, John M. Levesque, Steven M. Day, Philip Maechling |
| 2009 | CCGRID | Natively Supporting True One-Sided Communication in. | Gopalakrishnan Santhanaraman, Pavan Balaji, K. Gopalakrishnan, Rajeev Thakur, William Gropp, Dhabaleswar K. Panda |
| 2009 | CLUSTER | Reducing network contention with mixed workloads on modern multicore, clusters. | Matthew J. Koop, Miao Luo, Dhabaleswar K. Panda |
| 2009 | CLUSTER | Design alternatives for implementing fence synchronization in MPI-2 one-sided communication for InfiniBand clusters. | Gopalakrishnan Santhanaraman, Tejus Gangadharappa, Sundeep Narravula, Amith R. Mamidala, Dhabaleswar K. Panda |
| 2009 | CLUSTER | RDMA over Ethernet - A preliminary study. | Hari Subramoni, Ping Lai, Miao Luo, Dhabaleswar K. Panda |
| 2009 | CLUSTER | An efficient hardware-software approach to network fault tolerance with InfiniBand. | Abhinav Vishnu, Manojkumar Krishnan, Dhabaleswar K. Panda |
| 2009 | HiPC | Fast checkpointing by Write Aggregation with Dynamic Buffer and Interleaving on multicore architecture. | Xiangyong Ouyang, Karthik Gopalakrishnan, Tejus Gangadharappa, Dhabaleswar K. Panda |
| 2009 | HOTI | Tutorial: Infiniband and 10-Gigabit Ethernet for Dummies. | Dhabaleswar K. Panda, Matthew J. Koop, Pavan Balaji |
| 2009 | HOTI | Tutorial: Designing High-End Computing Systems with Infiniband and 10-Gigabit Ethernet. | Dhabaleswar K. Panda, Matthew J. Koop, Pavan Balaji |
| 2009 | HOTI | Designing Next Generation Clusters: Evaluation of InfiniBand DDR/QDR on Intel Computing Platforms. | Hari Subramoni, Matthew J. Koop, Dhabaleswar K. Panda |
| 2009 | ICPP | CIFTS: A Coordinated Infrastructure for Fault-Tolerant Systems. | Rinku Gupta, Peter H. Beckman, Byung-Hoon Park, Ewing L. Lusk, Paul Hargrove, Al Geist, Dhabaleswar K. Panda, Andrew Lumsdaine, Jack J. Dongarra |
| 2009 | ICPP | Designing Efficient FTP Mechanisms for High Performance Data-Transfer over InfiniBand. | Ping Lai, Hari Subramoni, Sundeep Narravula, Amith R. Mamidala, Dhabaleswar K. Panda |
| 2009 | ICPP | Accelerating Checkpoint Operation by Node-Level Write Aggregation on Multicore Systems. | Xiangyong Ouyang, Karthik Gopalakrishnan, Dhabaleswar K. Panda |
| 2008 | CCGRID | Advanced RDMA-Based Admission Control for Modern Data-Centers. | Ping Lai, Sundeep Narravula, Karthikeyan Vaidyanathan, Dhabaleswar K. Panda |
| 2008 | CCGRID | MPI Collectives on Modern Multicore Clusters: Performance Optimizations and Communication Characteristics. | Amith R. Mamidala, Rahul Kumar, Debraj De, Dhabaleswar K. Panda |
| 2008 | CCGRID | Optimized Distributed Data Sharing Substrate in Multi-core Commodity Clusters: A Comprehensive Study with Applications. | Karthikeyan Vaidyanathan, Ping Lai, Sundeep Narravula, Dhabaleswar K. Panda |
| 2008 | CLUSTER | Efficient one-copy MPI shared memory communication in Virtual Machines. | Wei Huang, Matthew J. Koop, Dhabaleswar K. Panda |
| 2008 | CLUSTER | Scalable MPI design over InfiniBand using eXtended Reliable Connection. | Matthew J. Koop, Jaidev K. Sridhar, Dhabaleswar K. Panda |
| 2008 | CLUSTER | Designing next generation clusters with InfiniBand and 10GE/iWARP: Opportunities and challenges. | Dhabaleswar K. Panda |
| 2008 | HiPC | Sockets Direct Protocol for Hybrid Network Stacks: A Case Study with iWARP over 10G Ethernet. | Pavan Balaji, Sitha Bhagvat, Rajeev Thakur, Dhabaleswar K. Panda |
| 2008 | HiPC | Designing a High-Performance Clustered NAS: A Case Study with pNFS over RDMA on InfiniBand. | Ranjit Noronha, Xiangyong Ouyang, Dhabaleswar K. Panda |
| 2008 | HiPC | ScELA: Scalable and Extensible Launching Architecture for Clusters. | Jaidev K. Sridhar, Matthew J. Koop, Jonathan L. Perkins, Dhabaleswar K. Panda |
| 2008 | HOTI | Performance Analysis and Evaluation of PCIe 2.0 and Quad-Data Rate InfiniBand. | Matthew J. Koop, Wei Huang, Karthik Gopalakrishnan, Dhabaleswar K. Panda |
| 2008 | ICPP | Designing an Efficient Kernel-Level and User-Level Hybrid Approach for MPI Intra-Node Communication on Multi-Core Systems. | Lei Chai, Ping Lai, Hyun-Wook Jin, Dhabaleswar K. Panda |
| 2008 | ICPP | Performance of HPC Middleware over InfiniBand WAN. | Sundeep Narravula, Hari Subramoni, Ping Lai, Ranjit Noronha, Dhabaleswar K. Panda |
| 2008 | ICPP | IMCa: A High Performance Caching Front-End for GlusterFS on InfiniBand. | Ranjit Noronha, Dhabaleswar K. Panda |
| 2008 | ICS | Can software reliability outperform hardware reliability on high performance interconnects?: a case study with MPI over infiniband. | Matthew J. Koop, Rahul Kumar, Dhabaleswar K. Panda |
| 2007 | CCGRID | Understanding the Impact of Multi-Core Architecture in Cluster Computing: A Case Study with Intel Dual-Core System. | Lei Chai, Qi Gao, Dhabaleswar K. Panda |
| 2007 | CCGRID | Reducing Connection Memory Requirements of MPI for InfiniBand Clusters: A Message Coalescing Approach. | Matthew J. Koop, Terry R. Jones, Dhabaleswar K. Panda |
| 2007 | CCGRID | High Performance Distributed Lock Management Services using Network-based Remote Atomic Operations. | Sundeep Narravula, A. Marnidala, Abhinav Vishnu, Karthikeyan Vaidyanathan, Dhabaleswar K. Panda |
| 2007 | CCGRID | Hot-Spot Avoidance With Multi-Pathing Over InfiniBand: An MPI Perspective. | Abhinav Vishnu, Matthew J. Koop, Adam Moody, Amith R. Mamidala, Sundeep Narravula, Dhabaleswar K. Panda |
| 2007 | CLUSTER | High performance virtual machine migration with RDMA over modern interconnects. | Wei Huang, Qi Gao, Jiuxing Liu, Dhabaleswar K. Panda |
| 2007 | CLUSTER | Lightweight kernel-level primitives for high-performance MPI intra-node communication over multi-core systems. | Hyun-Wook Jin, Sayantan Sur, Lei Chai, Dhabaleswar K. Panda |
| 2007 | CLUSTER | Zero-copy protocol for MPI using infiniband unreliable datagram. | Matthew J. Koop, Sayantan Sur, Dhabaleswar K. Panda |
| 2007 | CLUSTER | Designing high-end computing systems with InfiniBand and10-Gigabit Ethernet iWARP. | Dhabaleswar K. Panda, Pavan Balaji |
| 2007 | CLUSTER | Efficient asynchronous memory copy operations on multi-core systems and I/OAT. | Karthikeyan Vaidyanathan, Lei Chai, Wei Huang, Dhabaleswar K. Panda |
| 2007 | HOTI | Performance Analysis and Evaluation of Mellanox ConnectX InfiniBand Architecture with Multi-Core Platforms. | Sayantan Sur, Matthew J. Koop, Lei Chai, Dhabaleswar K. Panda |
| 2007 | ICPP | Advanced Flow-control Mechanisms for the Sockets Direct Protocol over InfiniBand. | Pavan Balaji, Sitha Bhagvat, Dhabaleswar K. Panda, Rajeev Thakur, William Gropp |
| 2007 | ICPP | Group-based Coordinated Checkpointing for MPI: A Case Study on InfiniBand. | Qi Gao, Wei Huang, Matthew J. Koop, Dhabaleswar K. Panda |
| 2007 | ICPP | High Performance MPI over iWARP: Early Experiences. | Sundeep Narravula, Amith R. Mamidala, Abhinav Vishnu, Gopalakrishnan Santhanaraman, Dhabaleswar K. Panda |
| 2007 | ICPP | Designing NFS with RDMA for Security, Performance and Scalability. | Ranjit Noronha, Lei Chai, Thomas Talpey, Dhabaleswar K. Panda |
| 2007 | ICS | High performance MPI design using unreliable datagram for ultra-scale InfiniBand clusters. | Matthew J. Koop, Sayantan Sur, Qi Gao, Dhabaleswar K. Panda |
| 2007 | ISPASS | Benefits of I/O Acceleration Technology (I/OAT) in Clusters. | Karthikeyan Vaidyanathan, Dhabaleswar K. Panda |
| 2007 | PPoPP | On using connection-oriented vs. connection-less transport for performance and scalability of collective and one-sided operations: trade-offs and impact. | Amith R. Mamidala, Sundeep Narravula, Abhinav Vishnu, Gopalakrishnan Santhanaraman, Dhabaleswar K. Panda |
| 2007 | SC | Analyzing the impact of supporting out-of-order communication on in-order performance with iWARP. | Pavan Balaji, Wu-chun Feng, Sitha Bhagvat, Dhabaleswar K. Panda, Rajeev Thakur, William Gropp |
| 2007 | SC | pNFS/PVFS2 over InfiniBand: early experiences. | Lei Chai, Xiangyong Ouyang, Ranjit Noronha, Dhabaleswar K. Panda |
| 2007 | SC | DMTracker: finding bugs in large-scale parallel programs by detecting anomaly in data movements. | Qi Gao, Feng Qin, Dhabaleswar K. Panda |
| 2007 | SC | Virtual machine aware communication libraries for high performance computing. | Wei Huang, Matthew J. Koop, Qi Gao, Dhabaleswar K. Panda |
| 2006 | CCGRID | MPI over uDAPL: Can High Performance and Portability Exist Across Architectures?. | Lei Chai, Ranjit Noronha, Dhabaleswar K. Panda |
| 2006 | CCGRID | Design of High Performance MVAPICH2: MPI2 over InfiniBand. | Wei Huang, Gopalakrishnan Santhanaraman, Hyun-Wook Jin, Qi Gao, Dhabaleswar K. Panda |
| 2006 | CCGRID | Designing Efficient Cooperative Caching Schemes for Multi-Tier Data-Centers over RDMA-enabled Networks. | Sundeep Narravula, Hyun-Wook Jin, Karthikeyan Vaidyanathan, Dhabaleswar K. Panda |
| 2006 | CLUSTER | Designing High Performance and Scalable MPI Intra-node Communication Support for Clusters. | Lei Chai, Albert Hartono, Dhabaleswar K. Panda |
| 2006 | CLUSTER | Exploiting RDMA operations for Providing Efficient Fine-Grained Resource Monitoring in Cluster-based Servers. | Karthikeyan Vaidyanathan, Hyun-Wook Jin, Dhabaleswar K. Panda |
| 2006 | HiPC | DDSS: A Low-Overhead Distributed Data Sharing Substrate for Cluster-Based Data-Centers over Modern Interconnects. | Karthikeyan Vaidyanathan, Sundeep Narravula, Dhabaleswar K. Panda |
| 2006 | HOTI | Memory Scalability Evaluation of the Next-Generation Intel Bensley Platform with InfiniBand. | Matthew J. Koop, Wei Huang, Abhinav Vishnu, Dhabaleswar K. Panda |
| 2006 | ICCCN | NemC: A Network Emulator for Cluster-of-Clusters. | Hyun-Wook Jin, Sundeep Narravula, Karthikeyan Vaidyanathan, Dhabaleswar K. Panda |
| 2006 | ICPP | Application-Transparent Checkpoint/Restart for MPI Programs over InfiniBand. | Qi Gao, Weikuan Yu, Wei Huang, Dhabaleswar K. Panda |
| 2006 | ICPP | High Performance Block I/O for Global File System (GFS) with InfiniBand RDMA. | Shuang Liang, Weikuan Yu, Dhabaleswar K. Panda |
| 2006 | ICS | A case for high performance computing with virtual machines. | Wei Huang, Jiuxing Liu, Blent Abali, Dhabaleswar K. Panda |
| 2006 | PPoPP | RDMA read based rendezvous protocol for MPI over InfiniBand: design alternatives and benefits. | Sayantan Sur, Hyun-Wook Jin, Lei Chai, Dhabaleswar K. Panda |
| 2006 | SC | Panel: Data intensive computing. | Leslie S. Perkins, Phil Andrews, Dhabaleswar K. Panda, Dave Morton, Ron Bonica, Nick Henry Werstiuk, Randy Kreiser |
| 2005 | CCGRID | Architecture for caching responses with multiple dynamic dependencies in multi-tier data-centers over InfiniBand. | Sundeep Narravula, Pavan Balaji, Karthikeyan Vaidyanathan, Hyun-Wook Jin, Dhabaleswar K. Panda |
| 2005 | CCGRID | Can high performance software DSM systems designed with InfiniBand features benefit from PCI-Express? | Ranjit Noronha, Dhabaleswar K. Panda |
| 2005 | CLUSTER | Head-to-TOE Evaluation of High-Performance Sockets over Protocol Offload Engines. | Pavan Balaji, Wu-chun Feng, Qi Gao, Ranjit Noronha, Weikuan Yu, Dhabaleswar K. Panda |
| 2005 | CLUSTER | Supporting iWARP Compatibility and Features for Regular Network Adapters. | Pavan Balaji, Hyun-Wook Jin, Karthikeyan Vaidyanathan, Dhabaleswar K. Panda |
| 2005 | CLUSTER | Swapping to Remote Memory over InfiniBand: An Approach using a High Performance Network Block Device. | Shuang Liang, Ranjit Noronha, Dhabaleswar K. Panda |
| 2005 | EuroPar | Performance Evaluation of MM5 on Clusters with Modern Interconnects: Scalability and Impact. | Ranjit Noronha, Dhabaleswar K. Panda |
| 2005 | HiPC | High Performance RDMA Based All-to-All Broadcast for InfiniBand Clusters. | Sayantan Sur, Uday Bondhugula, Amith R. Mamidala, Hyun-Wook Jin, Dhabaleswar K. Panda |
| 2005 | HiPC | Supporting MPI-2 One Sided Communication on Multi-rail InfiniBand Clusters: Design Challenges and Performance Benefits. | Abhinav Vishnu, Gopalakrishnan Santhanaraman, Wei Huang, Hyun-Wook Jin, Dhabaleswar K. Panda |
| 2005 | HOTI | Performance Characterization of a 10-Gigabit Ethernet TOE. | Wu-chun Feng, Pavan Balaji, Christopher Baron, Laxmi N. Bhuyan, Dhabaleswar K. Panda |
| 2005 | HOTI | Can Memory-Less Network Adapters Benefit Next-Generation InfiniBand Systems?. | Sayantan Sur, Abhinav Vishnu, Hyun-Wook Jin, Wei Huang, Dhabaleswar K. Panda |
| 2005 | ICPP | LiMIC: Support for High-Performance MPI Intra-node Communication on Linux Cluster. | Hyun-Wook Jin, Sayantan Sur, Lei Chai, Dhabaleswar K. Panda |
| 2005 | ICS | High performance support of parallel virtual file system (PVFS2) over Quadrics. | Weikuan Yu, Shuang Liang, Dhabaleswar K. Panda |
| 2005 | ISPASS | On the provision of prioritization and soft qos in dynamically reconfigurable shared data-centers over infiniband. | Pavan Balaji, Sundeep Narravula, Karthikeyan Vaidyanathan, Hyun-Wook Jin, Dhabaleswar K. Panda |
| 2004 | CCGRID | High performance MPI-2 one-sided communication over InfiniBand. | Weihang Jiang, Jiuxing Liu, Hyun-Wook Jin, Dhabaleswar K. Panda, William Gropp, Rajeev Thakur |
| 2004 | CCGRID | Designing high performance DSM systems using InfiniBand features. | Ranjit Noronha, Dhabaleswar K. Panda |
| 2004 | CCGRID | Unifier: unifying cache management and communication buffer management for PVFS over InfiniBand. | Jiesheng Wu, Pete Wyckoff, Dhabaleswar K. Panda, Robert B. Ross |
| 2004 | CLUSTER | Towards provision of quality of service guarantees in job scheduling. | Mohammad Islam, Pavan Balaji, P. Sadayappan, Dhabaleswar K. Panda |
| 2004 | CLUSTER | Efficient Barrier and Allreduce on Infiniband clusters using multicast and adaptive algorithms. | Amith R. Mamidala, Jiuxing Liu, Dhabaleswar K. Panda |
| 2004 | CLUSTER | State of InfiniBand in designing HPC clusters, storage/file systems, and datacenters [datacenters read as data centers]. | Dhabaleswar K. Panda |
| 2004 | CLUSTER | NIC-based offload of dynamic user-defined modules for Myrinet clusters. | Adam Wagner, Hyun-Wook Jin, Dhabaleswar K. Panda, Rolf Riesen |
| 2004 | CLUSTER | Scalable, high-performance NIC-based all-to-all broadcast over Myrinet/GM. | Weikuan Yu, Dhabaleswar K. Panda, Darius Buntinas |
| 2004 | HiPC | Fast and Scalable Startup of MPI Programs in InfiniBand Clusters. | Weikuan Yu, Jiesheng Wu, Dhabaleswar K. Panda |
| 2004 | HOTI | Performance evaluation of InfiniBand with PCI Express. | Jiuxing Liu, Amith R. Mamidala, Abhinav Vishnu, Dhabaleswar K. Panda |
| 2004 | ICPP | Efficient and Scalable All-to-All Personalized Exchange for InfiniBand-Based Clusters. | Sayantan Sur, Hyun-Wook Jin, Dhabaleswar K. Panda |
| 2004 | ISPASS | Sockets Direct Protocol over InfiniBand in clusters: is it beneficial? | Pavan Balaji, Sundeep Narravula, Karthikeyan Vaidyanathan, Savitha Krishnamoorthy, Jiesheng Wu, Dhabaleswar K. Panda |
| 2004 | SC | Building Multirail InfiniBand Clusters: MPI-Level Design and Performance Evaluation. | Jiuxing Liu, Abhinav Vishnu, Dhabaleswar K. Panda |
| 2003 | CCGRID | Application-Bypas Broadcast in MPICH over GM. | Darius Buntinas, Dhabaleswar K. Panda, Ron Brightwell |
| 2003 | CLUSTER | Optimizing Mechanisms for Latency Tolerance in Remote Memory Access Communication on Clusters. | Jarek Nieplocha, Vinod Tipparaju, Manojkumar Krishnan, Gopalakrishnan Santhanaraman, Dhabaleswar K. Panda |
| 2003 | CLUSTER | Designing Next Generation Clusters with Infiniband: Opportunities and Challenges. | Dhabaleswar K. Panda |
| 2003 | CLUSTER | Application-Bypass Reduction for Large-Scale Clusters. | Adam Wagner, Darius Buntinas, Dhabaleswar K. Panda, Ron Brightwell |
| 2003 | CLUSTER | Supporting Efficient Noncontiguous Access in PVFS over InfiniBand. | Jiesheng Wu, Pete Wyckoff, Dhabaleswar K. Panda |
| 2003 | HiPC | Exploiting Non-blocking Remote Memory Access Communication in Scientific Benchmarks. | Vinod Tipparaju, Manojkumar Krishnan, Jarek Nieplocha, Gopalakrishnan Santhanaraman, Dhabaleswar K. Panda |
| 2003 | HOTI | Micro-benchmark level performance comparison of high-speed cluster interconnects. | Jiuxing Liu, Balasubramanian Chandrasekaran, Weikuan Yu, Jiesheng Wu, Darius Buntinas, Sushmitha P. Kini, Pete Wyckoff, Dhabaleswar K. Panda |
| 2003 | HPDC | Impact of High Performance Sockets on Data Intensive Applications. | Pavan Balaji, Jiesheng Wu, Tahsin M. Kur, mit V. atalyrek, Dhabaleswar K. Panda, Joel H. Saltz |
| 2003 | HPDC | QoS-Aware Middleware for Cluster-Based Servers to support Interactive and Resource-Adaptive Applications. | S. Senapathi, B. Chandrasekaran, Don Stredney, Han-Wei Shen, Dhabaleswar K. Panda |
| 2003 | ICPP | PVFS over InfiniBand: Design and Performance Evaluation. | Jiesheng Wu, Pete Wyckoff, Dhabaleswar K. Panda |
| 2003 | ICPP | High Performance and Reliable NIC-Based Multicast over Myrinet/GM-2. | Weikuan Yu, Darius Buntinas, Dhabaleswar K. Panda |
| 2003 | ICS | High performance RDMA-based MPI implementation over InfiniBand. | Jiuxing Liu, Jiesheng Wu, Sushmitha P. Kini, Pete Wyckoff, Dhabaleswar K. Panda |
| 2003 | JSSPP | QoPS: A QoS Based Scheme for Parallel Job Scheduling. | Mohammad Islam, Pavan Balaji, P. Sadayappan, Dhabaleswar K. Panda |
| 2003 | KDD | Towards NIC-based intrusion detection. | Matthew Eric Otey, Srinivasan Parthasarathy, Amol Ghoting, G. Li, Sundeep Narravula, Dhabaleswar K. Panda |
| 2003 | SC | Performance Comparison of MPI Implementations over InfiniBand, Myrinet and Quadrics. | Jiuxing Liu, B. Chandrasekaran, Jiesheng Wu, Weihang Jiang, Sushmitha P. Kini, Weikuan Yu, Darius Buntinas, Pete Wyckoff, Dhabaleswar K. Panda |
| 2003 | SC | Scalable NIC-based Reduction on Large-scale Clusters. | Adam Moody, Juan Fernndez, Fabrizio Petrini, Dhabaleswar K. Panda |
| 2002 | CLUSTER | High Performance User Level Sockets over Gigabit Ethernet. | Pavan Balaji, Piyush Shivam, Pete Wyckoff, Dhabaleswar K. Panda |
| 2002 | CLUSTER | Efficient Barrier Using Remote Memory Operations on VIA-Based Clusters. | Rinku Gupta, Vinod Tipparaju, Jarek Nieplocha, Dhabaleswar K. Panda |
| 2002 | CLUSTER | Impact of On-Demand Connection Management in MPI over VIA. | Jiesheng Wu, Jiuxing Liu, Pete Wyckoff, Dhabaleswar K. Panda |
| 2002 | HOTI | Tutorial 2: InfiniBand Architecture and Where it is Headed. | Dhabaleswar K. Panda |
| 2002 | ICDCS | A Reliable Multicast Algorithm for Mobile Ad Hoc Networks. | Thiagaraja Gopalsamy, Mukesh Singhal, Dhabaleswar K. Panda, P. Sadayappan |
| 2002 | LCN | Active Network Interface: Opportunities and Challenges. | Dhabaleswar K. Panda |
| 2001 | ICPP | Implementing TreadMarksover VIA on Myrinet and Gigabit Ethernet: Challenges, Design Experience, and Performance Evaluation. | Mohammad Banikazemi, Jiuxing Liu, Dhabaleswar K. Panda, P. Sadayappan |
| 2001 | ICPP | NIC-Based Rate Control for Proportional Bandwidth Allocation in Myrinet Clusters. | Abhishek Gulati, Dhabaleswar K. Panda, P. Sadayappan, Pete Wyckoff |
| 2001 | SC | EMP: zero-copy OS-bypass NIC-driven gigabit ethernet message passing. | Piyush Shivam, Pete Wyckoff, Dhabaleswar K. Panda |
| 2000 | HiPC | Can Scatter Communication Take Advantage of Multidestination Message Passing? | Mohammad Banikazemi, Dhabaleswar K. Panda |
| 2000 | HiPC | Characterization and enhancement of Static Mapping Heuristics for Heterogeneous Systems. | Praveen Holenarsipur, Vladimir Yarmolenko, Jos Duato, Dhabaleswar K. Panda, P. Sadayappan |
| 1999 | HCW | Communication Modeling of Heterogeneous Networks of Workstations for Performance Characterization of Collective Operations. | Mohammad Banikazemi, Jayanthi Sampathkumar, Sandeep Prabhu, Dhabaleswar K. Panda, P. Sadayappan |
| 1998 | ICPP | Efficient Collective Communication on Heterogeneous Networks of Workstations. | Mohammad Banikazemi, Vijay Moorthy, Dhabaleswar K. Panda |
| 1998 | ICPP | Impact of Adaptivity on the Behaviour of Networks of Workstations under Bursty Traffic. | Federico Silla, Manuel P. Malumbres, Jos Duato, Donglai Dai, Dhabaleswar K. Panda |
| 1998 | ICPP | Where to Provide Support for Efficient Multicasting in Irregular Networks: Network Interface or Switch? | Rajeev Sivaram, Ram Kesavan, Dhabaleswar K. Panda, Craig B. Stunkel |
| 1997 | HiPC | Prioritized demand multiplexing (PDM): a low-latency virtual channel flow control framework for prioritized traffic. | Abdel-Halim Smai, Dhabaleswar K. Panda, Lars-Erik Thorelli |
| 1997 | HPCA | Multicast on Irregular Switch-Based Networks with Wormhole Routing. | Ram Kesavan, Kiran Bondalapati, Dhabaleswar K. Panda |
| 1997 | ICPP | How Much Does Network Contention Affect Distributed Shared Memory Performance? | Donglai Dai, Dhabaleswar K. Panda |
| 1997 | ICPP | Optimal Multicast with Packetization and Network Interface Support. | Ram Kesavan, Dhabaleswar K. Panda |
| 1997 | ISCA | Implementing Multidestination Worms in Switch-Based Parallel Systems: Architectural Alternatives and their Impact. | Craig B. Stunkel, Rajeev Sivaram, Dhabaleswar K. Panda |
| 1997 | WSC | Simulation of Modern Parallel Systems: A CSIM-based Approach. | Dhabaleswar K. Panda, Debashis Basak, Donglai Dai, Ram Kesavan, Rajeev Sivaram, Mohammad Banikazemi, Vijay Moorthy |
| 1996 | ICPP | Designing Processor-Cluster Based Systems: Interplay Between Organizations and Broadcasting Algorithms. | Debashis Basak, Dhabaleswar K. Panda |
| 1996 | ICPP | Reducing Cache Invalidation Overheads in Wormhole Routed DSMs Using Multidestination Message Passing. | Donglai Dai, Dhabaleswar K. Panda |
| 1996 | ICPP | Minimizing Node Contention in Multiple Multicast on Wormhole k-ary N-Cube Networks. | Ram Kesavan, Dhabaleswar K. Panda |
| 1996 | ICS | Hybrid Algorithms for Complete Exchange in 2D Meshes. | N. S. Sundar, Doddaballapur Narasimha-Murthy Jayasimha, Dhabaleswar K. Panda, P. Sadayappan |
| 1995 | HPCA | Fast Barrier Synchronization in Wormhole k-ary n-cube Networks with Multidestination Worms. | Dhabaleswar K. Panda |
| 1994 | ICPP | Designing Large Hierarchical Multiprocessor Systems under Processor, Interconnection, and Packaging Advancements. | Debashis Basak, Dhabaleswar K. Panda |
| 1991 | ICPP | Message Vectorization for Converting Multicomputer Programs to Shared-Memory Multiprocessors. | Dhabaleswar K. Panda, Kai Hwang |
| 1990 | ICPP | Algorithm-Driven Simulation and Performance Projection of a RISC-based Orthogonal Multiprocessor. | Sharad Mehrotra, Chien-Ming Cheng, Kai Hwang, Michel Dubois, Dhabaleswar K. Panda |
| 1990 | ICS | OMP: a RISC-based multiprocessor using orthogonal-access memories and multiple spanning buses. | Kai Hwang, Michel Dubois, Dhabaleswar K. Panda, S. Rao, Shisheng Shang, Aydin resin, W. Mao, H. Nair, M. Lytwyn, F. Hsieh, J. Liu, Sharad Mehrotra, Chien-Ming Cheng |
| 1989 | ARITH | Optical arithmetic using high-radix symbolic substitution rules. | Kai Hwang, Dhabaleswar K. Panda |