Skip to content

Dhabaleswar K. Panda

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

316

Venues

28

Active years

1989–2026

Best venue rank

A*

Where they publish

Papers

Showing the 300 most recent indexed papers.

YearVenueTitleAuthors
2026HPDCHAT-MPI: Hierarchical Auto Tuning of MPI Inter-Node Communication on InfiniBand Clusters.Sungjae Lee, Sam Tilford, Dhabaleswar K. Panda
2026WACVSupporting Ultra-High-Resolution Digital Agriculture Tasks with Fully Synthetic Curriculum Learning.Jacob Hatef, Quentin Anthony, Nawras Alnaasan, Dhabaleswar K. Panda
2025CLUSTERTowards Dynamic Message Passing Protocols for Stencil-Based Communication Patterns.Kaushik Kandadi Suresh, Bharath Ramesh, Goutham Kalikrishna Reddy Kuncham, Hari Subramoni, Dhabaleswar K. Panda
2025HiPCPerformance Characterization of Data Transfer and Allocation Strategies on AMD MI300A APUs: Early Experiences.Goutham Kalikrishna Reddy Kuncham, Siyuan Zhang, Bharath Ramesh, Kaushik Kandadi Suresh, Dhabaleswar K. Panda
2025HiPCEnhanced MPI Intra-Node Communication Framework: A Hybrid Approach with Cooperative DMA Channel-Based Data Transfer.Shulei Xu, Tu Tran, Dhabaleswar K. Panda
2025HOTICharacterizing Communication Patterns in Distributed Large Language Model Inference.Lang Xu, Kaushik Kandadi Suresh, Quentin Anthony, Nawras Alnaasan, Dhabaleswar K. Panda
2025ICPPDesign and Optimization of GPU-Aware MPI Allreduce Using Direct Sendrecv Communication.Chen-Chun Chen, Jinghan Yao, Hari Subramoni, Dhabaleswar K. Panda
2025SCOpenSHMEM MLIR: A Dialect for Compile-Time Optimization of One-Sided Communications.Michael Beebe, Benjamin Michalowicz, Andrew McNamara, Yash Kumar, Dhabaleswar K. Panda, Yong Chen, Wendy Poole, Steve Poole
2025SCA Streaming Collectives Interface Targeting Dataflow Acceleration and HPC Workloads.Nicholas Contini, Jake Queiser, Bharath Ramesh, Hari Subramoni, Dhabaleswar K. Panda
2025SCMPI Communication Performance on AMD MI300A: Microbenchmarks and Applications.Goutham Kalikrishna Reddy Kuncham, Siyuan Zhang, Shoaib Mohammad, Chen-Chun Chen, Dhabaleswar K. Panda
2024HiPCHyperSack: Distributed Hyperparameter Optimization for Deep Learning using Resource-Aware Scheduling on Heterogeneous GPU Systems.Nawras Alnaasan, Bharath Ramesh, Jinghan Yao, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda
2024HiPCDesign and Implementation of Kernel-based MPI Reduction Operations for Intel GPU s.Chen-Chun Chen, Goutham Kalikrishna Reddy Kuncham, Hari Subramoni, Dhabaleswar K. Panda
2024HiPCEffective and Efficient Offloading Designs for One-Sided Communication to SmartNICs.Benjamin Michalowicz, Kaushik Kandadi Suresh, Hari Subramoni, Mustafa Abduljabbar, Dhabaleswar K. Panda, Steve Poole
2024HiPCUsing BlueField-3 SmartNICs to Offload Vector Operations in Krylov Subspace Methods.Kaushik Kandadi Suresh, Benjamin Michalowicz, Nick Contini, Bharath Ramesh, Mustafa Abduljabbar, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda
2024HiPCScaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning.Lang Xu, Quentin Anthony, Jacob Hatef, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda
2024HOTICharacterizing Communication in Distributed Parameter-Efficient Fine-Tuning for Large Language Models.Nawras Alnaasan, Horng-Ruey Huang, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda
2024HOTIDemystifying the Communication Characteristics for Distributed Transformer Models.Quentin Anthony, Benjamin Michalowicz, Jacob Hatef, Lang Xu, Mustafa Abdul Jabbar, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda
2024HOTIOHIO: Improving RDMA Network Scalability in MPI_Alltoall Through Optimized Hierarchical and Intra/Inter-Node Communication Overlap Design.Tu Tran, Goutham Kalikrishna Reddy Kuncham, Bharath Ramesh, Shulei Xu, Hari Subramoni, Mustafa Abduljabbar, Dhabaleswar K. Panda
2024ICPPThe Case for Co-Designing Model Architectures with Hardware.Quentin Anthony, Jacob Hatef, Deepak Narayanan, Stella Biderman, Stas Bekman, Junqi Yin, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda
2023CCGRIDScaMP: Scalable Meta-Parallelism for Deep Learning Search.Quentin Anthony, Lang Xu, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda
2023CCGRIDScaMP: Scalable Meta-Parallelism for Deep Learning Search.Quentin Anthony, Lang Xu, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda
2023CCGRIDImplementing and Optimizing a GPU-aware MPI Library for Intel GPUs: Early Experiences.Chen-Chun Chen, Kawthar Shafie Khorassani, Goutham Kalikrishna Reddy Kuncham, Rahul Vaidya, Mustafa Abduljabbar, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda
2023ICFECPerformance Characterization of Using Quantization for DNN Inference on Edge Devices.Hyunho Ahn, Tian Chen, Nawras Alnaasan, Aamir Shafi, Mustafa Abduljabbar, Hari Subramoni, Dhabaleswar K. Panda
2023ICSEnabling Reconfigurable HPC through MPI-based Inter-FPGA Communication.Nicholas Contini, Bharath Ramesh, Kaushik Kandadi Suresh, Tu Tran, Benjamin Michalowicz, Mustafa Abduljabbar, Hari Subramoni, Dhabaleswar K. Panda
2023SCMPI-xCCL: A Portable MPI Library over Collective Communication Libraries for Various Accelerators.Chen-Chun Chen, Kawthar Shafie Khorassani, Pouya Kousha, Qinghua Zhou, Jinghan Yao, Hari Subramoni, Dhabaleswar K. Panda
2023SCDemocratizing HPC Access and Use with Knowledge Graphs.Pouya Kousha, Vivekananda Sathu, Matthew Lieber, Hari Subramoni, Dhabaleswar K. Panda
2022CLUSTERSpark Meets MPI: Towards High-Performance Communication Framework for Spark using MPI.Kinan Al-Attar, Aamir Shafi, Mustafa Abduljabbar, Hari Subramoni, Dhabaleswar K. Panda
2022HiPCAccDP: Accelerated Data-Parallel Distributed DNN Training for Modern GPU-Based HPC Clusters.Nawras Alnaasan, Arpan Jain, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda
2022HiPCDesigning Efficient Pipelined Communication Schemes using Compression in MPI Libraries.Bharath Ramesh, Qinghua Zhou, Aamir Shafi, Mustafa Abduljabbar, Hari Subramoni, Dhabaleswar K. Panda
2022HiPCEfficient Personalized and Non-Personalized Alltoall Communication for Modern Multi-HCA GPU-Based Clusters.Kaushik Kandadi Suresh, Akshay Paniraja Guptha, Benjamin Michalowicz, Bharath Ramesh, Mustafa Abduljabbar, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda
2022HiPCAccelerating Broadcast Communication with GPU Compression for Deep Learning Workloads.Qinghua Zhou, Quentin Anthony, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda
2022HOTINetwork Assisted Non-Contiguous Transfers for GPU-Aware MPI Libraries.Kaushik Kandadi Suresh, Kawthar Shafie Khorassani, Chen-Chun Chen, Bharath Ramesh, Mustafa Abduljabbar, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda
2021CCGRIDAdaptive and Hierarchical Large Message All-to-all Communication Algorithms for Large-scale Dense GPU Systems.Kawthar Shafie Khorassani, Ching-Hsiang Chu, Quentin G. Anthony, Hari Subramoni, Dhabaleswar K. Panda
2021HiPCTowards Architecture-aware Hierarchical Communication Trees on Modern HPC Systems.Bharath Ramesh, Jahanzeb Maqbool Hashmi, Shulei Xu, Aamir Shafi, Seyedeh Mahdieh Ghazimirsaeed, Mohammadreza Bayatpour, Hari Subramoni, Dhabaleswar K. Panda
2021HiPCDistMILE: A Distributed Multi-Level Framework for Scalable Graph Embedding.Yuntian He, Saket Gurukar, Pouya Kousha, Hari Subramoni, Dhabaleswar K. Panda, Srinivasan Parthasarathy
2021HiPCLarge-Message Nonblocking MPI_Iallgather and MPI Ibcast Offload via BlueField-2 DPU.Nick Sarkauskas, Mohammadreza Bayatpour, Tu Tran, Bharath Ramesh, Hari Subramoni, Dhabaleswar K. Panda
2021HiPCLayout-aware Hardware-assisted Designs for Derived Data Types in MPI.Kaushik Kandadi Suresh, Bharath Ramesh, Chen-Chun Chen, Seyedeh Mahdieh Ghazimirsaeed, Mohammadreza Bayatpour, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda
2021HOTIAccelerating CPU-based Distributed DNN Training on Modern HPC Clusters using BlueField-2 DPUs.Arpan Jain, Nawras Alnaasan, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda
2020CCGRIDDesign and Characterization of InfiniBand Hardware Tag Matching in MPI.Mohammadreza Bayatpour, Seyedeh Mahdieh Ghazimirsaeed, Shulei Xu, Hari Subramoni, Dhabaleswar K. Panda
2020CLUSTERDynamic Kernel Fusion for Bulk Non-contiguous Data Transfer on GPU Clusters.Ching-Hsiang Chu, Kawthar Shafie Khorassani, Qinghua Zhou, Hari Subramoni, Dhabaleswar K. Panda
2020HiPCBlink: Towards Efficient RDMA-based Communication Coroutines for Parallel Python Applications.Aamir Shafi, Jahanzeb Maqbool Hashmi, Hari Subramoni, Dhabaleswar K. Panda
2020SCScalable MPI Collectives using SHARP: Large Scale Performance Evaluation on the TACC Frontera System.Bharath Ramesh, Kaushik Kandadi Suresh, Nick Sarkauskas, Mohammadreza Bayatpour, Jahanzeb Maqbool Hashmi, Hari Subramoni, Dhabaleswar K. Panda
2020SCGEMS: GPU-enabled memory-aware model-parallelism system for distributed DNN training.Arpan Jain, Ammar Ahmad Awan, Asmaa M. Aljuhani, Jahanzeb Maqbool Hashmi, Quentin G. Anthony, Hari Subramoni, Dhabaleswar K. Panda, Raghu Machiraju, Anil Parwani
2020SCExploring Hybrid MPI+Kokkos Tasks Programming Model.Samuel Khuvis, Karen Tomko, Jahanzeb Maqbool Hashmi, Dhabaleswar K. Panda
2019ASPLOSCharacterizing CUDA Unified Memory (UM)-Aware MPI Designs on Modern GPU Architectures.Karthik Vadambacheri Manian, A. A. Ammar, Amit Ruhela, Ching-Hsiang Chu, Hari Subramoni, Dhabaleswar K. Panda
2019CCGRIDScalable Distributed DNN Training using TensorFlow and CUDA-Aware MPI: Characterization, Designs, and Performance Evaluation.Ammar Ahmad Awan, Jeroen Bdorf, Ching-Hsiang Chu, Hari Subramoni, Dhabaleswar K. Panda
2019CCGRIDDesign and Characterization of Shared Address Space MPI Collectives on Modern Architectures.Jahanzeb Maqbool Hashmi, Sourav Chakraborty, Mohammadreza Bayatpour, Hari Subramoni, Dhabaleswar K. Panda
2019CLUSTERPerformance Characterization of DNN Training using TensorFlow and PyTorch on Modern Clusters.Arpan Jain, Ammar Ahmad Awan, Quentin Anthony, Hari Subramoni, Dhabaleswar K. Panda
2019HiPCHigh-Performance Adaptive MPI Derived Datatype Communication for Modern Multi-GPU Systems.Ching-Hsiang Chu, Jahanzeb Maqbool Hashmi, Kawthar Shafie Khorassani, Hari Subramoni, Dhabaleswar K. Panda
2019HiPCDesigning a Profiling and Visualization Tool for Scalable and In-depth Analysis of High-Performance GPU Clusters.Pouya Kousha, Bharath Ramesh, Kaushik Kandadi Suresh, Ching-Hsiang Chu, Arpan Jain, Nick Sarkauskas, Hari Subramoni, Dhabaleswar K. Panda
2019HiPCSCOR-KV: SIMD-Aware Client-Centric and Optimistic RDMA-Based Key-Value Store for Emerging CPU Architectures.Dipti Shankar, Xiaoyi Lu, Dhabaleswar K. Panda
2019HOTIDesigning Scalable and High-Performance MPI Libraries on Amazon Elastic Fabric Adapter.Sourav Chakraborty, Shulei Xu, Hari Subramoni, Dhabaleswar K. Panda
2019HOTICommunication Profiling and Characterization of Deep Learning Workloads on Clusters with High-Performance Interconnects.Ammar Ahmad Awan, Arpan Jain, Ching-Hsiang Chu, Hari Subramoni, Dhabaleswar K. Panda
2019HPDCUMR-EC: A Unified and Multi-Rail Erasure Coding Library for High-Performance Distributed Storage Systems.Haiyang Shi, Xiaoyi Lu, Dipti Shankar, Dhabaleswar K. Panda
2019PPoPPHigh performance distributed deep learning: a beginner's guide.Dhabaleswar K. Panda, Ammar Ahmad Awan, Hari Subramoni
2019SCScaling TensorFlow, PyTorch, and MXNet using MVAPICH2 for High-Performance Deep Learning on Frontera.Arpan Jain, Ammar Ahmad Awan, Hari Subramoni, Dhabaleswar K. Panda
2019SCLeveraging Network-level parallelism with Multiple Process-Endpoints for MPI Broadcast.Amit Ruhela, Bharath Ramesh, Sourav Chakraborty, Hari Subramoni, Jahanzeb Maqbool Hashmi, Dhabaleswar K. Panda
2018CLOUDHigh-Performance Multi-Rail Erasure Coding Library over Modern Data Center Architectures: Early Experiences.Haiyang Shi, Xiaoyi Lu, Dipti Shankar, Dhabaleswar K. Panda
2018CLUSTERSALaR: Scalable and Adaptive Designs for Large Message Reduction Collectives.Mohammadreza Bayatpour, Jahanzeb Maqbool Hashmi, Sourav Chakraborty, Hari Subramoni, Pouya Kousha, Dhabaleswar K. Panda
2018CLUSTERCutting the Tail: Designing High Performance Message Brokers to Reduce Tail Latencies in Stream Processing.M. Haseeb Javed, Xiaoyi Lu, Dhabaleswar K. Panda
2018HiPCOC-DNN: Exploiting Advanced Unified Memory Capabilities in CUDA 9 and Volta GPUs for Out-of-Core DNN Training.Ammar Ahmad Awan, Ching-Hsiang Chu, Hari Subramoni, Xiaoyi Lu, Dhabaleswar K. Panda
2018HiPCAccelerating TensorFlow with Adaptive RDMA-Based gRPC.Rajarshi Biswas, Xiaoyi Lu, Dhabaleswar K. Panda
2018SCCooperative rendezvous protocols for improved performance and overlap.Sourav Chakraborty, Mohammadreza Bayatpour, Jahanzeb Maqbool Hashmi, Hari Subramoni, Dhabaleswar K. Panda
2017CCGRIDSwift-X: Accelerating OpenStack Swift with RDMA for Building an Efficient HPC Cloud.Shashank Gugnani, Xiaoyi Lu, Dhabaleswar K. Panda
2017CLUSTERContention-Aware Kernel-Assisted MPI Collectives for Multi-/Many-Core Systems.Sourav Chakraborty, Hari Subramoni, Dhabaleswar K. Panda
2017CLUSTERA Scalable Network-Based Performance Analysis Tool for MPI on Large-Scale HPC Systems.Hari Subramoni, Xiaoyi Lu, Dhabaleswar K. Panda
2017HiPCMPI-LiFE: Designing High-Performance Linear Fascicle Evaluation of Brain Connectome with MPI.Shashank Gugnani, Xiaoyi Lu, Franco Pestilli, Cesar F. Caiafa, Dhabaleswar K. Panda
2017HiPCKernel-Assisted Communication Engine for MPI on Emerging Manycore Processors.Jahanzeb Maqbool Hashmi, Khaled Hamidouche, Hari Subramoni, Dhabaleswar K. Panda
2017HiPCDesigning Registration Caching Free High-Performance MPI Library with Implicit On-Demand Paging (ODP) of InfiniBand.Mingzhe Li, Xiaoyi Lu, Hari Subramoni, Dhabaleswar K. Panda
2017HOTICharacterizing Deep Learning over Big Data (DLoBD) Stacks on RDMA-Capable Networks.Xiaoyi Lu, Haiyang Shi, M. Haseeb Javed, Rajarshi Biswas, Dhabaleswar K. Panda
2017ICDCSHigh-Performance and Resilient Key-Value Store with Online Erasure Coding for Big Data Workloads.Dipti Shankar, Xiaoyi Lu, Dhabaleswar K. Panda
2017ICPPEfficient and Scalable Multi-Source Streaming Broadcast on GPU Clusters for Deep Learning.Ching-Hsiang Chu, Xiaoyi Lu, Ammar Ahmad Awan, Hari Subramoni, Jahanzeb Maqbool Hashmi, Bracy Elton, Dhabaleswar K. Panda
2017ICPPMPI-GDS: High Performance MPI Designs with GPUDirect-aSync for CPU-GPU Control Flow Decoupling.Akshay Venkatesh, Khaled Hamidouche, Sreeram Potluri, Davide Rossetti, Ching-Hsiang Chu, Dhabaleswar K. Panda
2017PPoPPS-Caffe: Co-designing MPI Runtimes and Caffe for Scalable Deep Learning on Modern GPU Clusters.Ammar Ahmad Awan, Khaled Hamidouche, Jahanzeb Maqbool Hashmi, Dhabaleswar K. Panda
2017SCAn In-depth Performance Characterization of CPU- and GPU-based DNN Training on Modern Architectures.Ammar Ahmad Awan, Hari Subramoni, Dhabaleswar K. Panda
2017SCScalable reduction collectives with data partitioning-based multi-leader design.Mohammadreza Bayatpour, Sourav Chakraborty, Hari Subramoni, Xiaoyi Lu, Dhabaleswar K. Panda
2016CCGRIDSHMEMPMI - Shared Memory Based PMI for Improved Performance and Scalability.Sourav Chakraborty, Hari Subramoni, Jonathan L. Perkins, Dhabaleswar K. Panda
2016CCGRIDCUDA Kernel Based Collective Reduction Operations on Large-scale GPU Clusters.Ching-Hsiang Chu, Khaled Hamidouche, Akshay Venkatesh, Ammar Ahmad Awan, Dhabaleswar K. Panda
2016CloudComRe-Designing CNTK Deep Learning Framework on Modern GPU Enabled Clusters.Dip Sankar Banerjee, Khaled Hamidouche, Dhabaleswar K. Panda
2016CloudComDesigning Virtualization-Aware and Automatic Topology Detection Schemes for Accelerating Hadoop on SR-IOV-Enabled Clouds.Shashank Gugnani, Xiaoyi Lu, Dhabaleswar K. Panda
2016CloudComImpact of HPC Cloud Networking Technologies on Accelerating Hadoop RPC and HBase.Xiaoyi Lu, Dipti Shankar, Shashank Gugnani, Hari Subramoni, Dhabaleswar K. Panda
2016CLUSTERAdaptive and Dynamic Design for MPI Tag Matching.Mohammadreza Bayatpour, Hari Subramoni, Sourav Chakraborty, Dhabaleswar K. Panda
2016EuroParSlurm-V: Extending Slurm for Building Efficient HPC Cloud with SR-IOV and IVShmem.Jie Zhang, Xiaoyi Lu, Sourav Chakraborty, Dhabaleswar K. Panda
2016HiPCCUDA M3: Designing Efficient CUDA Managed Memory-Aware MPI by Exploiting GDR and IPC.Khaled Hamidouche, Ammar Ahmad Awan, Akshay Venkatesh, Dhabaleswar K. Panda
2016HiPCMizan-RMA: Accelerating Mizan Graph Processing Framework with MPI RMA.Mingzhe Li, Xiaoyi Lu, Khaled Hamidouche, Jie Zhang, Dhabaleswar K. Panda
2016HPCCEnabling Performance Efficient Runtime Support for Hybrid MPI+UPC++ Programming Models.Jahanzeb Maqbool Hashmi, Khaled Hamidouche, Dhabaleswar K. Panda
2016ICPADSSystem-Level Scalable Checkpoint-Restart for Petascale Computing.Jiajun Cao, Kapil Arya, Rohan Garg, L. Shawn Matott, Dhabaleswar K. Panda, Hari Subramoni, Jrme Vienne, Gene Cooperman
2016ICPPHigh Performance MPI Library for Container-Based HPC Cloud on InfiniBand Clusters.Jie Zhang, Xiaoyi Lu, Dhabaleswar K. Panda
2016ICSHigh Performance Design for HDFS with Byte-Addressability of NVM and RDMA.Nusrat Sharmin Islam, Md. Wasi-ur-Rahman, Xiaoyi Lu, Dhabaleswar K. Panda
2016PPoPPDesigning high performance communication runtime for GPU managed memory: early experiences.Dip Sankar Banerjee, Khaled Hamidouche, Dhabaleswar K. Panda
2016SCEfficient Reliability Support for Hardware Multicast-Based Broadcast in GPU-enabled Streaming Applications.Ching-Hsiang Chu, Khaled Hamidouche, Hari Subramoni, Akshay Venkatesh, Bracy Elton, Dhabaleswar K. Panda
2016SCOpenSHMEM Non-blocking Data Movement Operations with MVAPICH2-X: Early Experiences.Khaled Hamidouche, Jie Zhang, Dhabaleswar K. Panda, Karen Tomko
2016SCDesigning MPI library with on-demand paging (ODP) of infiniband: challenges and benefits.Mingzhe Li, Khaled Hamidouche, Xiaoyi Lu, Hari Subramoni, Jie Zhang, Dhabaleswar K. Panda
2015CCGRIDNon-Blocking PMI Extensions for Fast MPI Startup.Sourav Chakraborty, Hari Subramoni, Adam Moody, Akshay Venkatesh, Jonathan L. Perkins, Dhabaleswar K. Panda
2015CCGRIDPower-Check: An Energy-Efficient Checkpointing Framework for HPC Clusters.Raghunath Raja Chandrasekar, Akshay Venkatesh, Khaled Hamidouche, Dhabaleswar K. Panda
2015CCGRIDTriple-H: A Hybrid Approach to Accelerate HDFS on HPC Clusters with Heterogeneous Storage Architecture.Nusrat Sharmin Islam, Xiaoyi Lu, Md. Wasi-ur-Rahman, Dipti Shankar, Dhabaleswar K. Panda
2015CCGRIDMVAPICH2 over OpenStack with SR-IOV: An Efficient Approach to Build HPC Clouds.Jie Zhang, Xiaoyi Lu, Mark Daniel Arnold, Dhabaleswar K. Panda
2015CLUSTERExploiting GPUDirect RDMA in Designing High Performance OpenSHMEM for NVIDIA GPU Clusters.Khaled Hamidouche, Akshay Venkatesh, Ammar Ahmad Awan, Hari Subramoni, Ching-Hsiang Chu, Dhabaleswar K. Panda
2015CLUSTERHigh Performance MPI Datatype Support with User-Mode Memory Registration: Challenges, Designs, and Benefits.Mingzhe Li, Hari Subramoni, Khaled Hamidouche, Xiaoyi Lu, Dhabaleswar K. Panda
2015EuroParHigh-Performance and Scalable Design of MPI-3 RMA on Xeon Phi Clusters.Mingzhe Li, Khaled Hamidouche, Xiaoyi Lu, Jian Lin, Dhabaleswar K. Panda
2015HiPCHigh Performance OpenSHMEM Strided Communication Support with InfiniBand UMR.Mingzhe Li, Khaled Hamidouche, Xiaoyi Lu, Jie Zhang, Jian Lin, Dhabaleswar K. Panda
2015HiPCOffloaded GPU Collectives Using CORE-Direct and CUDA Capabilities on InfiniBand Clusters.Akshay Venkatesh, Khaled Hamidouche, Hari Subramoni, Dhabaleswar K. Panda
2015HOTIImpact of InfiniBand DC Transport Protocol on Energy Consumption of All-to-All Collective Algorithms.Hari Subramoni, Akshay Venkatesh, Khaled Hamidouche, Karen Tomko, Dhabaleswar K. Panda
2015ICPPAccelerating I/O Performance of Big Data Analytics on HPC Clusters through RDMA-Based Key-Value Store.Nusrat Sharmin Islam, Dipti Shankar, Xiaoyi Lu, Md. Wasi-ur-Rahman, Dhabaleswar K. Panda
2015ISPASSCan RDMA benefit online data processing workloads on memcached and MySQL?Dipti Shankar, Xiaoyi Lu, Jithin Jose, Md. Wasi-ur-Rahman, Nusrat S. Islam, Dhabaleswar K. Panda
2014CLUSTERHigh performance OpenSHMEM for Xeon Phi clusters: Extensions, runtime designs and application co-design.Jithin Jose, Khaled Hamidouche, Xiaoyi Lu, Sreeram Potluri, Jie Zhang, Karen Tomko, Dhabaleswar K. Panda
2014CLUSTERScalable Graph500 design with MPI-3 RMA.Mingzhe Li, Xiaoyi Lu, Sreeram Potluri, Khaled Hamidouche, Jithin Jose, Karen Tomko, Dhabaleswar K. Panda
2014EuroParMapReduce over Lustre: Can RDMA-Based Approach Benefit?Md. Wasi-ur-Rahman, Xiaoyi Lu, Nusrat Sharmin Islam, Raghunath Rajachandrasekar, Dhabaleswar K. Panda
2014EuroParCan Inter-VM Shmem Benefit MPI Applications on SR-IOV Based Virtualized Infiniband Clusters?Jie Zhang, Xiaoyi Lu, Jithin Jose, Rong Shi, Dhabaleswar K. Panda
2014HiPCDesigning efficient small message transfer mechanism for inter-node MPI communication on InfiniBand GPU clusters.Rong Shi, Sreeram Potluri, Khaled Hamidouche, Jonathan L. Perkins, Mingzhe Li, Davide Rossetti, Dhabaleswar K. Panda
2014HiPCA high performance broadcast design with hardware multicast and GPUDirect RDMA for streaming applications on Infiniband clusters.Akshay Venkatesh, Hari Subramoni, Khaled Hamidouche, Dhabaleswar K. Panda
2014HiPCHigh performance MPI library over SR-IOV enabled infiniband clusters.Jie Zhang, Xiaoyi Lu, Jithin Jose, Mingzhe Li, Rong Shi, Dhabaleswar K. Panda
2014HOTIAccelerating Spark with RDMA for Big Data Processing: Early Experiences.Xiaoyi Lu, Md. Wasi-ur-Rahman, Nusrat S. Islam, Dipti Shankar, Dhabaleswar K. Panda
2014HPDCSOR-HDFS: a SEDA-based approach to maximize overlapping in RDMA-enhanced HDFS.Nusrat S. Islam, Xiaoyi Lu, Md. Wasi-ur-Rahman, Dhabaleswar K. Panda
2014HPDCMIC-Check: a distributed check pointing framework for the intel many integrated cores architecture.Raghunath Rajachandrasekar, Sreeram Potluri, Akshay Venkatesh, Khaled Hamidouche, Md. Wasi-ur-Rahman, Dhabaleswar K. Panda
2014ICPADSMessage from the general co-chairs IEEE ICPADS 2014.Dhabaleswar K. Panda, Jang-Ping Sheu
2014ICPPHAND: A Hybrid Approach to Accelerate Non-contiguous Data Movement Using MPI Datatypes on GPU Clusters.Rong Shi, Xiaoyi Lu, Sreeram Potluri, Khaled Hamidouche, Jie Zhang, Dhabaleswar K. Panda
2014ICPPDesigning Topology-Aware Communication Schedules for Alltoall Operations in Large InfiniBand Clusters.Hari Subramoni, Krishna Chaitanya Kandalla, Jithin Jose, Karen Tomko, Karl W. Schulz, Dmitry Pekurovsky, Dhabaleswar K. Panda
2014ICPPPerformance Modeling for RDMA-Enhanced Hadoop MapReduce.Md. Wasi-ur-Rahman, Xiaoyi Lu, Nusrat Sharmin Islam, Dhabaleswar K. Panda
2014ICSHOMR: a hybrid approach to exploit maximum overlapping in MapReduce over high performance interconnects.Md. Wasi-ur-Rahman, Xiaoyi Lu, Nusrat Sharmin Islam, Dhabaleswar K. Panda
2014PPoPPInitial study of multi-endpoint runtime for MPI+OpenMP hybrid programming model on multi-core systems.Miao Luo, Xiaoyi Lu, Khaled Hamidouche, Krishna Chaitanya Kandalla, Dhabaleswar K. Panda
2013CCGRIDSR-IOV Support for Virtualization on InfiniBand Clusters: Early Experience.Jithin Jose, Mingzhe Li, Xiaoyi Lu, Krishna Chaitanya Kandalla, Mark Daniel Arnold, Dhabaleswar K. Panda
2013CCGRIDEfficient Intra-node Communication on Intel-MIC Clusters.Sreeram Potluri, Akshay Venkatesh, Devendar Bureddy, Krishna Chaitanya Kandalla, Dhabaleswar K. Panda
2013CLOUDDoes RDMA-based enhanced Hadoop MapReduce need a new performance model?Md. Wasi-ur-Rahman, Xiaoyi Lu, Nusrat S. Islam, Dhabaleswar K. Panda
2013CLUSTERA scalable and portable approach to accelerate hybrid HPL on heterogeneous CPU-GPU clusters.Rong Shi, Sreeram Potluri, Khaled Hamidouche, Xiaoyi Lu, Karen Tomko, Dhabaleswar K. Panda
2013CLUSTERDesign of network topology aware scheduling services for large InfiniBand clusters.Hari Subramoni, Devendar Bureddy, Krishna Chaitanya Kandalla, Karl W. Schulz, Bill Barth, Jonathan L. Perkins, Mark Daniel Arnold, Dhabaleswar K. Panda
2013HOTITutorials.Dhabaleswar K. Panda, Xiaoyi Lu
2013HOTICan Parallel Replication Benefit Hadoop Distributed File System for High Performance Interconnects?Nusrat S. Islam, Xiaoyi Lu, Md. Wasi-ur-Rahman, Dhabaleswar K. Panda
2013HOTIDesigning Optimized MPI Broadcast and Allreduce for Many Integrated Core (MIC) InfiniBand Clusters.Krishna Chaitanya Kandalla, Akshay Venkatesh, Khaled Hamidouche, Sreeram Potluri, Devendar Bureddy, Dhabaleswar K. Panda
2013HPDCA 1 PB/s file system to checkpoint three million MPI tasks.Raghunath Rajachandrasekar, Adam Moody, Kathryn Mohror, Dhabaleswar K. Panda
2013ICPPA Novel Functional Partitioning Approach to Design High-Performance MPI-3 Non-blocking Alltoallv Collective on Multi-core Systems.Krishna Chaitanya Kandalla, Hari Subramoni, Karen Tomko, Dmitry Pekurovsky, Dhabaleswar K. Panda
2013ICPPHigh-Performance Design of Hadoop RPC with RDMA over InfiniBand.Xiaoyi Lu, Nusrat S. Islam, Md. Wasi-ur-Rahman, Jithin Jose, Hari Subramoni, Hao Wang, Dhabaleswar K. Panda
2013ICPPEfficient Inter-node MPI Communication Using GPUDirect RDMA for InfiniBand Clusters with NVIDIA GPUs.Sreeram Potluri, Khaled Hamidouche, Akshay Venkatesh, Devendar Bureddy, Dhabaleswar K. Panda
2013ICSMIC-RO: enabling efficient remote offload on heterogeneous many integrated core (MIC) clusters with InfiniBand.Khaled Hamidouche, Sreeram Potluri, Hari Subramoni, Krishna Chaitanya Kandalla, Dhabaleswar K. Panda
2013SCMVAPICH-PRISM: a proxy-based communication framework using InfiniBand and SCIF for intel MIC clusters.Sreeram Potluri, Devendar Bureddy, Khaled Hamidouche, Akshay Venkatesh, Krishna Chaitanya Kandalla, Hari Subramoni, Dhabaleswar K. Panda
2012CCGRIDScalable Memcached Design for InfiniBand Clusters Using Hybrid Transports.Jithin Jose, Hari Subramoni, Krishna Chaitanya Kandalla, Md. Wasi-ur-Rahman, Hao Wang, Sundeep Narravula, Dhabaleswar K. Panda
2012CLUSTERCan Network-Offload Based Non-blocking Neighborhood MPI Collectives Improve Communication Overheads of Irregular Graph Algorithms?Krishna Chaitanya Kandalla, Aydin Bulu, Hari Subramoni, Karen Tomko, Jrme Vienne, Leonid Oliker, Dhabaleswar K. Panda
2012CLUSTERMinimizing Network Contention in InfiniBand Clusters with a QoS-Aware Data-Staging Framework.Raghunath Rajachandrasekar, Jai Jaswani, Hari Subramoni, Dhabaleswar K. Panda
2012EuroParA Scalable InfiniBand Network Topology-Aware Performance Analysis Tool for MPI.Hari Subramoni, Jrme Vienne, Dhabaleswar K. Panda
2012HOTIPerformance Analysis and Evaluation of InfiniBand FDR and 40GigE RoCE on HPC and Cloud Computing Systems.Jrme Vienne, Jitong Chen, Md. Wasi-ur-Rahman, Nusrat S. Islam, Hari Subramoni, Dhabaleswar K. Panda
2012ICPPSupporting Hybrid MPI and OpenSHMEM over InfiniBand: Design and Performance Evaluation.Jithin Jose, Krishna Chaitanya Kandalla, Miao Luo, Dhabaleswar K. Panda
2012ICPPSSD-Assisted Hybrid Memory to Accelerate Memcached over High Performance Networks.Xiangyong Ouyang, Nusrat S. Islam, Raghunath Rajachandrasekar, Jithin Jose, Miao Luo, Hao Wang, Dhabaleswar K. Panda
2012ICSCongestion avoidance on manycore high performance computing systems.Miao Luo, Dhabaleswar K. Panda, Khaled Z. Ibrahim, Costin Iancu
2012ISPASSUnderstanding the communication characteristics in HBase: What are the fundamental bottlenecks?Md. Wasi-ur-Rahman, Jian Huang, Jithin Jose, Xiangyong Ouyang, Hao Wang, Nusrat S. Islam, Hari Subramoni, Chet Murthy, Dhabaleswar K. Panda
2012SCHigh performance RDMA-based design of HDFS over InfiniBand.Nusrat S. Islam, Md. Wasi-ur-Rahman, Jithin Jose, Raghunath Rajachandrasekar, Hao Wang, Hari Subramoni, Chet Murthy, Dhabaleswar K. Panda
2012SCDesign of a scalable InfiniBand topology service to enable network-topology-aware placement of processes.Hari Subramoni, Sreeram Potluri, Krishna Chaitanya Kandalla, Bill Barth, Jrme Vienne, Jeff Keasler, Karen A. Tomko, Karl W. Schulz, Adam Moody, Dhabaleswar K. Panda
2011CCGRIDHigh Performance Pipelined Process Migration with RDMA.Xiangyong Ouyang, Raghunath Rajachandrasekar, Xavier Besseron, Dhabaleswar K. Panda
2011CLUSTERCan a Decentralized Metadata Service Layer Benefit Parallel Filesystems?Vilobh Meshram, Xavier Besseron, Xiangyong Ouyang, Raghunath Rajachandrasekar, Ravi Prakash, Dhabaleswar K. Panda
2011CLUSTERMPI Alltoall Personalized Exchange on GPGPU Clusters: Design Alternatives and Benefit.Ashish Kumar Singh, Sreeram Potluri, Hao Wang, Krishna Chaitanya Kandalla, Sayantan Sur, Dhabaleswar K. Panda
2011CLUSTERDesign and Evaluation of Network Topology-/Speed- Aware Broadcast Algorithms for InfiniBand Clusters.Hari Subramoni, Krishna Chaitanya Kandalla, Jrme Vienne, Sayantan Sur, Bill Barth, Karen A. Tomko, Robert T. McLay, Karl W. Schulz, Dhabaleswar K. Panda
2011CLUSTEROptimized Non-contiguous MPI Datatype Communication for GPU Clusters: Design, Implementation and Evaluation with MVAPICH2.Hao Wang, Sreeram Potluri, Miao Luo, Ashish Kumar Singh, Xiangyong Ouyang, Sayantan Sur, Dhabaleswar K. Panda
2011EuroParINAM - A Scalable InfiniBand Network Analysis and Monitoring Tool.N. Dandapanthula, Hari Subramoni, Jrme Vienne, Krishna Chaitanya Kandalla, Sayantan Sur, Dhabaleswar K. Panda, Ron Brightwell
2011EuroParCan Checkpoint/Restart Mechanisms Benefit from Hierarchical Data Staging?Raghunath Rajachandrasekar, Xiangyong Ouyang, Xavier Besseron, Vilobh Meshram, Dhabaleswar K. Panda
2011HiPCMulti-threaded UPC runtime with network endpoints: Design alternatives and evaluation on multi-core architectures.Miao Luo, Jithin Jose, Sayantan Sur, Dhabaleswar K. Panda
2011HOTIDesigning Non-blocking Broadcast with Collective Offload on InfiniBand Clusters: A Case Study with HPL.Krishna Chaitanya Kandalla, Hari Subramoni, Jrme Vienne, S. Pai Raikar, Karen Tomko, Sayantan Sur, Dhabaleswar K. Panda
2011HPCABeyond block I/O: Rethinking traditional storage primitives.Xiangyong Ouyang, David W. Nellans, Robert Wipfel, David Flynn, Dhabaleswar K. Panda
2011ICPPMemcached Design on High Performance RDMA Capable Interconnects.Jithin Jose, Hari Subramoni, Miao Luo, Minjia Zhang, Jian Huang, Md. Wasi-ur-Rahman, Nusrat S. Islam, Xiangyong Ouyang, Hao Wang, Sayantan Sur, Dhabaleswar K. Panda
2011ICPPCRFS: A Lightweight User-Level Filesystem for Generic Checkpoint/Restart.Xiangyong Ouyang, Raghunath Rajachandrasekar, Xavier Besseron, Hao Wang, Jian Huang, Dhabaleswar K. Panda
2010CCGRIDAn MPI-Stream Hybrid Programming Model for Computational Clusters.Emilio Pasquale Mancini, Gregory Marsh, Dhabaleswar K. Panda
2010CCGRIDHigh Performance Data Transfer in Grid Environment Using GridFTP over InfiniBand.Hari Subramoni, Ping Lai, Rajkumar Kettimuthu, Dhabaleswar K. Panda
2010CLUSTERRDMA-Based Job Migration Framework for MPI over InfiniBand.Xiangyong Ouyang, Sonya Marcarelli, Raghunath Rajachandrasekar, Dhabaleswar K. Panda
2010HOTIDesigning High-End Computing Systems with InfiniBand and High-Speed Ethernet.Dhabaleswar K. Panda, Sayantan Sur, Pavan Balaji
2010HOTIDesign and Evaluation of Generalized Collective Communication Primitives with Overlap Using ConnectX-2 Offload Engine.Hari Subramoni, Krishna Chaitanya Kandalla, Sayantan Sur, Dhabaleswar K. Panda
2010ICPPDesigning Power-Aware Collective Communication Algorithms for InfiniBand Clusters.Krishna Chaitanya Kandalla, Emilio Pasquale Mancini, Sayantan Sur, Dhabaleswar K. Panda
2010ICPPImproving Application Performance and Predictability Using Multiple Virtual Lanes in Modern Multi-core InfiniBand Clusters.Hari Subramoni, Ping Lai, Sayantan Sur, Dhabaleswar K. Panda
2010ICSQuantifying performance benefits of overlap using MPI-2 in a seismic modeling application.Sreeram Potluri, Ping Lai, Karen A. Tomko, Sayantan Sur, Yifeng Cui, Mahidhar Tatineni, Karl W. Schulz, William L. Barth, Amitava Majumdar, Dhabaleswar K. Panda
2010SCScalable Earthquake Simulation on Petascale Supercomputers.Yifeng Cui, Kim B. Olsen, Thomas H. Jordan, Kwangyoon Lee, Jun Zhou, Patrick Small, Daniel Roten, Geoffrey Ely, Dhabaleswar K. Panda, Amit Chourasia, John M. Levesque, Steven M. Day, Philip Maechling
2009CCGRIDNatively Supporting True One-Sided Communication in.Gopalakrishnan Santhanaraman, Pavan Balaji, K. Gopalakrishnan, Rajeev Thakur, William Gropp, Dhabaleswar K. Panda
2009CLUSTERReducing network contention with mixed workloads on modern multicore, clusters.Matthew J. Koop, Miao Luo, Dhabaleswar K. Panda
2009CLUSTERDesign alternatives for implementing fence synchronization in MPI-2 one-sided communication for InfiniBand clusters.Gopalakrishnan Santhanaraman, Tejus Gangadharappa, Sundeep Narravula, Amith R. Mamidala, Dhabaleswar K. Panda
2009CLUSTERRDMA over Ethernet - A preliminary study.Hari Subramoni, Ping Lai, Miao Luo, Dhabaleswar K. Panda
2009CLUSTERAn efficient hardware-software approach to network fault tolerance with InfiniBand.Abhinav Vishnu, Manojkumar Krishnan, Dhabaleswar K. Panda
2009HiPCFast checkpointing by Write Aggregation with Dynamic Buffer and Interleaving on multicore architecture.Xiangyong Ouyang, Karthik Gopalakrishnan, Tejus Gangadharappa, Dhabaleswar K. Panda
2009HOTITutorial: Infiniband and 10-Gigabit Ethernet for Dummies.Dhabaleswar K. Panda, Matthew J. Koop, Pavan Balaji
2009HOTITutorial: Designing High-End Computing Systems with Infiniband and 10-Gigabit Ethernet.Dhabaleswar K. Panda, Matthew J. Koop, Pavan Balaji
2009HOTIDesigning Next Generation Clusters: Evaluation of InfiniBand DDR/QDR on Intel Computing Platforms.Hari Subramoni, Matthew J. Koop, Dhabaleswar K. Panda
2009ICPPCIFTS: A Coordinated Infrastructure for Fault-Tolerant Systems.Rinku Gupta, Peter H. Beckman, Byung-Hoon Park, Ewing L. Lusk, Paul Hargrove, Al Geist, Dhabaleswar K. Panda, Andrew Lumsdaine, Jack J. Dongarra
2009ICPPDesigning Efficient FTP Mechanisms for High Performance Data-Transfer over InfiniBand.Ping Lai, Hari Subramoni, Sundeep Narravula, Amith R. Mamidala, Dhabaleswar K. Panda
2009ICPPAccelerating Checkpoint Operation by Node-Level Write Aggregation on Multicore Systems.Xiangyong Ouyang, Karthik Gopalakrishnan, Dhabaleswar K. Panda
2008CCGRIDAdvanced RDMA-Based Admission Control for Modern Data-Centers.Ping Lai, Sundeep Narravula, Karthikeyan Vaidyanathan, Dhabaleswar K. Panda
2008CCGRIDMPI Collectives on Modern Multicore Clusters: Performance Optimizations and Communication Characteristics.Amith R. Mamidala, Rahul Kumar, Debraj De, Dhabaleswar K. Panda
2008CCGRIDOptimized Distributed Data Sharing Substrate in Multi-core Commodity Clusters: A Comprehensive Study with Applications.Karthikeyan Vaidyanathan, Ping Lai, Sundeep Narravula, Dhabaleswar K. Panda
2008CLUSTEREfficient one-copy MPI shared memory communication in Virtual Machines.Wei Huang, Matthew J. Koop, Dhabaleswar K. Panda
2008CLUSTERScalable MPI design over InfiniBand using eXtended Reliable Connection.Matthew J. Koop, Jaidev K. Sridhar, Dhabaleswar K. Panda
2008CLUSTERDesigning next generation clusters with InfiniBand and 10GE/iWARP: Opportunities and challenges.Dhabaleswar K. Panda
2008HiPCSockets Direct Protocol for Hybrid Network Stacks: A Case Study with iWARP over 10G Ethernet.Pavan Balaji, Sitha Bhagvat, Rajeev Thakur, Dhabaleswar K. Panda
2008HiPCDesigning a High-Performance Clustered NAS: A Case Study with pNFS over RDMA on InfiniBand.Ranjit Noronha, Xiangyong Ouyang, Dhabaleswar K. Panda
2008HiPCScELA: Scalable and Extensible Launching Architecture for Clusters.Jaidev K. Sridhar, Matthew J. Koop, Jonathan L. Perkins, Dhabaleswar K. Panda
2008HOTIPerformance Analysis and Evaluation of PCIe 2.0 and Quad-Data Rate InfiniBand.Matthew J. Koop, Wei Huang, Karthik Gopalakrishnan, Dhabaleswar K. Panda
2008ICPPDesigning an Efficient Kernel-Level and User-Level Hybrid Approach for MPI Intra-Node Communication on Multi-Core Systems.Lei Chai, Ping Lai, Hyun-Wook Jin, Dhabaleswar K. Panda
2008ICPPPerformance of HPC Middleware over InfiniBand WAN.Sundeep Narravula, Hari Subramoni, Ping Lai, Ranjit Noronha, Dhabaleswar K. Panda
2008ICPPIMCa: A High Performance Caching Front-End for GlusterFS on InfiniBand.Ranjit Noronha, Dhabaleswar K. Panda
2008ICSCan software reliability outperform hardware reliability on high performance interconnects?: a case study with MPI over infiniband.Matthew J. Koop, Rahul Kumar, Dhabaleswar K. Panda
2007CCGRIDUnderstanding the Impact of Multi-Core Architecture in Cluster Computing: A Case Study with Intel Dual-Core System.Lei Chai, Qi Gao, Dhabaleswar K. Panda
2007CCGRIDReducing Connection Memory Requirements of MPI for InfiniBand Clusters: A Message Coalescing Approach.Matthew J. Koop, Terry R. Jones, Dhabaleswar K. Panda
2007CCGRIDHigh Performance Distributed Lock Management Services using Network-based Remote Atomic Operations.Sundeep Narravula, A. Marnidala, Abhinav Vishnu, Karthikeyan Vaidyanathan, Dhabaleswar K. Panda
2007CCGRIDHot-Spot Avoidance With Multi-Pathing Over InfiniBand: An MPI Perspective.Abhinav Vishnu, Matthew J. Koop, Adam Moody, Amith R. Mamidala, Sundeep Narravula, Dhabaleswar K. Panda
2007CLUSTERHigh performance virtual machine migration with RDMA over modern interconnects.Wei Huang, Qi Gao, Jiuxing Liu, Dhabaleswar K. Panda
2007CLUSTERLightweight kernel-level primitives for high-performance MPI intra-node communication over multi-core systems.Hyun-Wook Jin, Sayantan Sur, Lei Chai, Dhabaleswar K. Panda
2007CLUSTERZero-copy protocol for MPI using infiniband unreliable datagram.Matthew J. Koop, Sayantan Sur, Dhabaleswar K. Panda
2007CLUSTERDesigning high-end computing systems with InfiniBand and10-Gigabit Ethernet iWARP.Dhabaleswar K. Panda, Pavan Balaji
2007CLUSTEREfficient asynchronous memory copy operations on multi-core systems and I/OAT.Karthikeyan Vaidyanathan, Lei Chai, Wei Huang, Dhabaleswar K. Panda
2007HOTIPerformance Analysis and Evaluation of Mellanox ConnectX InfiniBand Architecture with Multi-Core Platforms.Sayantan Sur, Matthew J. Koop, Lei Chai, Dhabaleswar K. Panda
2007ICPPAdvanced Flow-control Mechanisms for the Sockets Direct Protocol over InfiniBand.Pavan Balaji, Sitha Bhagvat, Dhabaleswar K. Panda, Rajeev Thakur, William Gropp
2007ICPPGroup-based Coordinated Checkpointing for MPI: A Case Study on InfiniBand.Qi Gao, Wei Huang, Matthew J. Koop, Dhabaleswar K. Panda
2007ICPPHigh Performance MPI over iWARP: Early Experiences.Sundeep Narravula, Amith R. Mamidala, Abhinav Vishnu, Gopalakrishnan Santhanaraman, Dhabaleswar K. Panda
2007ICPPDesigning NFS with RDMA for Security, Performance and Scalability.Ranjit Noronha, Lei Chai, Thomas Talpey, Dhabaleswar K. Panda
2007ICSHigh performance MPI design using unreliable datagram for ultra-scale InfiniBand clusters.Matthew J. Koop, Sayantan Sur, Qi Gao, Dhabaleswar K. Panda
2007ISPASSBenefits of I/O Acceleration Technology (I/OAT) in Clusters.Karthikeyan Vaidyanathan, Dhabaleswar K. Panda
2007PPoPPOn using connection-oriented vs. connection-less transport for performance and scalability of collective and one-sided operations: trade-offs and impact.Amith R. Mamidala, Sundeep Narravula, Abhinav Vishnu, Gopalakrishnan Santhanaraman, Dhabaleswar K. Panda
2007SCAnalyzing the impact of supporting out-of-order communication on in-order performance with iWARP.Pavan Balaji, Wu-chun Feng, Sitha Bhagvat, Dhabaleswar K. Panda, Rajeev Thakur, William Gropp
2007SCpNFS/PVFS2 over InfiniBand: early experiences.Lei Chai, Xiangyong Ouyang, Ranjit Noronha, Dhabaleswar K. Panda
2007SCDMTracker: finding bugs in large-scale parallel programs by detecting anomaly in data movements.Qi Gao, Feng Qin, Dhabaleswar K. Panda
2007SCVirtual machine aware communication libraries for high performance computing.Wei Huang, Matthew J. Koop, Qi Gao, Dhabaleswar K. Panda
2006CCGRIDMPI over uDAPL: Can High Performance and Portability Exist Across Architectures?.Lei Chai, Ranjit Noronha, Dhabaleswar K. Panda
2006CCGRIDDesign of High Performance MVAPICH2: MPI2 over InfiniBand.Wei Huang, Gopalakrishnan Santhanaraman, Hyun-Wook Jin, Qi Gao, Dhabaleswar K. Panda
2006CCGRIDDesigning Efficient Cooperative Caching Schemes for Multi-Tier Data-Centers over RDMA-enabled Networks.Sundeep Narravula, Hyun-Wook Jin, Karthikeyan Vaidyanathan, Dhabaleswar K. Panda
2006CLUSTERDesigning High Performance and Scalable MPI Intra-node Communication Support for Clusters.Lei Chai, Albert Hartono, Dhabaleswar K. Panda
2006CLUSTERExploiting RDMA operations for Providing Efficient Fine-Grained Resource Monitoring in Cluster-based Servers.Karthikeyan Vaidyanathan, Hyun-Wook Jin, Dhabaleswar K. Panda
2006HiPCDDSS: A Low-Overhead Distributed Data Sharing Substrate for Cluster-Based Data-Centers over Modern Interconnects.Karthikeyan Vaidyanathan, Sundeep Narravula, Dhabaleswar K. Panda
2006HOTIMemory Scalability Evaluation of the Next-Generation Intel Bensley Platform with InfiniBand.Matthew J. Koop, Wei Huang, Abhinav Vishnu, Dhabaleswar K. Panda
2006ICCCNNemC: A Network Emulator for Cluster-of-Clusters.Hyun-Wook Jin, Sundeep Narravula, Karthikeyan Vaidyanathan, Dhabaleswar K. Panda
2006ICPPApplication-Transparent Checkpoint/Restart for MPI Programs over InfiniBand.Qi Gao, Weikuan Yu, Wei Huang, Dhabaleswar K. Panda
2006ICPPHigh Performance Block I/O for Global File System (GFS) with InfiniBand RDMA.Shuang Liang, Weikuan Yu, Dhabaleswar K. Panda
2006ICSA case for high performance computing with virtual machines.Wei Huang, Jiuxing Liu, Blent Abali, Dhabaleswar K. Panda
2006PPoPPRDMA read based rendezvous protocol for MPI over InfiniBand: design alternatives and benefits.Sayantan Sur, Hyun-Wook Jin, Lei Chai, Dhabaleswar K. Panda
2006SCPanel: Data intensive computing.Leslie S. Perkins, Phil Andrews, Dhabaleswar K. Panda, Dave Morton, Ron Bonica, Nick Henry Werstiuk, Randy Kreiser
2005CCGRIDArchitecture for caching responses with multiple dynamic dependencies in multi-tier data-centers over InfiniBand.Sundeep Narravula, Pavan Balaji, Karthikeyan Vaidyanathan, Hyun-Wook Jin, Dhabaleswar K. Panda
2005CCGRIDCan high performance software DSM systems designed with InfiniBand features benefit from PCI-Express?Ranjit Noronha, Dhabaleswar K. Panda
2005CLUSTERHead-to-TOE Evaluation of High-Performance Sockets over Protocol Offload Engines.Pavan Balaji, Wu-chun Feng, Qi Gao, Ranjit Noronha, Weikuan Yu, Dhabaleswar K. Panda
2005CLUSTERSupporting iWARP Compatibility and Features for Regular Network Adapters.Pavan Balaji, Hyun-Wook Jin, Karthikeyan Vaidyanathan, Dhabaleswar K. Panda
2005CLUSTERSwapping to Remote Memory over InfiniBand: An Approach using a High Performance Network Block Device.Shuang Liang, Ranjit Noronha, Dhabaleswar K. Panda
2005EuroParPerformance Evaluation of MM5 on Clusters with Modern Interconnects: Scalability and Impact.Ranjit Noronha, Dhabaleswar K. Panda
2005HiPCHigh Performance RDMA Based All-to-All Broadcast for InfiniBand Clusters.Sayantan Sur, Uday Bondhugula, Amith R. Mamidala, Hyun-Wook Jin, Dhabaleswar K. Panda
2005HiPCSupporting MPI-2 One Sided Communication on Multi-rail InfiniBand Clusters: Design Challenges and Performance Benefits.Abhinav Vishnu, Gopalakrishnan Santhanaraman, Wei Huang, Hyun-Wook Jin, Dhabaleswar K. Panda
2005HOTIPerformance Characterization of a 10-Gigabit Ethernet TOE.Wu-chun Feng, Pavan Balaji, Christopher Baron, Laxmi N. Bhuyan, Dhabaleswar K. Panda
2005HOTICan Memory-Less Network Adapters Benefit Next-Generation InfiniBand Systems?.Sayantan Sur, Abhinav Vishnu, Hyun-Wook Jin, Wei Huang, Dhabaleswar K. Panda
2005ICPPLiMIC: Support for High-Performance MPI Intra-node Communication on Linux Cluster.Hyun-Wook Jin, Sayantan Sur, Lei Chai, Dhabaleswar K. Panda
2005ICSHigh performance support of parallel virtual file system (PVFS2) over Quadrics.Weikuan Yu, Shuang Liang, Dhabaleswar K. Panda
2005ISPASSOn the provision of prioritization and soft qos in dynamically reconfigurable shared data-centers over infiniband.Pavan Balaji, Sundeep Narravula, Karthikeyan Vaidyanathan, Hyun-Wook Jin, Dhabaleswar K. Panda
2004CCGRIDHigh performance MPI-2 one-sided communication over InfiniBand.Weihang Jiang, Jiuxing Liu, Hyun-Wook Jin, Dhabaleswar K. Panda, William Gropp, Rajeev Thakur
2004CCGRIDDesigning high performance DSM systems using InfiniBand features.Ranjit Noronha, Dhabaleswar K. Panda
2004CCGRIDUnifier: unifying cache management and communication buffer management for PVFS over InfiniBand.Jiesheng Wu, Pete Wyckoff, Dhabaleswar K. Panda, Robert B. Ross
2004CLUSTERTowards provision of quality of service guarantees in job scheduling.Mohammad Islam, Pavan Balaji, P. Sadayappan, Dhabaleswar K. Panda
2004CLUSTEREfficient Barrier and Allreduce on Infiniband clusters using multicast and adaptive algorithms.Amith R. Mamidala, Jiuxing Liu, Dhabaleswar K. Panda
2004CLUSTERState of InfiniBand in designing HPC clusters, storage/file systems, and datacenters [datacenters read as data centers].Dhabaleswar K. Panda
2004CLUSTERNIC-based offload of dynamic user-defined modules for Myrinet clusters.Adam Wagner, Hyun-Wook Jin, Dhabaleswar K. Panda, Rolf Riesen
2004CLUSTERScalable, high-performance NIC-based all-to-all broadcast over Myrinet/GM.Weikuan Yu, Dhabaleswar K. Panda, Darius Buntinas
2004HiPCFast and Scalable Startup of MPI Programs in InfiniBand Clusters.Weikuan Yu, Jiesheng Wu, Dhabaleswar K. Panda
2004HOTIPerformance evaluation of InfiniBand with PCI Express.Jiuxing Liu, Amith R. Mamidala, Abhinav Vishnu, Dhabaleswar K. Panda
2004ICPPEfficient and Scalable All-to-All Personalized Exchange for InfiniBand-Based Clusters.Sayantan Sur, Hyun-Wook Jin, Dhabaleswar K. Panda
2004ISPASSSockets Direct Protocol over InfiniBand in clusters: is it beneficial?Pavan Balaji, Sundeep Narravula, Karthikeyan Vaidyanathan, Savitha Krishnamoorthy, Jiesheng Wu, Dhabaleswar K. Panda
2004SCBuilding Multirail InfiniBand Clusters: MPI-Level Design and Performance Evaluation.Jiuxing Liu, Abhinav Vishnu, Dhabaleswar K. Panda
2003CCGRIDApplication-Bypas Broadcast in MPICH over GM.Darius Buntinas, Dhabaleswar K. Panda, Ron Brightwell
2003CLUSTEROptimizing Mechanisms for Latency Tolerance in Remote Memory Access Communication on Clusters.Jarek Nieplocha, Vinod Tipparaju, Manojkumar Krishnan, Gopalakrishnan Santhanaraman, Dhabaleswar K. Panda
2003CLUSTERDesigning Next Generation Clusters with Infiniband: Opportunities and Challenges.Dhabaleswar K. Panda
2003CLUSTERApplication-Bypass Reduction for Large-Scale Clusters.Adam Wagner, Darius Buntinas, Dhabaleswar K. Panda, Ron Brightwell
2003CLUSTERSupporting Efficient Noncontiguous Access in PVFS over InfiniBand.Jiesheng Wu, Pete Wyckoff, Dhabaleswar K. Panda
2003HiPCExploiting Non-blocking Remote Memory Access Communication in Scientific Benchmarks.Vinod Tipparaju, Manojkumar Krishnan, Jarek Nieplocha, Gopalakrishnan Santhanaraman, Dhabaleswar K. Panda
2003HOTIMicro-benchmark level performance comparison of high-speed cluster interconnects.Jiuxing Liu, Balasubramanian Chandrasekaran, Weikuan Yu, Jiesheng Wu, Darius Buntinas, Sushmitha P. Kini, Pete Wyckoff, Dhabaleswar K. Panda
2003HPDCImpact of High Performance Sockets on Data Intensive Applications.Pavan Balaji, Jiesheng Wu, Tahsin M. Kur, mit V. atalyrek, Dhabaleswar K. Panda, Joel H. Saltz
2003HPDCQoS-Aware Middleware for Cluster-Based Servers to support Interactive and Resource-Adaptive Applications.S. Senapathi, B. Chandrasekaran, Don Stredney, Han-Wei Shen, Dhabaleswar K. Panda
2003ICPPPVFS over InfiniBand: Design and Performance Evaluation.Jiesheng Wu, Pete Wyckoff, Dhabaleswar K. Panda
2003ICPPHigh Performance and Reliable NIC-Based Multicast over Myrinet/GM-2.Weikuan Yu, Darius Buntinas, Dhabaleswar K. Panda
2003ICSHigh performance RDMA-based MPI implementation over InfiniBand.Jiuxing Liu, Jiesheng Wu, Sushmitha P. Kini, Pete Wyckoff, Dhabaleswar K. Panda
2003JSSPPQoPS: A QoS Based Scheme for Parallel Job Scheduling.Mohammad Islam, Pavan Balaji, P. Sadayappan, Dhabaleswar K. Panda
2003KDDTowards NIC-based intrusion detection.Matthew Eric Otey, Srinivasan Parthasarathy, Amol Ghoting, G. Li, Sundeep Narravula, Dhabaleswar K. Panda
2003SCPerformance Comparison of MPI Implementations over InfiniBand, Myrinet and Quadrics.Jiuxing Liu, B. Chandrasekaran, Jiesheng Wu, Weihang Jiang, Sushmitha P. Kini, Weikuan Yu, Darius Buntinas, Pete Wyckoff, Dhabaleswar K. Panda
2003SCScalable NIC-based Reduction on Large-scale Clusters.Adam Moody, Juan Fernndez, Fabrizio Petrini, Dhabaleswar K. Panda
2002CLUSTERHigh Performance User Level Sockets over Gigabit Ethernet.Pavan Balaji, Piyush Shivam, Pete Wyckoff, Dhabaleswar K. Panda
2002CLUSTEREfficient Barrier Using Remote Memory Operations on VIA-Based Clusters.Rinku Gupta, Vinod Tipparaju, Jarek Nieplocha, Dhabaleswar K. Panda
2002CLUSTERImpact of On-Demand Connection Management in MPI over VIA.Jiesheng Wu, Jiuxing Liu, Pete Wyckoff, Dhabaleswar K. Panda
2002HOTITutorial 2: InfiniBand Architecture and Where it is Headed.Dhabaleswar K. Panda
2002ICDCSA Reliable Multicast Algorithm for Mobile Ad Hoc Networks.Thiagaraja Gopalsamy, Mukesh Singhal, Dhabaleswar K. Panda, P. Sadayappan
2002LCNActive Network Interface: Opportunities and Challenges.Dhabaleswar K. Panda
2001ICPPImplementing TreadMarksover VIA on Myrinet and Gigabit Ethernet: Challenges, Design Experience, and Performance Evaluation.Mohammad Banikazemi, Jiuxing Liu, Dhabaleswar K. Panda, P. Sadayappan
2001ICPPNIC-Based Rate Control for Proportional Bandwidth Allocation in Myrinet Clusters.Abhishek Gulati, Dhabaleswar K. Panda, P. Sadayappan, Pete Wyckoff
2001SCEMP: zero-copy OS-bypass NIC-driven gigabit ethernet message passing.Piyush Shivam, Pete Wyckoff, Dhabaleswar K. Panda
2000HiPCCan Scatter Communication Take Advantage of Multidestination Message Passing?Mohammad Banikazemi, Dhabaleswar K. Panda
2000HiPCCharacterization and enhancement of Static Mapping Heuristics for Heterogeneous Systems.Praveen Holenarsipur, Vladimir Yarmolenko, Jos Duato, Dhabaleswar K. Panda, P. Sadayappan
1999HCWCommunication Modeling of Heterogeneous Networks of Workstations for Performance Characterization of Collective Operations.Mohammad Banikazemi, Jayanthi Sampathkumar, Sandeep Prabhu, Dhabaleswar K. Panda, P. Sadayappan
1998ICPPEfficient Collective Communication on Heterogeneous Networks of Workstations.Mohammad Banikazemi, Vijay Moorthy, Dhabaleswar K. Panda
1998ICPPImpact of Adaptivity on the Behaviour of Networks of Workstations under Bursty Traffic.Federico Silla, Manuel P. Malumbres, Jos Duato, Donglai Dai, Dhabaleswar K. Panda
1998ICPPWhere to Provide Support for Efficient Multicasting in Irregular Networks: Network Interface or Switch?Rajeev Sivaram, Ram Kesavan, Dhabaleswar K. Panda, Craig B. Stunkel
1997HiPCPrioritized demand multiplexing (PDM): a low-latency virtual channel flow control framework for prioritized traffic.Abdel-Halim Smai, Dhabaleswar K. Panda, Lars-Erik Thorelli
1997HPCAMulticast on Irregular Switch-Based Networks with Wormhole Routing.Ram Kesavan, Kiran Bondalapati, Dhabaleswar K. Panda
1997ICPPHow Much Does Network Contention Affect Distributed Shared Memory Performance?Donglai Dai, Dhabaleswar K. Panda
1997ICPPOptimal Multicast with Packetization and Network Interface Support.Ram Kesavan, Dhabaleswar K. Panda
1997ISCAImplementing Multidestination Worms in Switch-Based Parallel Systems: Architectural Alternatives and their Impact.Craig B. Stunkel, Rajeev Sivaram, Dhabaleswar K. Panda
1997WSCSimulation of Modern Parallel Systems: A CSIM-based Approach.Dhabaleswar K. Panda, Debashis Basak, Donglai Dai, Ram Kesavan, Rajeev Sivaram, Mohammad Banikazemi, Vijay Moorthy
1996ICPPDesigning Processor-Cluster Based Systems: Interplay Between Organizations and Broadcasting Algorithms.Debashis Basak, Dhabaleswar K. Panda
1996ICPPReducing Cache Invalidation Overheads in Wormhole Routed DSMs Using Multidestination Message Passing.Donglai Dai, Dhabaleswar K. Panda
1996ICPPMinimizing Node Contention in Multiple Multicast on Wormhole k-ary N-Cube Networks.Ram Kesavan, Dhabaleswar K. Panda
1996ICSHybrid Algorithms for Complete Exchange in 2D Meshes.N. S. Sundar, Doddaballapur Narasimha-Murthy Jayasimha, Dhabaleswar K. Panda, P. Sadayappan
1995HPCAFast Barrier Synchronization in Wormhole k-ary n-cube Networks with Multidestination Worms.Dhabaleswar K. Panda
1994ICPPDesigning Large Hierarchical Multiprocessor Systems under Processor, Interconnection, and Packaging Advancements.Debashis Basak, Dhabaleswar K. Panda
1991ICPPMessage Vectorization for Converting Multicomputer Programs to Shared-Memory Multiprocessors.Dhabaleswar K. Panda, Kai Hwang
1990ICPPAlgorithm-Driven Simulation and Performance Projection of a RISC-based Orthogonal Multiprocessor.Sharad Mehrotra, Chien-Ming Cheng, Kai Hwang, Michel Dubois, Dhabaleswar K. Panda
1990ICSOMP: a RISC-based multiprocessor using orthogonal-access memories and multiple spanning buses.Kai Hwang, Michel Dubois, Dhabaleswar K. Panda, S. Rao, Shisheng Shang, Aydin resin, W. Mao, H. Nair, M. Lytwyn, F. Hsieh, J. Liu, Sharad Mehrotra, Chien-Ming Cheng
1989ARITHOptical arithmetic using high-radix symbolic substitution rules.Kai Hwang, Dhabaleswar K. Panda