Skip to content

Khaled Hamidouche

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

39

Venues

14

Active years

2011–2026

Best venue rank

A*

Where they publish

Papers

39 indexed papers, newest first.

YearVenueTitleAuthors
2026CCGRIDKernel-Initiated One-Sided Networking for GPU-Accelerated AI Workloads.Khaled Hamidouche, John Bachan, Pak Markthub, Peter-Jan Gootzen, Elena Agostini, Sylvain Jeaugey, Aamir Shafi, Georgios Theodorakis, Manjunath Gorentla Venkata
2024SCOptimizing Distributed ML Communication with Fused Computation-Collective Operations.Kishore Punniyamurthy, Khaled Hamidouche, Bradford M. Beckmann
2023ISCAA Research Retrospective on AMD's Exascale Computing Journey.Gabriel H. Loh, Michael J. Schulte, Mike Ignatowski, Vignesh Adhinarayanan, Shaizeen Aga, Derrick Aguren, Varun Agrawal, Ashwin M. Aji, Johnathan Alsop, Paul T. Bauman, Bradford M. Beckmann, Majed Valad Beigi, Sergey Blagodurov, Travis Boraten, Michael Boyer, William C. Brantley, Noel Chalmers, Shaoming Chen, Kevin Cheng, Michael L. Chu, David Cownie, Nicholas Curtis, Joris Del Pino, Nam Duong, Alexandru Dutu, Yasuko Eckert, Christopher Erb, Chip Freitag, Joseph L. Greathouse, Sudhanva Gurumurthi, Anthony Gutierrez, Khaled Hamidouche, Sachin Hossamani, Wei Huang, Mahzabeen Islam, Nuwan Jayasena, John Kalamatianos, Onur Kayiran, Jagadish Kotra, Alan Lee, Daniel Lowell, Niti Madan, Abhinandan Majumdar, Nicholas Malaya, Srilatha Manne, Susumu Mashimo, Damon McDougall, Elliot Mednick, Michael Mishkin, Mark Nutter, Indrani Paul, Matthew Poremba, Brandon Potter, Kishore Punniyamurthy, Sooraj Puthoor, Steven E. Raasch, Karthik Rao, Gregory Rodgers, Marko Scrbak, Mohammad Seyedzadeh, John Slice, Vilas Sridharan, Ren van Oostrum, Eric Van Tassell, Abhinav Vishnu, Samuel Wasmundt, Mark Wilkening, Noah Wolfe, Mark Wyse, Adithya Yalavarti, Dmitri Yudanov
2020PPoPP<u>G</u>PU <u>i</u>nitiated <u>O</u>penSHMEM: correct and efficient intra-kernel networking for dGPUs.Khaled Hamidouche, Michael LeBeane
2017HiPCKernel-Assisted Communication Engine for MPI on Emerging Manycore Processors.Jahanzeb Maqbool Hashmi, Khaled Hamidouche, Hari Subramoni, Dhabaleswar K. Panda
2017ICPPMPI-GDS: High Performance MPI Designs with GPUDirect-aSync for CPU-GPU Control Flow Decoupling.Akshay Venkatesh, Khaled Hamidouche, Sreeram Potluri, Davide Rossetti, Ching-Hsiang Chu, Dhabaleswar K. Panda
2017PPoPPS-Caffe: Co-designing MPI Runtimes and Caffe for Scalable Deep Learning on Modern GPU Clusters.Ammar Ahmad Awan, Khaled Hamidouche, Jahanzeb Maqbool Hashmi, Dhabaleswar K. Panda
2017SCGPU triggered networking for intra-kernel communications.Michael LeBeane, Khaled Hamidouche, Brad Benton, Maurcio Breternitz, Steven K. Reinhardt, Lizy K. John
2016CCGRIDCUDA Kernel Based Collective Reduction Operations on Large-scale GPU Clusters.Ching-Hsiang Chu, Khaled Hamidouche, Akshay Venkatesh, Ammar Ahmad Awan, Dhabaleswar K. Panda
2016CloudComRe-Designing CNTK Deep Learning Framework on Modern GPU Enabled Clusters.Dip Sankar Banerjee, Khaled Hamidouche, Dhabaleswar K. Panda
2016HiPCCUDA M3: Designing Efficient CUDA Managed Memory-Aware MPI by Exploiting GDR and IPC.Khaled Hamidouche, Ammar Ahmad Awan, Akshay Venkatesh, Dhabaleswar K. Panda
2016HiPCMizan-RMA: Accelerating Mizan Graph Processing Framework with MPI RMA.Mingzhe Li, Xiaoyi Lu, Khaled Hamidouche, Jie Zhang, Dhabaleswar K. Panda
2016HPCCEnabling Performance Efficient Runtime Support for Hybrid MPI+UPC++ Programming Models.Jahanzeb Maqbool Hashmi, Khaled Hamidouche, Dhabaleswar K. Panda
2016PPoPPDesigning high performance communication runtime for GPU managed memory: early experiences.Dip Sankar Banerjee, Khaled Hamidouche, Dhabaleswar K. Panda
2016SCEfficient Reliability Support for Hardware Multicast-Based Broadcast in GPU-enabled Streaming Applications.Ching-Hsiang Chu, Khaled Hamidouche, Hari Subramoni, Akshay Venkatesh, Bracy Elton, Dhabaleswar K. Panda
2016SCOpenSHMEM Non-blocking Data Movement Operations with MVAPICH2-X: Early Experiences.Khaled Hamidouche, Jie Zhang, Dhabaleswar K. Panda, Karen Tomko
2016SCDesigning MPI library with on-demand paging (ODP) of infiniband: challenges and benefits.Mingzhe Li, Khaled Hamidouche, Xiaoyi Lu, Hari Subramoni, Jie Zhang, Dhabaleswar K. Panda
2016SBAC-PADDesigning High Performance Heterogeneous Broadcast for Streaming Applications on GPU Clusters.Ching-Hsiang Chu, Khaled Hamidouche, Hari Subramoni, Akshay Venkatesh, Bracy Elton, Dhabaleswar K. Panda
2015CCGRIDPower-Check: An Energy-Efficient Checkpointing Framework for HPC Clusters.Raghunath Raja Chandrasekar, Akshay Venkatesh, Khaled Hamidouche, Dhabaleswar K. Panda
2015CLUSTERExploiting GPUDirect RDMA in Designing High Performance OpenSHMEM for NVIDIA GPU Clusters.Khaled Hamidouche, Akshay Venkatesh, Ammar Ahmad Awan, Hari Subramoni, Ching-Hsiang Chu, Dhabaleswar K. Panda
2015CLUSTERHigh Performance MPI Datatype Support with User-Mode Memory Registration: Challenges, Designs, and Benefits.Mingzhe Li, Hari Subramoni, Khaled Hamidouche, Xiaoyi Lu, Dhabaleswar K. Panda
2015EuroParHigh-Performance and Scalable Design of MPI-3 RMA on Xeon Phi Clusters.Mingzhe Li, Khaled Hamidouche, Xiaoyi Lu, Jian Lin, Dhabaleswar K. Panda
2015HiPCHigh Performance OpenSHMEM Strided Communication Support with InfiniBand UMR.Mingzhe Li, Khaled Hamidouche, Xiaoyi Lu, Jie Zhang, Jian Lin, Dhabaleswar K. Panda
2015HiPCOffloaded GPU Collectives Using CORE-Direct and CUDA Capabilities on InfiniBand Clusters.Akshay Venkatesh, Khaled Hamidouche, Hari Subramoni, Dhabaleswar K. Panda
2015HOTIImpact of InfiniBand DC Transport Protocol on Energy Consumption of All-to-All Collective Algorithms.Hari Subramoni, Akshay Venkatesh, Khaled Hamidouche, Karen Tomko, Dhabaleswar K. Panda
2015SCA case for application-oblivious energy-efficient MPI runtime.Akshay Venkatesh, Abhinav Vishnu, Khaled Hamidouche, Nathan R. Tallent, Dhabaleswar K. Panda, Darren J. Kerbyson, Adolfy Hoisie
2014CLUSTERHigh performance OpenSHMEM for Xeon Phi clusters: Extensions, runtime designs and application co-design.Jithin Jose, Khaled Hamidouche, Xiaoyi Lu, Sreeram Potluri, Jie Zhang, Karen Tomko, Dhabaleswar K. Panda
2014CLUSTERScalable Graph500 design with MPI-3 RMA.Mingzhe Li, Xiaoyi Lu, Sreeram Potluri, Khaled Hamidouche, Jithin Jose, Karen Tomko, Dhabaleswar K. Panda
2014HiPCDesigning efficient small message transfer mechanism for inter-node MPI communication on InfiniBand GPU clusters.Rong Shi, Sreeram Potluri, Khaled Hamidouche, Jonathan L. Perkins, Mingzhe Li, Davide Rossetti, Dhabaleswar K. Panda
2014HiPCA high performance broadcast design with hardware multicast and GPUDirect RDMA for streaming applications on Infiniband clusters.Akshay Venkatesh, Hari Subramoni, Khaled Hamidouche, Dhabaleswar K. Panda
2014HPDCMIC-Check: a distributed check pointing framework for the intel many integrated cores architecture.Raghunath Rajachandrasekar, Sreeram Potluri, Akshay Venkatesh, Khaled Hamidouche, Md. Wasi-ur-Rahman, Dhabaleswar K. Panda
2014ICPPHAND: A Hybrid Approach to Accelerate Non-contiguous Data Movement Using MPI Datatypes on GPU Clusters.Rong Shi, Xiaoyi Lu, Sreeram Potluri, Khaled Hamidouche, Jie Zhang, Dhabaleswar K. Panda
2014PPoPPInitial study of multi-endpoint runtime for MPI+OpenMP hybrid programming model on multi-core systems.Miao Luo, Xiaoyi Lu, Khaled Hamidouche, Krishna Chaitanya Kandalla, Dhabaleswar K. Panda
2013CLUSTERA scalable and portable approach to accelerate hybrid HPL on heterogeneous CPU-GPU clusters.Rong Shi, Sreeram Potluri, Khaled Hamidouche, Xiaoyi Lu, Karen Tomko, Dhabaleswar K. Panda
2013HOTIDesigning Optimized MPI Broadcast and Allreduce for Many Integrated Core (MIC) InfiniBand Clusters.Krishna Chaitanya Kandalla, Akshay Venkatesh, Khaled Hamidouche, Sreeram Potluri, Devendar Bureddy, Dhabaleswar K. Panda
2013ICPPEfficient Inter-node MPI Communication Using GPUDirect RDMA for InfiniBand Clusters with NVIDIA GPUs.Sreeram Potluri, Khaled Hamidouche, Akshay Venkatesh, Devendar Bureddy, Dhabaleswar K. Panda
2013ICSMIC-RO: enabling efficient remote offload on heterogeneous many integrated core (MIC) clusters with InfiniBand.Khaled Hamidouche, Sreeram Potluri, Hari Subramoni, Krishna Chaitanya Kandalla, Dhabaleswar K. Panda
2013SCMVAPICH-PRISM: a proxy-based communication framework using InfiniBand and SCIF for intel MIC clusters.Sreeram Potluri, Devendar Bureddy, Khaled Hamidouche, Akshay Venkatesh, Krishna Chaitanya Kandalla, Hari Subramoni, Dhabaleswar K. Panda
2011SBAC-PADParallel Biological Sequence Comparison on Heterogeneous High Performance Computing Platforms with BSP++.Khaled Hamidouche, Fernando Machado Mendonca, Jol Falcou, Daniel Etiemble