Khaled Hamidouche
Publication record assembled from the DBLP archive of ranked conferences.
Papers indexed
39
Venues
14
Active years
2011–2026
Best venue rank
A*
Where they publish
Papers
39 indexed papers, newest first.
| Year | Venue | Title | Authors |
|---|---|---|---|
| 2026 | CCGRID | Kernel-Initiated One-Sided Networking for GPU-Accelerated AI Workloads. | Khaled Hamidouche, John Bachan, Pak Markthub, Peter-Jan Gootzen, Elena Agostini, Sylvain Jeaugey, Aamir Shafi, Georgios Theodorakis, Manjunath Gorentla Venkata |
| 2024 | SC | Optimizing Distributed ML Communication with Fused Computation-Collective Operations. | Kishore Punniyamurthy, Khaled Hamidouche, Bradford M. Beckmann |
| 2023 | ISCA | A Research Retrospective on AMD's Exascale Computing Journey. | Gabriel H. Loh, Michael J. Schulte, Mike Ignatowski, Vignesh Adhinarayanan, Shaizeen Aga, Derrick Aguren, Varun Agrawal, Ashwin M. Aji, Johnathan Alsop, Paul T. Bauman, Bradford M. Beckmann, Majed Valad Beigi, Sergey Blagodurov, Travis Boraten, Michael Boyer, William C. Brantley, Noel Chalmers, Shaoming Chen, Kevin Cheng, Michael L. Chu, David Cownie, Nicholas Curtis, Joris Del Pino, Nam Duong, Alexandru Dutu, Yasuko Eckert, Christopher Erb, Chip Freitag, Joseph L. Greathouse, Sudhanva Gurumurthi, Anthony Gutierrez, Khaled Hamidouche, Sachin Hossamani, Wei Huang, Mahzabeen Islam, Nuwan Jayasena, John Kalamatianos, Onur Kayiran, Jagadish Kotra, Alan Lee, Daniel Lowell, Niti Madan, Abhinandan Majumdar, Nicholas Malaya, Srilatha Manne, Susumu Mashimo, Damon McDougall, Elliot Mednick, Michael Mishkin, Mark Nutter, Indrani Paul, Matthew Poremba, Brandon Potter, Kishore Punniyamurthy, Sooraj Puthoor, Steven E. Raasch, Karthik Rao, Gregory Rodgers, Marko Scrbak, Mohammad Seyedzadeh, John Slice, Vilas Sridharan, Ren van Oostrum, Eric Van Tassell, Abhinav Vishnu, Samuel Wasmundt, Mark Wilkening, Noah Wolfe, Mark Wyse, Adithya Yalavarti, Dmitri Yudanov |
| 2020 | PPoPP | <u>G</u>PU <u>i</u>nitiated <u>O</u>penSHMEM: correct and efficient intra-kernel networking for dGPUs. | Khaled Hamidouche, Michael LeBeane |
| 2017 | HiPC | Kernel-Assisted Communication Engine for MPI on Emerging Manycore Processors. | Jahanzeb Maqbool Hashmi, Khaled Hamidouche, Hari Subramoni, Dhabaleswar K. Panda |
| 2017 | ICPP | MPI-GDS: High Performance MPI Designs with GPUDirect-aSync for CPU-GPU Control Flow Decoupling. | Akshay Venkatesh, Khaled Hamidouche, Sreeram Potluri, Davide Rossetti, Ching-Hsiang Chu, Dhabaleswar K. Panda |
| 2017 | PPoPP | S-Caffe: Co-designing MPI Runtimes and Caffe for Scalable Deep Learning on Modern GPU Clusters. | Ammar Ahmad Awan, Khaled Hamidouche, Jahanzeb Maqbool Hashmi, Dhabaleswar K. Panda |
| 2017 | SC | GPU triggered networking for intra-kernel communications. | Michael LeBeane, Khaled Hamidouche, Brad Benton, Maurcio Breternitz, Steven K. Reinhardt, Lizy K. John |
| 2016 | CCGRID | CUDA Kernel Based Collective Reduction Operations on Large-scale GPU Clusters. | Ching-Hsiang Chu, Khaled Hamidouche, Akshay Venkatesh, Ammar Ahmad Awan, Dhabaleswar K. Panda |
| 2016 | CloudCom | Re-Designing CNTK Deep Learning Framework on Modern GPU Enabled Clusters. | Dip Sankar Banerjee, Khaled Hamidouche, Dhabaleswar K. Panda |
| 2016 | HiPC | CUDA M3: Designing Efficient CUDA Managed Memory-Aware MPI by Exploiting GDR and IPC. | Khaled Hamidouche, Ammar Ahmad Awan, Akshay Venkatesh, Dhabaleswar K. Panda |
| 2016 | HiPC | Mizan-RMA: Accelerating Mizan Graph Processing Framework with MPI RMA. | Mingzhe Li, Xiaoyi Lu, Khaled Hamidouche, Jie Zhang, Dhabaleswar K. Panda |
| 2016 | HPCC | Enabling Performance Efficient Runtime Support for Hybrid MPI+UPC++ Programming Models. | Jahanzeb Maqbool Hashmi, Khaled Hamidouche, Dhabaleswar K. Panda |
| 2016 | PPoPP | Designing high performance communication runtime for GPU managed memory: early experiences. | Dip Sankar Banerjee, Khaled Hamidouche, Dhabaleswar K. Panda |
| 2016 | SC | Efficient Reliability Support for Hardware Multicast-Based Broadcast in GPU-enabled Streaming Applications. | Ching-Hsiang Chu, Khaled Hamidouche, Hari Subramoni, Akshay Venkatesh, Bracy Elton, Dhabaleswar K. Panda |
| 2016 | SC | OpenSHMEM Non-blocking Data Movement Operations with MVAPICH2-X: Early Experiences. | Khaled Hamidouche, Jie Zhang, Dhabaleswar K. Panda, Karen Tomko |
| 2016 | SC | Designing MPI library with on-demand paging (ODP) of infiniband: challenges and benefits. | Mingzhe Li, Khaled Hamidouche, Xiaoyi Lu, Hari Subramoni, Jie Zhang, Dhabaleswar K. Panda |
| 2016 | SBAC-PAD | Designing High Performance Heterogeneous Broadcast for Streaming Applications on GPU Clusters. | Ching-Hsiang Chu, Khaled Hamidouche, Hari Subramoni, Akshay Venkatesh, Bracy Elton, Dhabaleswar K. Panda |
| 2015 | CCGRID | Power-Check: An Energy-Efficient Checkpointing Framework for HPC Clusters. | Raghunath Raja Chandrasekar, Akshay Venkatesh, Khaled Hamidouche, Dhabaleswar K. Panda |
| 2015 | CLUSTER | Exploiting GPUDirect RDMA in Designing High Performance OpenSHMEM for NVIDIA GPU Clusters. | Khaled Hamidouche, Akshay Venkatesh, Ammar Ahmad Awan, Hari Subramoni, Ching-Hsiang Chu, Dhabaleswar K. Panda |
| 2015 | CLUSTER | High Performance MPI Datatype Support with User-Mode Memory Registration: Challenges, Designs, and Benefits. | Mingzhe Li, Hari Subramoni, Khaled Hamidouche, Xiaoyi Lu, Dhabaleswar K. Panda |
| 2015 | EuroPar | High-Performance and Scalable Design of MPI-3 RMA on Xeon Phi Clusters. | Mingzhe Li, Khaled Hamidouche, Xiaoyi Lu, Jian Lin, Dhabaleswar K. Panda |
| 2015 | HiPC | High Performance OpenSHMEM Strided Communication Support with InfiniBand UMR. | Mingzhe Li, Khaled Hamidouche, Xiaoyi Lu, Jie Zhang, Jian Lin, Dhabaleswar K. Panda |
| 2015 | HiPC | Offloaded GPU Collectives Using CORE-Direct and CUDA Capabilities on InfiniBand Clusters. | Akshay Venkatesh, Khaled Hamidouche, Hari Subramoni, Dhabaleswar K. Panda |
| 2015 | HOTI | Impact of InfiniBand DC Transport Protocol on Energy Consumption of All-to-All Collective Algorithms. | Hari Subramoni, Akshay Venkatesh, Khaled Hamidouche, Karen Tomko, Dhabaleswar K. Panda |
| 2015 | SC | A case for application-oblivious energy-efficient MPI runtime. | Akshay Venkatesh, Abhinav Vishnu, Khaled Hamidouche, Nathan R. Tallent, Dhabaleswar K. Panda, Darren J. Kerbyson, Adolfy Hoisie |
| 2014 | CLUSTER | High performance OpenSHMEM for Xeon Phi clusters: Extensions, runtime designs and application co-design. | Jithin Jose, Khaled Hamidouche, Xiaoyi Lu, Sreeram Potluri, Jie Zhang, Karen Tomko, Dhabaleswar K. Panda |
| 2014 | CLUSTER | Scalable Graph500 design with MPI-3 RMA. | Mingzhe Li, Xiaoyi Lu, Sreeram Potluri, Khaled Hamidouche, Jithin Jose, Karen Tomko, Dhabaleswar K. Panda |
| 2014 | HiPC | Designing efficient small message transfer mechanism for inter-node MPI communication on InfiniBand GPU clusters. | Rong Shi, Sreeram Potluri, Khaled Hamidouche, Jonathan L. Perkins, Mingzhe Li, Davide Rossetti, Dhabaleswar K. Panda |
| 2014 | HiPC | A high performance broadcast design with hardware multicast and GPUDirect RDMA for streaming applications on Infiniband clusters. | Akshay Venkatesh, Hari Subramoni, Khaled Hamidouche, Dhabaleswar K. Panda |
| 2014 | HPDC | MIC-Check: a distributed check pointing framework for the intel many integrated cores architecture. | Raghunath Rajachandrasekar, Sreeram Potluri, Akshay Venkatesh, Khaled Hamidouche, Md. Wasi-ur-Rahman, Dhabaleswar K. Panda |
| 2014 | ICPP | HAND: A Hybrid Approach to Accelerate Non-contiguous Data Movement Using MPI Datatypes on GPU Clusters. | Rong Shi, Xiaoyi Lu, Sreeram Potluri, Khaled Hamidouche, Jie Zhang, Dhabaleswar K. Panda |
| 2014 | PPoPP | Initial study of multi-endpoint runtime for MPI+OpenMP hybrid programming model on multi-core systems. | Miao Luo, Xiaoyi Lu, Khaled Hamidouche, Krishna Chaitanya Kandalla, Dhabaleswar K. Panda |
| 2013 | CLUSTER | A scalable and portable approach to accelerate hybrid HPL on heterogeneous CPU-GPU clusters. | Rong Shi, Sreeram Potluri, Khaled Hamidouche, Xiaoyi Lu, Karen Tomko, Dhabaleswar K. Panda |
| 2013 | HOTI | Designing Optimized MPI Broadcast and Allreduce for Many Integrated Core (MIC) InfiniBand Clusters. | Krishna Chaitanya Kandalla, Akshay Venkatesh, Khaled Hamidouche, Sreeram Potluri, Devendar Bureddy, Dhabaleswar K. Panda |
| 2013 | ICPP | Efficient Inter-node MPI Communication Using GPUDirect RDMA for InfiniBand Clusters with NVIDIA GPUs. | Sreeram Potluri, Khaled Hamidouche, Akshay Venkatesh, Devendar Bureddy, Dhabaleswar K. Panda |
| 2013 | ICS | MIC-RO: enabling efficient remote offload on heterogeneous many integrated core (MIC) clusters with InfiniBand. | Khaled Hamidouche, Sreeram Potluri, Hari Subramoni, Krishna Chaitanya Kandalla, Dhabaleswar K. Panda |
| 2013 | SC | MVAPICH-PRISM: a proxy-based communication framework using InfiniBand and SCIF for intel MIC clusters. | Sreeram Potluri, Devendar Bureddy, Khaled Hamidouche, Akshay Venkatesh, Krishna Chaitanya Kandalla, Hari Subramoni, Dhabaleswar K. Panda |
| 2011 | SBAC-PAD | Parallel Biological Sequence Comparison on Heterogeneous High Performance Computing Platforms with BSP++. | Khaled Hamidouche, Fernando Machado Mendonca, Jol Falcou, Daniel Etiemble |