Skip to content

John M. Mellor-Crummey

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

72

Venues

16

Active years

1989–2025

Best venue rank

A*

Where they publish

Papers

72 indexed papers, newest first.

YearVenueTitleAuthors
2025ICSAnalyzing the Performance of Applications at Exascale.Dragana Grbic, John M. Mellor-Crummey
2024SCMatrix-Free Finite Volume Kernels on a Dataflow Architecture.Ryuichi Sai, Franois P. Hamon, John M. Mellor-Crummey, Mauricio Araya-Polo
2024SCAutomated Code Generation of High-Order Stencils for a Dataflow Architecture.Ryuichi Sai, John M. Mellor-Crummey, Jinfan Xu, Mauricio Araya-Polo
2022ASPLOSValueExpert: exploring value patterns in GPU-accelerated applications.Keren Zhou, Yueming Hao, John M. Mellor-Crummey, Xiaozhu Meng, Xu Liu
2022ICSPreparing for performance analysis at exascale.Jonathon M. Anderson, Yumeng Liu, John M. Mellor-Crummey
2022ICSLow overhead and context sensitive profiling of CPU-accelerated applications.Keren Zhou, Jonathon M. Anderson, Xiaozhu Meng, John M. Mellor-Crummey
2021CGOGPA: A GPU Performance Advisor Based on Instruction Sampling.Keren Zhou, Xiaozhu Meng, Ryuichi Sai, John M. Mellor-Crummey
2021PPoPPParallel binary code analysis.Xiaozhu Meng, Jonathon M. Anderson, John M. Mellor-Crummey, Mark W. Krentel, Barton P. Miller, Srdan Milakovic
2021SCMeasurement and Analysis of GPU-Accelerated OpenCL Computations on Intel GPUs.Aaron Cherian, Keren Zhou, Dejan Grubisic, Xiaozhu Meng, John M. Mellor-Crummey
2020ICSTools for top-down performance analysis of GPU-accelerated applications.Keren Zhou, Mark W. Krentel, John M. Mellor-Crummey
2020PPoPPUsing sample-based time series data for automated diagnosis of scalability losses in parallel programs.Lai Wei, John M. Mellor-Crummey
2020PPoPPA tool for top-down performance analysis of GPU-accelerated applications.Keren Zhou, Mark Krentel, John M. Mellor-Crummey
2020SCGVProf: a value profiler for GPU-based clusters.Keren Zhou, Yueming Hao, John M. Mellor-Crummey, Xiaozhu Meng, Xu Liu
2019CGOA Tool for Performance Analysis of GPU-Accelerated Applications.Keren Zhou, John M. Mellor-Crummey
2019HOTILightweight, Packet-Centric Monitoring of Network Traffic and Congestion Implemented in P4.Philip Taffet, John M. Mellor-Crummey
2019SCUnderstanding congestion in high performance interconnection networks using sampling.Philip Taffet, John M. Mellor-Crummey
2018ICSAutomated Analysis of Time Series Data to Understand Parallel Program Behaviors.Lai Wei, John M. Mellor-Crummey
2018SCDynamic data race detection for OpenMP programs.Yizi Gu, John M. Mellor-Crummey
2016EuroParDesign and Verification of Distributed Phasers.Karthik Murthy, Sri Raj Paul, Kuldeep S. Meel, Tiago Cogumbreiro, John M. Mellor-Crummey
2016ICCSPerformance Analysis and Optimization of a Hybrid Seismic Imaging Application.Sri Raj Paul, Mauricio Araya-Polo, John M. Mellor-Crummey, Detlef Hohl
2016PPoPPContention-conscious, locality-preserving locks.Milind Chabbi, John M. Mellor-Crummey
2016PPoPPA wait-free queue as fast as fetch-and-add.Chaoran Yang, John M. Mellor-Crummey
2016SPAAA Practical Solution to the Cactus Stack Problem.Chaoran Yang, John M. Mellor-Crummey
2015PPoPPHigh performance locks for multi-level NUMA systems.Milind Chabbi, Michael W. Fagan, John M. Mellor-Crummey
2015PPoPPBarrier elision for production parallel programs.Milind Chabbi, Wim Lavrijsen, Wibe de Jong, Koushik Sen, John M. Mellor-Crummey, Costin Iancu
2014CGOCall Paths for Pin Tools.Milind Chabbi, Xu Liu, John M. Mellor-Crummey
2014ICSAuthor retrospective: compilation techniques for block-cyclic distributions.John M. Mellor-Crummey, Seema Hiranandani, Ajay Sethi
2014PLDITest-driven repair of data races in structured parallel programs.Rishi Surendran, Raghavan Raman, Swarat Chaudhuri, John M. Mellor-Crummey, Vivek Sarkar
2014PPoPPA tool to analyze the performance of multithreaded programs on NUMA architectures.Xu Liu, John M. Mellor-Crummey
2014PPoPPPortable, MPI-interoperable coarray fortran.Chaoran Yang, Wesley Bland, John M. Mellor-Crummey, Pavan Balaji
2013HPDCOn the efficacy of GPU-integrated MPI for scientific applications.Ashwin M. Aji, Lokendra S. Panwar, Feng Ji, Milind Chabbi, Karthik Murthy, Pavan Balaji, Keith R. Bisset, James Dinan, Wu-chun Feng, John M. Mellor-Crummey, Xiaosong Ma, Rajeev Thakur
2013ICSA new approach for performance analysis of openMP programs.Xu Liu, John M. Mellor-Crummey, Michael W. Fagan
2013ISPASSPinpointing data locality bottlenecks with low overhead.Xu Liu, John M. Mellor-Crummey
2013SCEffective sampling-driven performance tools for GPU-accelerated supercomputers.Milind Chabbi, Karthik Murthy, Michael W. Fagan, John M. Mellor-Crummey
2013SCA data-centric profiler for parallel programs.Xu Liu, John M. Mellor-Crummey
2012CGODeadSpy: a tool to pinpoint program inefficiencies.Milind Chabbi, John M. Mellor-Crummey
2011CGOPinpointing data locality problems using data-centric analysis.Xu Liu, John M. Mellor-Crummey
2011ICSScalable fine-grained call path tracing.Nathan R. Tallent, John M. Mellor-Crummey, Michael Franco, Reed Landrum, Laksono Adhianto
2010PPoPPAnalyzing lock contention in multithreaded applications.Nathan R. Tallent, John M. Mellor-Crummey, Allan Porterfield
2010SCScalable Identification of Load Imbalance in Parallel Executions Using Call Path Profiles.Nathan R. Tallent, Laksono Adhianto, John M. Mellor-Crummey
2009PLDIBinary analysis for measurement and attribution of program performance.Nathan R. Tallent, John M. Mellor-Crummey, Michael W. Fagan
2009PPoPPEffective performance measurement and analysis of multithreaded applications.Nathan R. Tallent, John M. Mellor-Crummey
2009SCDiagnosing performance bottlenecks in emerging petascale applications.Nathan R. Tallent, John M. Mellor-Crummey, Laksono Adhianto, Michael W. Fagan, Mark Krentel
2008ISPASSPinpointing and Exploiting Opportunities for Enhancing Data Reuse.Gabriel Marin, John M. Mellor-Crummey
2008PPoPPWhere will all the threads come from?John M. Mellor-Crummey
2007IPCCCApplication Insight Through Performance Modeling.Gabriel Marin, John M. Mellor-Crummey
2007ICSScalability analysis of SPMD codes using expectations.Cristian Coarfa, John M. Mellor-Crummey, Nathan Froyd, Yuri Dotsenko
2005HPDCScheduling strategies for mapping application workflows onto the grid.Anirban Mandal, Ken Kennedy, Charles Koelbel, Gabriel Marin, John M. Mellor-Crummey, Bo Liu, S. Lennart Johnsson
2005ICPADSPRec-I-DCM3: A Parallel Framework for Fast and Accurate Large Scale Phylogeny Reconstruction.Cristian Coarfa, Yuri Dotsenko, John M. Mellor-Crummey, Luay Nakhleh, Usman Roshan
2005ICSLow-overhead call path profiling of unmodified, optimized code.Nathan Froyd, John M. Mellor-Crummey, Robert J. Fowler
2005PPoPPEffective communication coalescing for data-parallel applications.Daniel G. Chavarra-Miranda, John M. Mellor-Crummey
2005PPoPPAn evaluation of global address space languages: co-array fortran and unified parallel C.Cristian Coarfa, Yuri Dotsenko, John M. Mellor-Crummey, Franois Cantonnet, Tarek A. El-Ghazawi, Ashrujit Mohanti, Yiyi Yao, Daniel G. Chavarra-Miranda
2004CCGRIDScheduling workflow applications in GrADS.Anirban Mandal, Anshuman Dasgupta, Ken Kennedy, Mark Mazina, Charles Koelbel, Gabriel Marin, Keith D. Cooper, John M. Mellor-Crummey, Bo Liu, S. Lennart Johnsson
2004SIGMETRICSCross-architecture performance predictions for scientific applications using parameterized models.Gabriel Marin, John M. Mellor-Crummey
2002ICSExperiences tuning SMG98: a semicoarsening multigrid benchmark based on the hypre library.Guohua Jin, John M. Mellor-Crummey
2001EuroParData-Parallel Compiler Support for Multipartitioning.Daniel G. Chavarra-Miranda, John M. Mellor-Crummey, Trushar Sarang
2001ICSTools for application-oriented performance tuning.John M. Mellor-Crummey, Robert J. Fowler, David B. Whalley
2001SCIncreasing temporal locality with skewing and recursive blocking.Guohua Jin, John M. Mellor-Crummey, Robert J. Fowler
2001SIGMETRICSOn providing useful information for analyzing and tuning applications.John M. Mellor-Crummey, Robert J. Fowler, David B. Whalley
1999ICSImproving memory hierarchy performance for irregular applications.John M. Mellor-Crummey, David B. Whalley, Ken Kennedy
1999PPoPPAn Evaluation of Computing Paradigms for N-Body Simulations on Distributed Memory Architectures.Collin McCurdy, John M. Mellor-Crummey
1998PLDIUsing Integer Sets for Data-Parallel Program Analysis and Optimization.Vikram S. Adve, John M. Mellor-Crummey
1998SCHigh Performance Fortran Compilation Techniques for Parallelizing Scientific Codes.Vikram S. Adve, Guohua Jin, John M. Mellor-Crummey, Qing Yi
1997SCCompiling Stencils in High Performance Fortran.Gerald Roth, John M. Mellor-Crummey, Ken Kennedy, R. Gregg Brickner
1995SCAn Integrated Compilation and Performance Analysis Environment for Data Parallel Programs.Vikram S. Adve, John M. Mellor-Crummey, Mark Anderson, Ken Kennedy, Jhy-Chun Wang, Daniel A. Reed
1994ICSCompilation techniques for block-cyclic distributions.Seema Hiranandani, Ken Kennedy, John M. Mellor-Crummey, Ajay Sethi
1992ICSAutomatic software cache coherence through vectorization.Ervan Darnell, John M. Mellor-Crummey, Ken Kennedy
1991ASPLOSSynchronization without Contention.John M. Mellor-Crummey, Michael L. Scott
1991PPoPPScalable Reader-Writer Synchronization for Shared-Memory Multiprocessors.John M. Mellor-Crummey, Michael L. Scott
1991SCOn-the-fly detection of data races for programs with nested fork-join parallelism.John M. Mellor-Crummey
1990SCParallel program debugging with on-the-fly anomaly detection.Robert Hood, Ken Kennedy, John M. Mellor-Crummey
1989ASPLOSA Software Instruction Counter.John M. Mellor-Crummey, Thomas J. LeBlanc