Skip to content

John D. Owens

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

55

Venues

22

Active years

1998–2026

Best venue rank

A*

Where they publish

Papers

55 indexed papers, newest first.

YearVenueTitleAuthors
2026SIGGRAPHFast Sparse Matrix Permutation for Mesh-Based Direct Solvers.Behrooz Zarebavani, Ahmed H. Mahmoud, Ana Dodik, Changcheng Yuan, Serban D. Porumbescu, John D. Owens, Maryam Mehri Dehnavi, Justin Solomon
2025SPAADecoupled Fallback: A Portable Single-Pass GPU Scan.Thomas Smith, Raph Levien, John D. Owens
2024SCAccelerating Multi-GPU Embedding Retrieval with PGAS-Style Communication for Deep Learning Recommendation Systems.Yuxin Chen, Aydin Bulu, Katherine A. Yelick, John D. Owens
2023ASPLOSAccelerating Sparse Data Orchestration via Dynamic Reflexive Tiling.Toluwanimi O. Odemuyiwa, Hadi Asghari Moghaddam, Michael Pellauer, Kartik Hegde, Po-An Tsai, Neal Clayton Crago, Aamer Jaleel, John D. Owens, Edgar Solomonik, Joel S. Emer, Christopher W. Fletcher
2023PPoPPStream-K: Work-Centric Parallel Decomposition for Dense Matrix-Matrix Multiplication on the GPU.Muhammad Osama, Duane Merrill, Cris Cecka, Michael Garland, John D. Owens
2023PPoPPA Programming Model for GPU Load Balancing.Muhammad Osama, Serban D. Porumbescu, John D. Owens
2023PPoPPHarmonic CUDA: Asynchronous Programming on GPUs.Jonathan D. Wapman, Sean Treichler, Serban D. Porumbescu, John D. Owens
2022HiPCBuilding a Performance Model for Deep Learning Recommendation Model Training on GPUs.Zhongyi Lin, Louis Feng, Ehsan K. Ardestani, Jaewon Lee, John Lundell, Changkyu Kim, Arun Kejariwal, John D. Owens
2022ICPPAtos: A Task-Parallel GPU Scheduler for Graph Analytics.Yuxin Chen, Benjamin Brock, Serban D. Porumbescu, Aydin Bulu, Katherine A. Yelick, John D. Owens
2022ISPASSBuilding a Performance Model for Deep Learning Recommendation Model Training on GPUs.Zhongyi Lin, Louis Feng, Ehsan K. Ardestani, Jaewon Lee, John Lundell, Changkyu Kim, Arun Kejariwal, John D. Owens
2022SCScalable Irregular Parallelism with GPUs: Getting CPUs Out of the Way.Yuxin Chen, Benjamin Brock, Serban D. Porumbescu, Aydin Bulu, Katherine A. Yelick, John D. Owens
2021EuroParTowards Flexible and Compiler-Friendly Layer Fusion for CNNs on Multicore CPUs.Zhongyi Lin, Evangelos Georganas, John D. Owens
2019PPoPPEngineering a high-performance GPU B-Tree.Muhammad A. Awad, Saman Ashkiani, Rob Johnson, Martin Farach-Colton, John D. Owens
2019SCRDMA vs. RPC for Implementing Distributed Data Structures.Benjamin A. Brock, Yuxin Chen, Jiakun Yan, John D. Owens, Aydin Bulu, Katherine A. Yelick
2018EuroParDesign Principles for Sparse Matrix Multiplication on the GPU.Carl Yang, Aydin Bulu, John D. Owens
2018ICPPImplementing Push-Pull Efficiently in GraphBLAS.Carl Yang, Aydin Bulu, John D. Owens
2016HPDCA Comparative Study on Exact Triangle Counting Algorithms on the GPU.Leyuan Wang, Yangzihao Wang, Carl Yang, John D. Owens
2016PPoPPGPU multisplit.Saman Ashkiani, Andrew A. Davidson, Ulrich Meyer, John D. Owens
2016PPoPPMultitasking Real-time Embedded GPU Computing Tasks.Pinar Muyan-zelik, John D. Owens
2016PPoPPGunrock: a high-performance graph processing library on the GPU.Yangzihao Wang, Andrew A. Davidson, Yuechao Pan, Yuduo Wu, Andy Riffel, John D. Owens
2016SPAAParallel Approaches to the String Matching Problem on the GPU.Saman Ashkiani, Nina Amenta, John D. Owens
2015EuroParFast Parallel Suffix Array on the GPU.Leyuan Wang, Sean Baxter, John D. Owens
2015PPoPPGunrock: a high-performance graph processing library on the GPU.Yangzihao Wang, Andrew A. Davidson, Yuechao Pan, Yuduo Wu, Andy Riffel, John D. Owens
2014ACCVA Comparative Study of GPU-Accelerated Multi-view Sequential Reconstruction Triangulation Methods for Large-Scale Scenes.Jason Mak, Mauricio Hess-Flores, Shawn Recker, John D. Owens, Kenneth I. Joy
2014WACVGPU-accelerated and efficient multi-view triangulation for scene reconstruction.Jason Mak, Mauricio Hess-Flores, Shawn Recker, John D. Owens, Kenneth I. Joy
2012IPASPlane-dependent error diffusion on a GPU.Yao Zhang, John Recker, Robert Ulichney, Ingeborg Tastl, John D. Owens
2011ASPLOSRegister packing for cyclic reduction: a case study.Andrew A. Davidson, John D. Owens
2011EuroParLessons Learned from Exploring the Backtracking Paradigm on the GPU.John Jenkins, Isha Arkatkar, John D. Owens, Alok N. Choudhary, Nagiza F. Samatova
2011HiPCCompute & memory optimizations for high-quality speech recognition on low-end GPU processors.Kshitij Gupta, John D. Owens
2011HPCAA quantitative performance analysis model for GPU architectures.Yao Zhang, John D. Owens
2010EuroParGPU-to-CPU Callbacks.Jeff A. Stuart, Michael Cox, John D. Owens
2010HPDCMulti-GPU volume rendering using MapReduce.Jeff A. Stuart, Cheng-Kai Chen, Kwan-Liu Ma, John D. Owens
2010PPoPPFast tridiagonal solvers on the GPU.Yao Zhang, Jonathan Cohen, John D. Owens
2009ASRUThree-layer optimizations for fast GMM computations on GPU-like parallel processors.Kshitij Gupta, John D. Owens
2009SSDBMData Parallel Bin-Based Indexing for Answering Queries on Multi-core Architectures.Luke J. Gosink, Kesheng Wu, E. Wes Bethel, John D. Owens, Kenneth I. Joy
2008ICCSAFast Deformable Registration on the GPU: A CUDA Implementation of Demons.Pinar Muyan-zelik, John D. Owens, Junyi Xia, Sanjiv S. Samant
2008ICSEfficient computation of sum-products on GPUs through software-managed cache.Mark Silberstein, Assaf Schuster, Dan Geiger, Anjul Patney, John D. Owens
2008SIGGRAPHBeyond programmable shading: fundamentals.Aaron E. Lefohn, Mike Houston, Chas Boyd, Kayvon Fatahalian, Tom Forsyth, David Luebke, John D. Owens
2008SIGGRAPHParallel programming models overview.John D. Owens
2007SIGGRAPHGPU architecture overview.John D. Owens
2007SIGGRAPHData-parallel algorithms and data structures.John D. Owens
2006DCOSSThe Virtual Pheromone Communication Primitive.Leo Szumel, John D. Owens
2006SCS07 - GPGPU: general-purpose computation on graphics hardware.David P. Luebke, Mark J. Harris, Naga K. Govindaraju, Aaron E. Lefohn, Mike Houston, John D. Owens, Mark Segal, Matthew Papakipos, Ian Buck
2005SIGGRAPHOctree textures on graphics hardware.Joe Kniss, Aaron E. Lefohn, Robert Strzodka, Shubhabrata Sengupta, John D. Owens
2005SIGGRAPHDynamic adaptive shadow maps on graphics hardware.Aaron E. Lefohn, Shubhabrata Sengupta, Joe Kniss, Robert Strzodka, John D. Owens
2005SIGGRAPHStreaming architectures and technology trends.John D. Owens
2003HPCAExploring the VLSI Scalability of Stream Processors.Brucek Khailany, William J. Dally, Scott Rixner, Ujval J. Kapasi, John D. Owens, Brian Towles
2002ICCDThe Imagine Stream Processor.Ujval J. Kapasi, William J. Dally, Scott Rixner, John D. Owens, Brucek Khailany
2002ICCDMedia Processing Applications on the Imagine Stream Processor.John D. Owens, Scott Rixner, Ujval J. Kapasi, Peter R. Mattson, Brian Towles, Ben Serebrin, William J. Dally
2002ICCDA Stream Processor Development Platform.Ben Serebrin, John D. Owens, Chen H. Chen, Stephen P. Crago, Ujval J. Kapasi, Peter R. Mattson, Jinyung Namkoong, Scott Rixner, William J. Dally
2000ASPLOSCommunication Scheduling.Peter R. Mattson, William J. Dally, Scott Rixner, Ujval J. Kapasi, John D. Owens
2000HPCARegister Organization for Media Processing.Scott Rixner, William J. Dally, Brucek Khailany, Peter R. Mattson, Ujval J. Kapasi, John D. Owens
2000ISCAMemory access scheduling.Scott Rixner, William J. Dally, Ujval J. Kapasi, Peter R. Mattson, John D. Owens
2000MICROEfficient conditional operations for data-parallel architectures.Ujval J. Kapasi, William J. Dally, Scott Rixner, Peter R. Mattson, John D. Owens, Brucek Khailany
1998MICROA Bandwidth-efficient Architecture for Media Processing.Scott Rixner, William J. Dally, Ujval J. Kapasi, Brucek Khailany, Abelardo Lpez-Lagunas, Peter R. Mattson, John D. Owens