John D. Owens
Publication record assembled from the DBLP archive of ranked conferences.
Papers indexed
55
Venues
22
Active years
1998–2026
Best venue rank
A*
Where they publish
- BPPoPP9 papers
- A*SIGGRAPH8 papers
- BEuroPar5 papers
- ASC4 papers
- A*ASPLOS3 papers
- A*HPCA3 papers
- CICCD3 papers
- BSPAA2 papers
- NationalHiPC2 papers
- BICPP2 papers
- AHPDC2 papers
- A*MICRO2 papers
- BISPASS1 paper
- BACCV1 paper
- AWACV1 paper
- CIPAS1 paper
- CASRU1 paper
- BSSDBM1 paper
- CICCSA1 paper
- AICS1 paper
- BDCOSS1 paper
- A*ISCA1 paper
Papers
55 indexed papers, newest first.
| Year | Venue | Title | Authors |
|---|---|---|---|
| 2026 | SIGGRAPH | Fast Sparse Matrix Permutation for Mesh-Based Direct Solvers. | Behrooz Zarebavani, Ahmed H. Mahmoud, Ana Dodik, Changcheng Yuan, Serban D. Porumbescu, John D. Owens, Maryam Mehri Dehnavi, Justin Solomon |
| 2025 | SPAA | Decoupled Fallback: A Portable Single-Pass GPU Scan. | Thomas Smith, Raph Levien, John D. Owens |
| 2024 | SC | Accelerating Multi-GPU Embedding Retrieval with PGAS-Style Communication for Deep Learning Recommendation Systems. | Yuxin Chen, Aydin Bulu, Katherine A. Yelick, John D. Owens |
| 2023 | ASPLOS | Accelerating Sparse Data Orchestration via Dynamic Reflexive Tiling. | Toluwanimi O. Odemuyiwa, Hadi Asghari Moghaddam, Michael Pellauer, Kartik Hegde, Po-An Tsai, Neal Clayton Crago, Aamer Jaleel, John D. Owens, Edgar Solomonik, Joel S. Emer, Christopher W. Fletcher |
| 2023 | PPoPP | Stream-K: Work-Centric Parallel Decomposition for Dense Matrix-Matrix Multiplication on the GPU. | Muhammad Osama, Duane Merrill, Cris Cecka, Michael Garland, John D. Owens |
| 2023 | PPoPP | A Programming Model for GPU Load Balancing. | Muhammad Osama, Serban D. Porumbescu, John D. Owens |
| 2023 | PPoPP | Harmonic CUDA: Asynchronous Programming on GPUs. | Jonathan D. Wapman, Sean Treichler, Serban D. Porumbescu, John D. Owens |
| 2022 | HiPC | Building a Performance Model for Deep Learning Recommendation Model Training on GPUs. | Zhongyi Lin, Louis Feng, Ehsan K. Ardestani, Jaewon Lee, John Lundell, Changkyu Kim, Arun Kejariwal, John D. Owens |
| 2022 | ICPP | Atos: A Task-Parallel GPU Scheduler for Graph Analytics. | Yuxin Chen, Benjamin Brock, Serban D. Porumbescu, Aydin Bulu, Katherine A. Yelick, John D. Owens |
| 2022 | ISPASS | Building a Performance Model for Deep Learning Recommendation Model Training on GPUs. | Zhongyi Lin, Louis Feng, Ehsan K. Ardestani, Jaewon Lee, John Lundell, Changkyu Kim, Arun Kejariwal, John D. Owens |
| 2022 | SC | Scalable Irregular Parallelism with GPUs: Getting CPUs Out of the Way. | Yuxin Chen, Benjamin Brock, Serban D. Porumbescu, Aydin Bulu, Katherine A. Yelick, John D. Owens |
| 2021 | EuroPar | Towards Flexible and Compiler-Friendly Layer Fusion for CNNs on Multicore CPUs. | Zhongyi Lin, Evangelos Georganas, John D. Owens |
| 2019 | PPoPP | Engineering a high-performance GPU B-Tree. | Muhammad A. Awad, Saman Ashkiani, Rob Johnson, Martin Farach-Colton, John D. Owens |
| 2019 | SC | RDMA vs. RPC for Implementing Distributed Data Structures. | Benjamin A. Brock, Yuxin Chen, Jiakun Yan, John D. Owens, Aydin Bulu, Katherine A. Yelick |
| 2018 | EuroPar | Design Principles for Sparse Matrix Multiplication on the GPU. | Carl Yang, Aydin Bulu, John D. Owens |
| 2018 | ICPP | Implementing Push-Pull Efficiently in GraphBLAS. | Carl Yang, Aydin Bulu, John D. Owens |
| 2016 | HPDC | A Comparative Study on Exact Triangle Counting Algorithms on the GPU. | Leyuan Wang, Yangzihao Wang, Carl Yang, John D. Owens |
| 2016 | PPoPP | GPU multisplit. | Saman Ashkiani, Andrew A. Davidson, Ulrich Meyer, John D. Owens |
| 2016 | PPoPP | Multitasking Real-time Embedded GPU Computing Tasks. | Pinar Muyan-zelik, John D. Owens |
| 2016 | PPoPP | Gunrock: a high-performance graph processing library on the GPU. | Yangzihao Wang, Andrew A. Davidson, Yuechao Pan, Yuduo Wu, Andy Riffel, John D. Owens |
| 2016 | SPAA | Parallel Approaches to the String Matching Problem on the GPU. | Saman Ashkiani, Nina Amenta, John D. Owens |
| 2015 | EuroPar | Fast Parallel Suffix Array on the GPU. | Leyuan Wang, Sean Baxter, John D. Owens |
| 2015 | PPoPP | Gunrock: a high-performance graph processing library on the GPU. | Yangzihao Wang, Andrew A. Davidson, Yuechao Pan, Yuduo Wu, Andy Riffel, John D. Owens |
| 2014 | ACCV | A Comparative Study of GPU-Accelerated Multi-view Sequential Reconstruction Triangulation Methods for Large-Scale Scenes. | Jason Mak, Mauricio Hess-Flores, Shawn Recker, John D. Owens, Kenneth I. Joy |
| 2014 | WACV | GPU-accelerated and efficient multi-view triangulation for scene reconstruction. | Jason Mak, Mauricio Hess-Flores, Shawn Recker, John D. Owens, Kenneth I. Joy |
| 2012 | IPAS | Plane-dependent error diffusion on a GPU. | Yao Zhang, John Recker, Robert Ulichney, Ingeborg Tastl, John D. Owens |
| 2011 | ASPLOS | Register packing for cyclic reduction: a case study. | Andrew A. Davidson, John D. Owens |
| 2011 | EuroPar | Lessons Learned from Exploring the Backtracking Paradigm on the GPU. | John Jenkins, Isha Arkatkar, John D. Owens, Alok N. Choudhary, Nagiza F. Samatova |
| 2011 | HiPC | Compute & memory optimizations for high-quality speech recognition on low-end GPU processors. | Kshitij Gupta, John D. Owens |
| 2011 | HPCA | A quantitative performance analysis model for GPU architectures. | Yao Zhang, John D. Owens |
| 2010 | EuroPar | GPU-to-CPU Callbacks. | Jeff A. Stuart, Michael Cox, John D. Owens |
| 2010 | HPDC | Multi-GPU volume rendering using MapReduce. | Jeff A. Stuart, Cheng-Kai Chen, Kwan-Liu Ma, John D. Owens |
| 2010 | PPoPP | Fast tridiagonal solvers on the GPU. | Yao Zhang, Jonathan Cohen, John D. Owens |
| 2009 | ASRU | Three-layer optimizations for fast GMM computations on GPU-like parallel processors. | Kshitij Gupta, John D. Owens |
| 2009 | SSDBM | Data Parallel Bin-Based Indexing for Answering Queries on Multi-core Architectures. | Luke J. Gosink, Kesheng Wu, E. Wes Bethel, John D. Owens, Kenneth I. Joy |
| 2008 | ICCSA | Fast Deformable Registration on the GPU: A CUDA Implementation of Demons. | Pinar Muyan-zelik, John D. Owens, Junyi Xia, Sanjiv S. Samant |
| 2008 | ICS | Efficient computation of sum-products on GPUs through software-managed cache. | Mark Silberstein, Assaf Schuster, Dan Geiger, Anjul Patney, John D. Owens |
| 2008 | SIGGRAPH | Beyond programmable shading: fundamentals. | Aaron E. Lefohn, Mike Houston, Chas Boyd, Kayvon Fatahalian, Tom Forsyth, David Luebke, John D. Owens |
| 2008 | SIGGRAPH | Parallel programming models overview. | John D. Owens |
| 2007 | SIGGRAPH | GPU architecture overview. | John D. Owens |
| 2007 | SIGGRAPH | Data-parallel algorithms and data structures. | John D. Owens |
| 2006 | DCOSS | The Virtual Pheromone Communication Primitive. | Leo Szumel, John D. Owens |
| 2006 | SC | S07 - GPGPU: general-purpose computation on graphics hardware. | David P. Luebke, Mark J. Harris, Naga K. Govindaraju, Aaron E. Lefohn, Mike Houston, John D. Owens, Mark Segal, Matthew Papakipos, Ian Buck |
| 2005 | SIGGRAPH | Octree textures on graphics hardware. | Joe Kniss, Aaron E. Lefohn, Robert Strzodka, Shubhabrata Sengupta, John D. Owens |
| 2005 | SIGGRAPH | Dynamic adaptive shadow maps on graphics hardware. | Aaron E. Lefohn, Shubhabrata Sengupta, Joe Kniss, Robert Strzodka, John D. Owens |
| 2005 | SIGGRAPH | Streaming architectures and technology trends. | John D. Owens |
| 2003 | HPCA | Exploring the VLSI Scalability of Stream Processors. | Brucek Khailany, William J. Dally, Scott Rixner, Ujval J. Kapasi, John D. Owens, Brian Towles |
| 2002 | ICCD | The Imagine Stream Processor. | Ujval J. Kapasi, William J. Dally, Scott Rixner, John D. Owens, Brucek Khailany |
| 2002 | ICCD | Media Processing Applications on the Imagine Stream Processor. | John D. Owens, Scott Rixner, Ujval J. Kapasi, Peter R. Mattson, Brian Towles, Ben Serebrin, William J. Dally |
| 2002 | ICCD | A Stream Processor Development Platform. | Ben Serebrin, John D. Owens, Chen H. Chen, Stephen P. Crago, Ujval J. Kapasi, Peter R. Mattson, Jinyung Namkoong, Scott Rixner, William J. Dally |
| 2000 | ASPLOS | Communication Scheduling. | Peter R. Mattson, William J. Dally, Scott Rixner, Ujval J. Kapasi, John D. Owens |
| 2000 | HPCA | Register Organization for Media Processing. | Scott Rixner, William J. Dally, Brucek Khailany, Peter R. Mattson, Ujval J. Kapasi, John D. Owens |
| 2000 | ISCA | Memory access scheduling. | Scott Rixner, William J. Dally, Ujval J. Kapasi, Peter R. Mattson, John D. Owens |
| 2000 | MICRO | Efficient conditional operations for data-parallel architectures. | Ujval J. Kapasi, William J. Dally, Scott Rixner, Peter R. Mattson, John D. Owens, Brucek Khailany |
| 1998 | MICRO | A Bandwidth-efficient Architecture for Media Processing. | Scott Rixner, William J. Dally, Ujval J. Kapasi, Brucek Khailany, Abelardo Lpez-Lagunas, Peter R. Mattson, John D. Owens |