| 2024 | SC | HPC with Enhanced User Separation. | Andrew Prout, Albert Reuther, Michael Houle, Michael Jones, Peter Michaleas, LaToya Anderson, William Arcand, Bill Bergeron, David Bestor, Alex Bonn, Daniel Burrill, Chansup Byun, Vijay Gadepally, Matthew Hubbell, Hayden Jananthan, Piotr Luszczek, Lauren Milechin, Guillermo Morales, Julie Mullen, Antonio Rosa, Charles Yee, Jeremy Kepner |
| 2023 | ICS | Using Additive Modifications in LU Factorization Instead of Pivoting. | Neil Lindquist, Piotr Luszczek, Jack J. Dongarra |
| 2023 | SC | GPU-based LU Factorization and Solve on Batches of Matrices with Band Structure. | Ahmad Abdelfattah, Stanimire Tomov, Piotr Luszczek, Hartwig Anzt, Jack J. Dongarra |
| 2022 | SC | Threshold Pivoting for Dense LU Factorization. | Neil Lindquist, Mark Gates, Piotr Luszczek, Jack J. Dongarra |
| 2022 | SC | Mixed-Precision Algorithm for Finding Selected Eigenvalues and Eigenvectors of Symmetric and Hermitian Matrices | Yaohung M. Tsai, Piotr Luszczek, Jack J. Dongarra |
| 2021 | ICS | Task-graph scheduling extensions for efficient synchronization and communication. | Seonmyeong Bak, Oscar R. Hernandez, Mark Gates, Piotr Luszczek, Vivek Sarkar |
| 2016 | SC | Performance-Portable Autotuning of OpenCL Kernels for Convolutional Layers of Deep Neural Networks. | Yaohung M. Tsai, Piotr Luszczek, Jakub Kurzak, Jack J. Dongarra |
| 2015 | HPCC | Flexible Linear Algebra Development and Scheduling with Cholesky Factorization. | Azzam Haidar, Asim YarKhan, Chongxiao Cao, Piotr Luszczek, Stanimire Tomov, Jack J. Dongarra |
| 2015 | PPoPP | Towards batched linear solvers on accelerated hardware platforms. | Azzam Haidar, Tingxing Dong, Piotr Luszczek, Stanimire Tomov, Jack J. Dongarra |
| 2015 | PPoPP | Optimization for performance and energy for batched matrix computations on GPUs. | Azzam Haidar, Tingxing Dong, Piotr Luszczek, Stanimire Tomov, Jack J. Dongarra |
| 2015 | SC | Weighted dynamic scheduling with many parallelism grains for offloading of numerical workloads to multiple varied accelerators. | Azzam Haidar, Yulu Jia, Piotr Luszczek, Stanimire Tomov, Asim YarKhan, Jack J. Dongarra |
| 2015 | SC | Performance of random sampling for computing low-rank approximations of a dense matrix on GPUs. | Tho Mary, Ichitaro Yamazaki, Jakub Kurzak, Piotr Luszczek, Stanimire Tomov, Jack J. Dongarra |
| 2015 | SC | Randomized algorithms to update partial singular value decomposition on a hybrid CPU/GPU cluster. | Ichitaro Yamazaki, Jakub Kurzak, Piotr Luszczek, Jack J. Dongarra |
| 2014 | HPCC | LU Factorization of Small Matrices: Accelerating Batched DGETRF on the GPU. | Tingxing Dong, Azzam Haidar, Piotr Luszczek, James Austin Harris, Stanimire Tomov, Jack J. Dongarra |
| 2014 | ICPP | Parallel Simulation of Superscalar Scheduling. | Blake Haugen, Jakub Kurzak, Asim YarKhan, Piotr Luszczek, Jack J. Dongarra |
| 2014 | SC | Performance and portability with OpenCL for throughput-oriented HPC workloads across accelerators, coprocessors, and multicore processors. | Chongxiao Cao, Mark Gates, Azzam Haidar, Piotr Luszczek, Stanimire Tomov, Ichitaro Yamazaki, Jack J. Dongarra |
| 2013 | EuroPar | Implementing a Systolic Algorithm for QR Factorization on Multicore Clusters with PaRSEC. | Guillaume Aupy, Mathieu Faverge, Yves Robert, Jakub Kurzak, Piotr Luszczek, Jack J. Dongarra |
| 2013 | PPAM | Portable HPC Programming on Intel Many-Integrated-Core Hardware with MAGMA Port to Xeon Phi. | Jack J. Dongarra, Mark Gates, Azzam Haidar, Yulu Jia, Khairul Kabir, Piotr Luszczek, Stanimire Tomov |
| 2013 | SC | An improved parallel singular value algorithm and its implementation for multicore hardware. | Azzam Haidar, Jakub Kurzak, Piotr Luszczek |
| 2013 | SC | Parallel reduction to hessenberg form with algorithm-based fault tolerance. | Yulu Jia, George Bosilca, Piotr Luszczek, Jack J. Dongarra |
| 2013 | SC | CPU-GPU hybrid bidiagonal reduction with soft error resilience. | Yulu Jia, Piotr Luszczek, George Bosilca, Jack J. Dongarra |
| 2012 | EuroPar | GPU-Accelerated Asynchronous Error Correction for Mixed Precision Iterative Refinement. | Hartwig Anzt, Piotr Luszczek, Jack J. Dongarra, Vincent Heuveline |
| 2011 | CLUSTER | High Performance Dense Linear System Solver with Soft Error Resilience. | Peng Du, Piotr Luszczek, Jack J. Dongarra |
| 2011 | EuroPar | Evaluation of the HPC Challenge Benchmarks in Virtualized Environments. | Piotr Luszczek, Eric Meek, Shirley Moore, Daniel Terpstra, Vincent M. Weaver, Jack J. Dongarra |
| 2011 | PPAM | Enhancing Parallelism of Tile Bidiagonal Transformation on Multicore Architectures Using Tree Reduction. | Hatem Ltaief, Piotr Luszczek, Jack J. Dongarra |
| 2011 | PPAM | Reducing the Time to Tune Parallel Dense Linear Algebra Routines with Partial Execution and Performance Modeling. | Piotr Luszczek, Jack J. Dongarra |
| 2011 | SC | High performance matrix inversion based on LU factorization for multicore architectures. | Jack J. Dongarra, Mathieu Faverge, Hatem Ltaief, Piotr Luszczek |
| 2011 | SC | Soft error resilient QR factorization for hybrid system with GPGPU. | Peng Du, Piotr Luszczek, Stanimire Tomov, Jack J. Dongarra |
| 2006 | SC | Tools and techniques for performance - Exploiting the performance of 32 bit floating point arithmetic in obtaining 64 bit accuracy (revisiting iterative refinement for linear systems). | Julie Langou, Julien Langou, Piotr Luszczek, Jakub Kurzak, Alfredo Buttari, Jack J. Dongarra |
| 2006 | SC | S12 - The HPC Challenge (HPCC) benchmark suite. | Piotr Luszczek, David H. Bailey, Jack J. Dongarra, Jeremy Kepner, Robert F. Lucas, Rolf Rabenseifner, Daisuke Takahashi |
| 2004 | ICCS | Design of Interactive Environment for Numerically Intensive Parallel Linear Algebra Calculations. | Piotr Luszczek, Jack J. Dongarra |
| 2003 | ICCS | Self-Adapting Software for Numerical Linear Algebra Library Routines on Clusters. | Zizhong Chen, Jack J. Dongarra, Piotr Luszczek, Kenneth Roche |