| 2025 | SC | Efficient Embedding Initialization via Dominant Eigenvector Projections. | Quentin R. Petit, Chong Li, Nahid Emad, Jack J. Dongarra |
| 2024 | ICCS | Trends in Computational Science: Natural Language Processing and Network Analysis of 23 Years of ICCS Publications. | Lijing Luo, Sergey V. Kovalchuk, Valeria V. Krzhizhanovskaya, Maciej Paszynski, Cllia de Mulatier, Jack J. Dongarra, Peter M. A. Sloot |
| 2023 | ICS | Using Additive Modifications in LU Factorization Instead of Pivoting. | Neil Lindquist, Piotr Luszczek, Jack J. Dongarra |
| 2023 | SC | GPU-based LU Factorization and Solve on Batches of Matrices with Band Structure. | Ahmad Abdelfattah, Stanimire Tomov, Piotr Luszczek, Hartwig Anzt, Jack J. Dongarra |
| 2023 | SC | Task-Based Polar Decomposition Using SLATE on Massively Parallel Systems with Hardware Accelerators. | Dalal Sukkari, Mark Gates, Mohammed A. Al Farhan, Hartwig Anzt, Jack J. Dongarra |
| 2022 | CLUSTER | Lossy all-to-all exchange for accelerating parallel 3-D FFTs on hybrid architectures with GPUs. | Sbastien Cayrols, Jiali Li, George Bosilca, Stanimire Tomov, Alan Ayala, Jack J. Dongarra |
| 2022 | HPCC | Message from the High Performance Computing and Communications 2022 General Chairs. | Jack J. Dongarra, Kenli Li, Hai Jin |
| 2022 | ICCS | Batch QR Factorization on GPUs: Design, Optimization, and Tuning. | Ahmad Abdelfattah, Stan Tomov, Jack J. Dongarra |
| 2022 | SC | Addressing Irregular Patterns of Matrix Computations on GPUs and Their Impact on Applications Powered by Sparse Direct Solvers. | Ahmad Abdelfattah, Pieter Ghysels, Wajih Boukaram, Stanimire Tomov, Xiaoye Sherry Li, Jack J. Dongarra |
| 2022 | SC | Reshaping Geostatistical Modeling and Prediction for Extreme-Scale Environmental Applications. | Qinglei Cao, Sameh Abdulah, Rabab Alomairy, Yu Pei, Pratik Nag, George Bosilca, Jack J. Dongarra, Marc G. Genton, David E. Keyes, Hatem Ltaief, Ying Sun |
| 2022 | SC | Threshold Pivoting for Dense LU Factorization. | Neil Lindquist, Mark Gates, Piotr Luszczek, Jack J. Dongarra |
| 2022 | SC | Mixed-Precision Algorithm for Finding Selected Eigenvalues and Eigenvectors of Symmetric and Hermitian Matrices | Yaohung M. Tsai, Piotr Luszczek, Jack J. Dongarra |
| 2021 | PACT | Scalability Issues in FFT Computation. | Alan Ayala, Stanimire Tomov, Miroslav Stoyanov, Jack J. Dongarra |
| 2020 | CCGRID | Using Arm Scalable Vector Extension to Optimize OPEN MPI. | Dong Zhong, Pavel Shamis, Qinglei Cao, George Bosilca, Shinji Sumimoto, Kenichi Miura, Jack J. Dongarra |
| 2020 | CLUSTER | Flexible Data Redistribution in a Task-Based Runtime System. | Qinglei Cao, George Bosilca, Wei Wu, Dong Zhong, Aurelien Bouteiller, Jack J. Dongarra |
| 2020 | CLUSTER | HAN: a Hierarchical AutotuNed Collective Communication Framework. | Xi Luo, Wei Wu, George Bosilca, Yu Pei, Qinglei Cao, Thananon Patinyasakdikul, Dong Zhong, Jack J. Dongarra |
| 2020 | ICCS | Investigating the Benefit of FP16-Enabled Mixed-Precision Solvers for Symmetric Positive Definite Matrices Using GPUs. | Ahmad Abdelfattah, Stan Tomov, Jack J. Dongarra |
| 2020 | ICCS | heFFTe: Highly Efficient FFT for Exascale. | Alan Ayala, Stanimire Tomov, Azzam Haidar, Jack J. Dongarra |
| 2019 | EuroPar | Linear Systems Solvers for Distributed-Memory Machines with GPU Accelerators. | Jakub Kurzak, Mark Gates, Ali Charara, Asim YarKhan, Ichitaro Yamazaki, Jack J. Dongarra |
| 2019 | ICPP | Massively Parallel Automated Software Tuning. | Jakub Kurzak, Yaohung M. Tsai, Mark Gates, Ahmad Abdelfattah, Jack J. Dongarra |
| 2019 | ICS | Least squares solvers for distributed-memory machines with GPU accelerators. | Jakub Kurzak, Mark Gates, Ali Charara, Asim YarKhan, Jack J. Dongarra |
| 2019 | SC | Towards Half-Precision Computation for Complex Matrices: A Case Study for Mixed Precision Solvers on GPUs. | Ahmad Abdelfattah, Stanimire Tomov, Jack J. Dongarra |
| 2019 | SC | Performance Analysis of Tile Low-Rank Cholesky Factorization Using PaRSEC Instrumentation Tools. | Qinglei Cao, Yu Pei, Thomas Hrault, Kadir Akbudak, Aleksandr Mikhalev, George Bosilca, Hatem Ltaief, David E. Keyes, Jack J. Dongarra |
| 2019 | SC | SLATE: design of a modern distributed and accelerated linear algebra library. | Mark Gates, Jakub Kurzak, Ali Charara, Asim YarKhan, Jack J. Dongarra |
| 2019 | SC | Generic Matrix Multiplication for Multi-GPU Accelerated Distributed-Memory Platforms over PaRSEC. | Thomas Hrault, Yves Robert, George Bosilca, Jack J. Dongarra |
| 2019 | SC | Evaluation of Programming Models to Address Load Imbalance on Distributed Multi-Core CPUs: A Case Study with Block Low-Rank Factorization. | Yu Pei, George Bosilca, Ichitaro Yamazaki, Akihiro Ida, Jack J. Dongarra |
| 2018 | EuroPar | Do Moldable Applications Perform Better on Failure-Prone HPC Platforms? | Valentin Le Fvre, George Bosilca, Aurlien Bouteiller, Thomas Hrault, Atsushi Hori, Yves Robert, Jack J. Dongarra |
| 2018 | HPDC | ADAPT: an event-based adaptive collective communication framework. | Xi Luo, Wei Wu, George Bosilca, Thananon Patinyasakdikul, Linnan Wang, Jack J. Dongarra |
| 2018 | ICCS | The Design of Fast and Energy-Efficient Linear Solvers: On the Potential of Half-Precision Arithmetic and Iterative Refinement Techniques. | Azzam Haidar, Ahmad Abdelfattah, Mawussi Zounon, Panruo Wu, Srikara Pranesh, Stanimire Tomov, Jack J. Dongarra |
| 2018 | SC | Harnessing GPU tensor cores for fast FP16 arithmetic to speed up mixed-precision iterative refinement solvers. | Azzam Haidar, Stanimire Tomov, Jack J. Dongarra, Nicholas J. Higham |
| 2018 | SBAC-PAD | A Jaccard Weights Kernel Leveraging Independent Thread Scheduling on GPUs. | Hartwig Anzt, Jack J. Dongarra |
| 2018 | SBAC-PAD | Variable-Size Batched Condition Number Calculation on GPUs. | Hartwig Anzt, Jack J. Dongarra, Goran Flegar, Thomas Grtzmacher |
| 2017 | EuroPar | Optimized Batched Linear Algebra for Modern Architectures. | Jack J. Dongarra, Sven Hammarling, Nicholas J. Higham, Samuel D. Relton, Mawussi Zounon |
| 2017 | ICCS | Factorization and Inversion of a Million Matrices using GPUs: Challenges and Countermeasures. | Ahmad Abdelfattah, Azzam Haidar, Stanimire Tomov, Jack J. Dongarra |
| 2017 | ICCS | Variable-Size Batched Gauss-Huard for Block-Jacobi Preconditioning. | Hartwig Anzt, Jack J. Dongarra, Goran Flegar, Enrique S. Quintana-Ort, Andrs E. Toms |
| 2017 | ICCS | The Design and Performance of Batched BLAS on Modern High-Performance Computing Systems. | Jack J. Dongarra, Sven Hammarling, Nicholas J. Higham, Samuel D. Relton, Pedro Valero-Lara, Mawussi Zounon |
| 2017 | ICCS | Optimizing the SVD Bidiagonalization Process for a Batch of Small Matrices. | Tingxing Dong, Azzam Haidar, Stanimire Tomov, Jack J. Dongarra |
| 2017 | ICCS | The Art of Computational Science, Bridging Gaps - Forming Alloys. Preface for ICCS 2017. | Petros Koumoutsakos, Eleni N. Chatzi, Valeria V. Krzhizhanovskaya, Michael Lees, Jack J. Dongarra, Peter M. A. Sloot |
| 2017 | ICPP | Variable-Size Batched LU for Small Matrices and Its Integration into Block-Jacobi Preconditioning. | Hartwig Anzt, Jack J. Dongarra, Goran Flegar, Enrique S. Quintana-Ort |
| 2017 | ICS | Novel HPC techniques to batch execution of many variable size BLAS computations on GPUs. | Ahmad Abdelfattah, Azzam Haidar, Stanimire Tomov, Jack J. Dongarra |
| 2017 | PPoPP | Batched Gauss-Jordan Elimination for Block-Jacobi Preconditioner Generation on GPUs. | Hartwig Anzt, Jack J. Dongarra, Goran Flegar, Enrique S. Quintana-Ort |
| 2017 | PPoPP | High-performance Cholesky factorization for GPU-only execution. | Azzam Haidar, Ahmad Abdelfattah, Stanimire Tomov, Jack J. Dongarra |
| 2017 | SC | Flexible batched sparse matrix-vector product on GPUs. | Hartwig Anzt, Gary Collins, Jack J. Dongarra, Goran Flegar, Enrique S. Quintana-Ort |
| 2017 | SC | Investigating half precision arithmetic to accelerate dense linear system solvers. | Azzam Haidar, Panruo Wu, Stanimire Tomov, Jack J. Dongarra |
| 2017 | SC | Dynamic task discovery in PaRSEC: a data-flow task-based runtime. | Reazul Hoque, Thomas Hrault, George Bosilca, Jack J. Dongarra |
| 2016 | EuroPar | High-Performance Matrix-Matrix Multiplications of Very Small Matrices. | Ian Masliah, Ahmad Abdelfattah, Azzam Haidar, Stanimire Tomov, Marc Baboulin, Jol Falcou, Jack J. Dongarra |
| 2016 | HPDC | With Extreme Scale Computing the Rules Have Changed. | Jack J. Dongarra |
| 2016 | HPDC | GPU-Aware Non-contiguous Data Movement In Open MPI. | Wei Wu, George Bosilca, Rolf Vandevaart, Sylvain Jeaugey, Jack J. Dongarra |
| 2016 | ICCS | High-Performance Tensor Contractions for GPUs. | Ahmad Abdelfattah, Marc Baboulin, Veselin Dobrev, Jack J. Dongarra, Christopher W. Earl, Joel Falcou, Azzam Haidar, Ian Karlin, Tzanio V. Kolev, Ian Masliah, Stanimire Tomov |
| 2016 | ICCS | Performance Tuning and Optimization Techniques of Fixed and Variable Size Batched Cholesky Factorization on GPUs. | Ahmad Abdelfattah, Azzam Haidar, Stanimire Tomov, Jack J. Dongarra |
| 2016 | ICCS | Data through the Computational Lens, Preface for ICCS 2016. | Ilkay Altintas, Michael Normal, Michael Lees, Valeria V. Krzhizhanovskaya, Jack J. Dongarra, Peter M. A. Sloot |
| 2016 | SC | Batched Generation of Incomplete Sparse Approximate Inverses on GPUs. | Hartwig Anzt, Edmond Chow, Thomas Huckle, Jack J. Dongarra |
| 2016 | SC | Failure detection and propagation in HPC systems. | George Bosilca, Aurlien Bouteiller, Amina Guermouche, Thomas Hrault, Yves Robert, Pierre Sens, Jack J. Dongarra |
| 2016 | SC | Towards Achieving Performance Portability Using Directives for Accelerators. | M. Graham Lopez, Vernica G. Vergara Larrea, Wayne Joubert, Oscar R. Hernandez, Azzam Haidar, Stanimire Tomov, Jack J. Dongarra |
| 2016 | SC | Performance-Portable Autotuning of OpenCL Kernels for Convolutional Layers of Deep Neural Networks. | Yaohung M. Tsai, Piotr Luszczek, Jakub Kurzak, Jack J. Dongarra |
| 2015 | CLUSTER | PaRSEC in Practice: Optimizing a Legacy Chemistry Application through Distributed Task-Based Execution. | Anthony Danalis, Heike Jagode, George Bosilca, Jack J. Dongarra |
| 2015 | EuroPar | Iterative Sparse Triangular Solves for Preconditioning. | Hartwig Anzt, Edmond Chow, Jack J. Dongarra |
| 2015 | HPCC | Flexible Linear Algebra Development and Scheduling with Cholesky Factorization. | Azzam Haidar, Asim YarKhan, Chongxiao Cao, Piotr Luszczek, Stanimire Tomov, Jack J. Dongarra |
| 2015 | ICCS | Performance Analysis and Optimisation of Two-sided Factorization Algorithms for Heterogeneous Platform. | Khairul Kabir, Azzam Haidar, Stanimire Tomov, Jack J. Dongarra |
| 2015 | PPAM | Dense Symmetric Indefinite Factorization on GPU Accelerated Architectures. | Marc Baboulin, Jack J. Dongarra, Adrien Rmy, Stanimire Tomov, Ichitaro Yamazaki |
| 2015 | PPAM | Accelerating NWChem Coupled Cluster Through Dataflow-Based Execution. | Heike Jagode, Anthony Danalis, George Bosilca, Jack J. Dongarra |
| 2015 | PPoPP | Energy efficiency and performance frontiers for sparse computations on GPU supercomputers. | Hartwig Anzt, Stanimire Tomov, Jack J. Dongarra |
| 2015 | PPoPP | Towards batched linear solvers on accelerated hardware platforms. | Azzam Haidar, Tingxing Dong, Piotr Luszczek, Stanimire Tomov, Jack J. Dongarra |
| 2015 | PPoPP | Optimization for performance and energy for batched matrix computations on GPUs. | Azzam Haidar, Tingxing Dong, Piotr Luszczek, Stanimire Tomov, Jack J. Dongarra |
| 2015 | SC | Adaptive precision solvers for sparse linear systems. | Hartwig Anzt, Jack J. Dongarra, Enrique S. Quintana-Ort |
| 2015 | SC | Tuning stationary iterative solvers for fault resilience. | Hartwig Anzt, Jack J. Dongarra, Enrique S. Quintana-Ort |
| 2015 | SC | GPU-accelerated co-design of induced dimension reduction: algorithmic fusion and kernel overlap. | Hartwig Anzt, Eduardo Ponce, Gregory D. Peterson, Jack J. Dongarra |
| 2015 | SC | Weighted dynamic scheduling with many parallelism grains for offloading of numerical workloads to multiple varied accelerators. | Azzam Haidar, Yulu Jia, Piotr Luszczek, Stanimire Tomov, Asim YarKhan, Jack J. Dongarra |
| 2015 | SC | Visualizing execution traces with task dependencies. | Blake Haugen, Stephen Richmond, Jakub Kurzak, Chad A. Steed, Jack J. Dongarra |
| 2015 | SC | Practical scalable consensus for pseudo-synchronous distributed systems. | Thomas Hrault, Aurlien Bouteiller, George Bosilca, Marc Gamell, Keita Teranishi, Manish Parashar, Jack J. Dongarra |
| 2015 | SC | Performance of random sampling for computing low-rank approximations of a dense matrix on GPUs. | Tho Mary, Ichitaro Yamazaki, Jakub Kurzak, Piotr Luszczek, Stanimire Tomov, Jack J. Dongarra |
| 2015 | SC | Efficient implementation of quantum materials simulations on distributed CPU-GPU systems. | Raffaele Solc, Anton Kozhevnikov, Azzam Haidar, Stanimire Tomov, Jack J. Dongarra, Thomas C. Schulthess |
| 2015 | SC | Randomized algorithms to update partial singular value decomposition on a hybrid CPU/GPU cluster. | Ichitaro Yamazaki, Jakub Kurzak, Piotr Luszczek, Jack J. Dongarra |
| 2015 | SC | Mixed-precision block gram Schmidt orthogonalization. | Ichitaro Yamazaki, Stanimire Tomov, Jakub Kurzak, Jack J. Dongarra, Jesse L. Barlow |
| 2014 | CLUSTER | Utilizing dataflow-based execution for coupled cluster methods. | Heike McCraw, Anthony Danalis, Thomas Hrault, George Bosilca, Jack J. Dongarra, Karol Kowalski, Theresa L. Windus |
| 2014 | CLUSTER | Power monitoring with PAPI for extreme scale architectures and dataflow-based programming models. | Heike McCraw, James Ralph, Anthony Danalis, Jack J. Dongarra |
| 2014 | HPCC | LU Factorization of Small Matrices: Accelerating Batched DGETRF on the GPU. | Tingxing Dong, Azzam Haidar, Piotr Luszczek, James Austin Harris, Stanimire Tomov, Jack J. Dongarra |
| 2014 | ICCS | Big Data Meets Computational Science, Preface for ICCS 2014. | David Abramson, Michael Lees, Valeria V. Krzhizhanovskaya, Jack J. Dongarra, Peter M. A. Sloot |
| 2014 | ICPP | A Fast Batched Cholesky Factorization on a GPU. | Tingxing Dong, Azzam Haidar, Stanimire Tomov, Jack J. Dongarra |
| 2014 | ICPP | Parallel Simulation of Superscalar Scheduling. | Blake Haugen, Jakub Kurzak, Asim YarKhan, Piotr Luszczek, Jack J. Dongarra |
| 2014 | ICS | Scaling up matrix computations on shared-memory manycore systems with 1000 CPU cores. | Fengguang Song, Jack J. Dongarra |
| 2014 | ISPASS | MIAMI: A framework for application performance diagnosis. | Gabriel Marin, Jack J. Dongarra, Daniel Terpstra |
| 2014 | SC | Performance and portability with OpenCL for throughput-oriented HPC workloads across accelerators, coprocessors, and multicore processors. | Chongxiao Cao, Mark Gates, Azzam Haidar, Piotr Luszczek, Stanimire Tomov, Ichitaro Yamazaki, Jack J. Dongarra |
| 2014 | SC | PTG: an abstraction for unhindered parallelism. | Anthony Danalis, George Bosilca, Aurlien Bouteiller, Thomas Hrault, Jack J. Dongarra |
| 2014 | SC | Deflation strategies to improve the convergence of communication-avoiding GMRES. | Ichitaro Yamazaki, Stanimire Tomov, Jack J. Dongarra |
| 2013 | EuroPar | Implementing a Systolic Algorithm for QR Factorization on Multicore Clusters with PaRSEC. | Guillaume Aupy, Mathieu Faverge, Yves Robert, Jakub Kurzak, Piotr Luszczek, Jack J. Dongarra |
| 2013 | EuroPar | Multi-criteria Checkpointing Strategies: Response-Time versus Resource Utilization. | Aurlien Bouteiller, Franck Cappello, Jack J. Dongarra, Amina Guermouche, Thomas Hrault, Yves Robert |
| 2013 | ICCS | Computation at the Frontiers of Science, preface for ICCS 2013. | Vassil Alexandrov, Michael Lees, Valeria V. Krzhizhanovskaya, Jack J. Dongarra, Peter M. A. Sloot |
| 2013 | ICCS | A Parallel Solver for Incompressible Fluid Flows. | Yushan Wang, Marc Baboulin, Jack J. Dongarra, Jol Falcou, Yann Fraigneau, Olivier P. Le Matre |
| 2013 | ICS | Toward a scalable multi-GPU eigensolver via compute-intensive kernels and efficient communication. | Azzam Haidar, Mark Gates, Stanimire Tomov, Jack J. Dongarra |
| 2013 | PPAM | Portable HPC Programming on Intel Many-Integrated-Core Hardware with MAGMA Port to Xeon Phi. | Jack J. Dongarra, Mark Gates, Azzam Haidar, Yulu Jia, Khairul Kabir, Piotr Luszczek, Stanimire Tomov |
| 2013 | SC | Optimal Checkpointing Period: Time vs. Energy. | Guillaume Aupy, Anne Benoit, Thomas Hrault, Yves Robert, Jack J. Dongarra |
| 2013 | SC | Parallel reduction to hessenberg form with algorithm-based fault tolerance. | Yulu Jia, George Bosilca, Piotr Luszczek, Jack J. Dongarra |
| 2013 | SC | CPU-GPU hybrid bidiagonal reduction with soft error resilience. | Yulu Jia, Piotr Luszczek, George Bosilca, Jack J. Dongarra |
| 2012 | EuroPar | GPU-Accelerated Asynchronous Error Correction for Mixed Precision Iterative Refinement. | Hartwig Anzt, Piotr Luszczek, Jack J. Dongarra, Vincent Heuveline |
| 2012 | EuroPar | Weighted Block-Asynchronous Iteration on GPU-Accelerated Systems. | Hartwig Anzt, Stanimire Tomov, Jack J. Dongarra, Vincent Heuveline |
| 2012 | EuroPar | A Checkpoint-on-Failure Protocol for Algorithm-Based Recovery in Standard MPI. | Wesley Bland, Peng Du, Aurlien Bouteiller, Thomas Hrault, George Bosilca, Jack J. Dongarra |
| 2012 | EuroPar | From Serial Loops to Parallel Execution on Distributed Systems. | George Bosilca, Aurlien Bouteiller, Anthony Danalis, Thomas Hrault, Jack J. Dongarra |
| 2012 | ICS | Enabling and scaling matrix computations on heterogeneous multi-core and multi-GPU systems. | Fengguang Song, Stanimire Tomov, Jack J. Dongarra |
| 2012 | PPoPP | Algorithm-based fault tolerance for dense matrix factorizations. | Peng Du, Aurlien Bouteiller, George Bosilca, Thomas Hrault, Jack J. Dongarra |
| 2012 | SC | Abstract: Matrices Over Runtime Systems at Exascale. | Emmanuel Agullo, George Bosilca, Brenger Bramas, Cedric Castagnede, Olivier Coulaud, Eric Darve, Jack J. Dongarra, Mathieu Faverge, Nathalie Furmento, Luc Giraud, Xavier Lacoste, Julien Langou, Hatem Ltaief, Matthias Messner, Raymond Namyst, Pierre Ramet, Toru Takahashi, Samuel Thibault, Stanimire Tomov, Ichitaro Yamazaki |
| 2012 | SC | Poster: Matrices over Runtime Systems at Exascale. | Emmanuel Agullo, George Bosilca, Brenger Bramas, Cedric Castagnede, Olivier Coulaud, Eric Darve, Jack J. Dongarra, Mathieu Faverge, Nathalie Furmento, Luc Giraud, Xavier Lacoste, Julien Langou, Hatem Ltaief, Matthias Messner, Raymond Namyst, Pierre Ramet, Toru Takahashi, Samuel Thibault, Stanimire Tomov, Ichitaro Yamazaki |
| 2012 | SC | Abstract: A Novel Hybrid CPU-GPU Generalized Eigensolver for Electronic Structure Calculations Based on Fine Grained Memory Aware Tasks. | Raffaele Solc, Azzam Haidar, Stanimire Tomov, Thomas C. Schulthess, Jack J. Dongarra |
| 2012 | SC | Poster: A Novel Hybrid CPU-GPU Generalized Eigensolver for Electronic Structure Calculations Based on Fine Grained Memory Aware Tasks. | Raffaele Solc, Azzam Haidar, Stanimire Tomov, Thomas C. Schulthess, Jack J. Dongarra |
| 2012 | SPAA | A scalable framework for heterogeneous GPU-based clusters. | Fengguang Song, Jack J. Dongarra |
| 2011 | AICCSA | LU factorization for accelerator-based systems. | Emmanuel Agullo, Cdric Augonnet, Jack J. Dongarra, Mathieu Faverge, Julien Langou, Hatem Ltaief, Stanimire Tomov |
| 2011 | CCGRID | EZTrace: A Generic Framework for Performance Analysis. | Franois Trahay, Franois Ru, Mathieu Faverge, Yutaka Ishikawa, Raymond Namyst, Jack J. Dongarra |
| 2011 | CLUSTER | Performance Portability of a GPU Enabled Factorization with the DAGuE Framework. | George Bosilca, Aurlien Bouteiller, Thomas Hrault, Pierre Lemarinier, Narapat Ohm Saengpatsa, Stanimire Tomov, Jack J. Dongarra |
| 2011 | CLUSTER | On Scalability for MPI Runtime Systems. | George Bosilca, Thomas Hrault, Ala Rezmerita, Jack J. Dongarra |
| 2011 | CLUSTER | High Performance Dense Linear System Solver with Soft Error Resilience. | Peng Du, Piotr Luszczek, Jack J. Dongarra |
| 2011 | CLUSTER | Process Distance-Aware Adaptive MPI Collective Communications. | Teng Ma, Thomas Hrault, George Bosilca, Jack J. Dongarra |
| 2011 | EuroPar | A Fully Empirical Autotuned Dense QR Factorization for Multicore Architectures. | Emmanuel Agullo, Jack J. Dongarra, Rajib Nath, Stanimire Tomov |
| 2011 | EuroPar | Correlated Set Coordination in Fault Tolerant Message Logging Protocols. | Aurlien Bouteiller, Thomas Hrault, George Bosilca, Jack J. Dongarra |
| 2011 | EuroPar | Evaluation of the HPC Challenge Benchmarks in Virtualized Environments. | Piotr Luszczek, Eric Meek, Shirley Moore, Daniel Terpstra, Vincent M. Weaver, Jack J. Dongarra |
| 2011 | ICPP | Kernel Assisted Collective Intra-node MPI Communication among Multi-Core and Many-Core CPUs. | Teng Ma, George Bosilca, Aurlien Bouteiller, Brice Goglin, Jeffrey M. Squyres, Jack J. Dongarra |
| 2011 | PPAM | Reducing the Amount of Pivoting in Symmetric Indefinite Systems. | Dulceneia Becker, Marc Baboulin, Jack J. Dongarra |
| 2011 | PPAM | Enhancing Parallelism of Tile Bidiagonal Transformation on Multicore Architectures Using Tree Reduction. | Hatem Ltaief, Piotr Luszczek, Jack J. Dongarra |
| 2011 | PPAM | Reducing the Time to Tune Parallel Dense Linear Algebra Routines with Partial Execution and Performance Modeling. | Piotr Luszczek, Jack J. Dongarra |
| 2011 | SC | High performance matrix inversion based on LU factorization for multicore architectures. | Jack J. Dongarra, Mathieu Faverge, Hatem Ltaief, Piotr Luszczek |
| 2011 | SC | Soft error resilient QR factorization for hybrid system with GPGPU. | Peng Du, Piotr Luszczek, Stanimire Tomov, Jack J. Dongarra |
| 2011 | SC | Parallel reduction to condensed forms for symmetric eigenvalue problems using aggregated fine-grained and memory-aware kernels. | Azzam Haidar, Hatem Ltaief, Jack J. Dongarra |
| 2011 | SC | Poster: new features of the PAPI hardware counter library. | Shirley Moore, Daniel Terpstra, Vincent M. Weaver, Heike Jagode, James Ralph, Jack J. Dongarra |
| 2011 | SC | Optimizing symmetric dense matrix-vector multiplication on GPUs. | Rajib Nath, Stanimire Tomov, Tingxing Dong, Jack J. Dongarra |
| 2011 | SC | Panel: many-task computing meets exascales. | Ioan Raicu, Daniel A. Reed, Jack J. Dongarra, Daniel S. Katz, David Abramson |
| 2010 | SC | Scalable Tile Communication-Avoiding QR Factorization on Multicore Cluster Systems. | Fengguang Song, Hatem Ltaief, Bilel Hadri, Jack J. Dongarra |
| 2009 | CLUSTER | Reasons for a pessimistic or optimistic message logging protocol in MPI uncoordinated failure, recovery. | Aurlien Bouteiller, Thomas Ropars, George Bosilca, Christine Morin, Jack J. Dongarra |
| 2009 | CLUSTER | Analytical modeling and optimization for affinity based thread scheduling on multicore systems. | Fengguang Song, Shirley Moore, Jack J. Dongarra |
| 2009 | ICCS | A Holistic Approach for Performance Measurement and Analysis for Petascale Applications. | Heike Jagode, Jack J. Dongarra, Sadaf R. Alam, Jeffrey S. Vetter, Wyatt Spear, Allen D. Malony |
| 2009 | ICCS | A Note on Auto-tuning GEMM for GPUs. | Yinan Li, Jack J. Dongarra, Stanimire Tomov |
| 2009 | ICCS | A Scalable Non-blocking Multicast Scheme for Distributed DAG Scheduling. | Fengguang Song, Jack J. Dongarra, Shirley Moore |
| 2009 | ICPP | CIFTS: A Coordinated Infrastructure for Fault-Tolerant Systems. | Rinku Gupta, Peter H. Beckman, Byung-Hoon Park, Ewing L. Lusk, Paul Hargrove, Al Geist, Dhabaleswar K. Panda, Andrew Lumsdaine, Jack J. Dongarra |
| 2009 | SC | Comparative study of one-sided factorizations with multiple software packages on multi-core hardware. | Emmanuel Agullo, Bilel Hadri, Hatem Ltaief, Jack J. Dongarra |
| 2009 | SC | Dynamic task scheduling for linear algebra algorithms on distributed-memory multicore systems. | Fengguang Song, Asim YarKhan, Jack J. Dongarra |
| 2008 | CLUSTER | A comparison of search heuristics for empirical code optimization. | Keith Seymour, Haihang You, Jack J. Dongarra |
| 2008 | HPDC | The impact of paravirtualized memory hierarchy on linear algebra computational kernels and software. | Lamia Youseff, Keith Seymour, Haihang You, Jack J. Dongarra, Richard Wolski |
| 2008 | ICCS | Fast and Small Short Vector SIMD Matrix Multiplication Kernels for the Synergistic Processing Element of the CELL Processor. | Wesley Alvaro, Jakub Kurzak, Jack J. Dongarra |
| 2008 | PPoPP | Matrix product on heterogeneous master-worker platforms. | Jack J. Dongarra, Jean-Francois Pineau, Yves Robert, Frdric Vivien |
| 2007 | CCGRID | Reliability Analysis of Self-Healing Network using Discrete-Event Simulation. | Thara Angskun, George Bosilca, Graham E. Fagg, Jelena Pjesivac-Grbovic, Jack J. Dongarra |
| 2007 | EuroPar | On Using Incremental Profiling for the Performance Analysis of Shared Memory Parallel Applications. | Karl Frlinger, Michael Gerndt, Jack J. Dongarra |
| 2007 | EuroPar | Decision Trees and MPI Collective Algorithm Selection Problem. | Jelena Pjesivac-Grbovic, George Bosilca, Graham E. Fagg, Thara Angskun, Jack J. Dongarra |
| 2007 | HPDC | Feedback-directed thread scheduling with memory considerations. | Fengguang Song, Shirley Moore, Jack J. Dongarra |
| 2007 | ICCS | Scalability Analysis of the SPEC OpenMP Benchmarks on Large-Scale Shared Memory Multiprocessors. | Karl Frlinger, Michael Gerndt, Jack J. Dongarra |
| 2007 | ICPP | L2 Cache Modeling for Scientific Applications on Chip Multi-Processors. | Fengguang Song, Shirley Moore, Jack J. Dongarra |
| 2007 | ISPA | Binomial Graph: A Scalable and Fault-Tolerant Logical Network Topology. | Thara Angskun, George Bosilca, Jack J. Dongarra |
| 2007 | PDCAT | Optimal Routing in Binomial Graph Networks. | Thara Angskun, George Bosilca, Bradley T. Vander Zanden, Jack J. Dongarra |
| 2007 | PPAM | Parallel Tiled QR Factorization for Multicore Architectures. | Alfredo Buttari, Julien Langou, Jakub Kurzak, Jack J. Dongarra |
| 2007 | SPAA | Bi-objective scheduling algorithms for optimizing makespan and reliability on heterogeneous systems. | Jack J. Dongarra, Emmanuel Jeannot, Erik Saule, Zhiao Shi |
| 2006 | CCGRID | Proposal of MPI Operation Level Checkpoint/Rollback and One Implementation. | Yuan Tang, Graham E. Fagg, Jack J. Dongarra |
| 2006 | CLUSTER | Robust task scheduling in non-deterministic heterogeneous computing systems. | Zhiao Shi, Emmanuel Jeannot, Jack J. Dongarra |
| 2006 | ICPP | The Impact of Multicore on Math Software and Exploiting Single Precision Computing to Obtain Double Precision Results. | Jack J. Dongarra |
| 2006 | ISPA | The Impact of Multicore on Math Software and Exploiting Single Precision Computing to Obtain Double Precision Results. | Jack J. Dongarra |
| 2006 | SC | Poster reception - Targeting multi-core architectures for linear algebra applications. | Alfredo Buttari, Jakub Kurzak, Jack J. Dongarra |
| 2006 | SC | HPC challenge - The 2006 HPC challenge awards. | Jack J. Dongarra, Jeremy Kepner |
| 2006 | SC | Tools and techniques for performance - Exploiting the performance of 32 bit floating point arithmetic in obtaining 64 bit accuracy (revisiting iterative refinement for linear systems). | Julie Langou, Julien Langou, Piotr Luszczek, Jakub Kurzak, Alfredo Buttari, Jack J. Dongarra |
| 2006 | SC | S12 - The HPC Challenge (HPCC) benchmark suite. | Piotr Luszczek, David H. Bailey, Jack J. Dongarra, Jeremy Kepner, Robert F. Lucas, Rolf Rabenseifner, Daisuke Takahashi |
| 2005 | CLUSTER | Processes Distribution of Homogeneous Parallel Linear Algebra Routines on Heterogeneous Clusters. | Javier Cuenca, Luis-Pedro Garca, Domingo Gimnez, Jack J. Dongarra |
| 2005 | ICCS | Numerically Stable Real Number Codes Based on Random Matrices. | Zizhong Chen, Jack J. Dongarra |
| 2005 | ICCS | Comparison of Nonlinear Conjugate-Gradient Methods for Computing the Electronic Properties of Nanostructure Architectures. | Stanimire Tomov, Julien Langou, Andrew Canning, Lin-Wang Wang, Jack J. Dongarra |
| 2005 | ICPP | Automatic Experimental Analysis of Communication Patterns in Virtual Topologies. | Nikhil Bhatia, Fengguang Song, Felix Wolf, Jack J. Dongarra, Bernd Mohr, Shirley Moore |
| 2005 | PPoPP | Fault tolerant high performance computing by a coding approach. | Zizhong Chen, Graham E. Fagg, Edgar Gabriel, Julien Langou, Thara Angskun, George Bosilca, Jack J. Dongarra |
| 2004 | EuroPar | Efficient Pattern Search in Large Traces Through Successive Refinement. | Felix Wolf, Bernd Mohr, Jack J. Dongarra, Shirley Moore |
| 2004 | ICCS | Accurate Cache and TLB Characterization Using Hardware Counters. | Jack J. Dongarra, Shirley Moore, Philip Mucci, Keith Seymour, Haihang You |
| 2004 | ICCS | Design of Interactive Environment for Numerically Intensive Parallel Linear Algebra Calculations. | Piotr Luszczek, Jack J. Dongarra |
| 2004 | ICPP | An Algebra for Cross-Experiment Performance Analysis. | Fengguang Song, Felix Wolf, Nikhil Bhatia, Jack J. Dongarra, Shirley Moore |
| 2004 | ISPA | Present and Future Supercomputer Architectures. | Jack J. Dongarra |
| 2003 | CCGRID | A Performance Oriented Migration Framework For The Grid. | Sathish S. Vadhiyar, Jack J. Dongarra |
| 2003 | EuroPar | GrADSolve - RPC for High Performance Computing on the Grid. | Sathish S. Vadhiyar, Jack J. Dongarra, Asim YarKhan |
| 2003 | GECCO | Distributed Probabilistic Model-Building Genetic Algorithm. | Tomoyuki Hiroyasu, Mitsunori Miki, Masaki Sano, Hisashi Shimosaka, Shigeyoshi Tsutsui, Jack J. Dongarra |
| 2003 | ICCS | Self-Adapting Software for Numerical Linear Algebra Library Routines on Clusters. | Zizhong Chen, Jack J. Dongarra, Piotr Luszczek, Kenneth Roche |
| 2003 | ICCS | Self-Adapting Numerical Software and Automatic Tuning of Heuristics. | Jack J. Dongarra, Victor Eijkhout |
| 2003 | ICCS | Performance Instrumentation and Measurement for Terascale Systems. | Jack J. Dongarra, Allen D. Malony, Shirley Moore, Philip Mucci, Sameer Shende |
| 2003 | ICCS | visPerf: Monitoring Tool for Grid Computing. | DongWoo Lee, Jack J. Dongarra, Rudrapatna S. Ramakrishna |
| 2003 | PDP | Automatic Optimisation of Parallel Linear Algebra Routines in Systems with Variable Load. | Javier Cuenca, Domingo Gimnez, Jos Gonzlez, Jack J. Dongarra, Kenneth Roche |
| 2002 | CCGRID | Three Tools to Help with Cluster and Grid Computing: SANS-Effort, PAPI, and NetSolve. | Jack J. Dongarra |
| 2002 | CLUSTER | Trends in High Performance Computing and Using Numerical Libraries on Cluster. | Jack J. Dongarra |
| 2002 | HPDC | A Metascheduler For The Grid. | Sathish S. Vadhiyar, Jack J. Dongarra |
| 2001 | CLUSTER | High Performance Computing and Trends: Connected Computational Requirements with Computing Resources. | Jack J. Dongarra |
| 2001 | EuroPar | High Performance Computing and Trends: Connecting Computational Requirements with Computing Resources. | Jack J. Dongarra |
| 2001 | ICCS | Fault Tolerant MPI for the HARNESS Meta-computing System. | Graham E. Fagg, Antonin Bukovsky, Jack J. Dongarra |
| 2001 | ICCS | Towards an Accurate Model for Collective Communications. | Sathish S. Vadhiyar, Graham E. Fagg, Jack J. Dongarra |
| 2001 | NCA | NetSolve and Its Applications. | Jack J. Dongarra |
| 2001 | SC | Numerical libraries and the grid: the GrADS experiments with ScaLAPACK. | Antoine Petitet, L. Susan Blackford, Jack J. Dongarra, Brett Ellis, Graham E. Fagg, Kenneth Roche, Sathish S. Vadhiyar |
| 2000 | EuroPar | Request Sequencing: Optimizing Communication for the Grid. | Dorian C. Arnold, Dieter Bachmann, Jack J. Dongarra |
| 2000 | SC | A Scalable Cross-Platform Infrastructure for Application Performance Tuning Using Hardware Counters. | Shirley Browne, Jack J. Dongarra, Nathan Garner, Kevin S. London, Philip Mucci |
| 2000 | SC | Automatically Tuned Collective Communications. | Sathish S. Vadhiyar, Graham E. Fagg, Jack J. Dongarra |
| 1999 | EuroPar | A Comparison of Parallel Solvers for Diagonally Dominant and General Narrow-Banded Linear Systems II. | Peter Arbenz, Andrew J. Cleary, Jack J. Dongarra, Markus Hegland |
| 1999 | EuroPar | Adaptive Scheduling for Task Farming with Grid Middleware. | Henri Casanova, MyungHo Kim, James S. Plank, Jack J. Dongarra |
| 1998 | HCW | NetSolve: A Network-Enabled Solver; Examples and Users. | Henri Casanova, Jack J. Dongarra |
| 1998 | HPDC | HARNESS: Heterogeneous Adaptable Reconfigurable NEtworked SystemS. | Jack J. Dongarra, Graham E. Fagg, Al Geist, James Arthur Kohl, Philip M. Papadopoulos, Stephen L. Scott, Vaidy S. Sunderam, M. Magliardi |
| 1998 | SC | Automatically Tuned Linear Algebra Software. | R. Clinton Whaley, Jack J. Dongarra |
| 1997 | SC | Scalable Networked Information Processing Environment (SNIPE). | Graham E. Fagg, Keith Moore, Jack J. Dongarra, Al Geist |
| 1996 | EuroPar | Selected Results from the ParkBench Benchmark. | Jack J. Dongarra, Tony Hey, Erich Strohmaier |
| 1996 | SC | ScaLAPACK: A Portable Linear Algebra Library for Distributed Memory Computers - Design Issues and Performance. | L. Susan Blackford, Jaeyoung Choi, Andrew J. Cleary, James Demmel, Inderjit S. Dhillon, Jack J. Dongarra, Sven Hammarling, Greg Henry, Antoine Petitet, Ken Stanley, David W. Walker, R. Clinton Whaley |
| 1996 | SC | NetSovle: A Network Server for Solving Computational Science Problems. | Henri Casanova, Jack J. Dongarra |
| 1995 | SC | Distributed Information Management in the National HPCC Software Exchange. | Shirley Browne, Jack J. Dongarra, Geoffrey C. Fox, Kenneth A. Hawick, Ken Kennedy, Rick Stevens, Robert Olson, Tom Rowan |
| 1994 | HPDC | Constructing Numerical Software Libraries for HPCC Environments. | Jack J. Dongarra |
| 1993 | SC | LAPACK++: a design overview of object-oriented extensions for high performance linear algebra. | Jack J. Dongarra, Roldan Pozo, David W. Walker |
| 1991 | SC | Graphical development tools for network-based concurrent supercomputing. | Adam Beguelin, Jack J. Dongarra |
| 1991 | SC | Gordon Bell prize lectures. | Jack J. Dongarra, Alan H. Karp, Ken Miura, Horst D. Simon |
| 1990 | SC | LAPACK: a portable linear algebra library for high-performance computers. | Edward C. Anderson, Zhaojun Bai, Jack J. Dongarra, Anne Greenbaum, A. McKenney, Jeremy Du Croz, Sven Hammarling, James Demmel, Christian H. Bischof, Danny C. Sorensen |
| 1989 | COMPSAC | A graphics tool to aid in the generation of parallel FORTRAN programs. | Orlie Brewer, Jack J. Dongarra, Danny C. Sorensen |
| 1988 | SC | Vectorizing compilers: a test suite and results. | David Callahan, Jack J. Dongarra, David Levine |
| 1987 | ICS | The LINPACK Benchmark: An Explanation. | Jack J. Dongarra |
| 1986 | SAC | Performance of advanced architectures. | Jack J. Dongarra |
| 1985 | ARITH | A fast algorithm for the symmetric eigenvalue problem. | Jack J. Dongarra, Danny C. Sorensen |