Stanimire Tomov
Publication record assembled from the DBLP archive of ranked conferences.
Papers indexed
53
Venues
15
Active years
2004–2023
Best venue rank
A*
Where they publish
Papers
53 indexed papers, newest first.
| Year | Venue | Title | Authors |
|---|---|---|---|
| 2023 | SC | GPU-based LU Factorization and Solve on Batches of Matrices with Band Structure. | Ahmad Abdelfattah, Stanimire Tomov, Piotr Luszczek, Hartwig Anzt, Jack J. Dongarra |
| 2022 | CLUSTER | Lossy all-to-all exchange for accelerating parallel 3-D FFTs on hybrid architectures with GPUs. | Sbastien Cayrols, Jiali Li, George Bosilca, Stanimire Tomov, Alan Ayala, Jack J. Dongarra |
| 2022 | CVPR | GPU-Based Homotopy Continuation for Minimal Problems in Computer Vision. | Chiang-Heng Chien, Hongyi Fan, Ahmad Abdelfattah, Elias P. Tsigaridas, Stanimire Tomov, Benjamin B. Kimia |
| 2022 | SC | Addressing Irregular Patterns of Matrix Computations on GPUs and Their Impact on Applications Powered by Sparse Direct Solvers. | Ahmad Abdelfattah, Pieter Ghysels, Wajih Boukaram, Stanimire Tomov, Xiaoye Sherry Li, Jack J. Dongarra |
| 2021 | PACT | Scalability Issues in FFT Computation. | Alan Ayala, Stanimire Tomov, Miroslav Stoyanov, Jack J. Dongarra |
| 2020 | ICCS | heFFTe: Highly Efficient FFT for Exascale. | Alan Ayala, Stanimire Tomov, Azzam Haidar, Jack J. Dongarra |
| 2019 | SC | Towards Half-Precision Computation for Complex Matrices: A Case Study for Mixed Precision Solvers on GPUs. | Ahmad Abdelfattah, Stanimire Tomov, Jack J. Dongarra |
| 2018 | HiPC | Optimizing the Fast Fourier Transform Using Mixed Precision on Tensor Core Hardware. | Anumeena Sorna, Xiaohe Cheng, Eduardo F. D'Azevedo, Kwai Wong, Stanimire Tomov |
| 2018 | ICCS | The Design of Fast and Energy-Efficient Linear Solvers: On the Potential of Half-Precision Arithmetic and Iterative Refinement Techniques. | Azzam Haidar, Ahmad Abdelfattah, Mawussi Zounon, Panruo Wu, Srikara Pranesh, Stanimire Tomov, Jack J. Dongarra |
| 2018 | SC | Harnessing GPU tensor cores for fast FP16 arithmetic to speed up mixed-precision iterative refinement solvers. | Azzam Haidar, Stanimire Tomov, Jack J. Dongarra, Nicholas J. Higham |
| 2017 | ICCS | Factorization and Inversion of a Million Matrices using GPUs: Challenges and Countermeasures. | Ahmad Abdelfattah, Azzam Haidar, Stanimire Tomov, Jack J. Dongarra |
| 2017 | ICCS | Optimizing the SVD Bidiagonalization Process for a Batch of Small Matrices. | Tingxing Dong, Azzam Haidar, Stanimire Tomov, Jack J. Dongarra |
| 2017 | ICS | Novel HPC techniques to batch execution of many variable size BLAS computations on GPUs. | Ahmad Abdelfattah, Azzam Haidar, Stanimire Tomov, Jack J. Dongarra |
| 2017 | PPoPP | High-performance Cholesky factorization for GPU-only execution. | Azzam Haidar, Ahmad Abdelfattah, Stanimire Tomov, Jack J. Dongarra |
| 2017 | SC | Investigating half precision arithmetic to accelerate dense linear system solvers. | Azzam Haidar, Panruo Wu, Stanimire Tomov, Jack J. Dongarra |
| 2016 | EuroPar | High-Performance Matrix-Matrix Multiplications of Very Small Matrices. | Ian Masliah, Ahmad Abdelfattah, Azzam Haidar, Stanimire Tomov, Marc Baboulin, Jol Falcou, Jack J. Dongarra |
| 2016 | ICCS | High-Performance Tensor Contractions for GPUs. | Ahmad Abdelfattah, Marc Baboulin, Veselin Dobrev, Jack J. Dongarra, Christopher W. Earl, Joel Falcou, Azzam Haidar, Ian Karlin, Tzanio V. Kolev, Ian Masliah, Stanimire Tomov |
| 2016 | ICCS | Performance Tuning and Optimization Techniques of Fixed and Variable Size Batched Cholesky Factorization on GPUs. | Ahmad Abdelfattah, Azzam Haidar, Stanimire Tomov, Jack J. Dongarra |
| 2016 | SC | Towards Achieving Performance Portability Using Directives for Accelerators. | M. Graham Lopez, Vernica G. Vergara Larrea, Wayne Joubert, Oscar R. Hernandez, Azzam Haidar, Stanimire Tomov, Jack J. Dongarra |
| 2015 | HPCC | Flexible Linear Algebra Development and Scheduling with Cholesky Factorization. | Azzam Haidar, Asim YarKhan, Chongxiao Cao, Piotr Luszczek, Stanimire Tomov, Jack J. Dongarra |
| 2015 | ICCS | Performance Analysis and Optimisation of Two-sided Factorization Algorithms for Heterogeneous Platform. | Khairul Kabir, Azzam Haidar, Stanimire Tomov, Jack J. Dongarra |
| 2015 | PPAM | Dense Symmetric Indefinite Factorization on GPU Accelerated Architectures. | Marc Baboulin, Jack J. Dongarra, Adrien Rmy, Stanimire Tomov, Ichitaro Yamazaki |
| 2015 | PPoPP | Energy efficiency and performance frontiers for sparse computations on GPU supercomputers. | Hartwig Anzt, Stanimire Tomov, Jack J. Dongarra |
| 2015 | PPoPP | Towards batched linear solvers on accelerated hardware platforms. | Azzam Haidar, Tingxing Dong, Piotr Luszczek, Stanimire Tomov, Jack J. Dongarra |
| 2015 | PPoPP | Optimization for performance and energy for batched matrix computations on GPUs. | Azzam Haidar, Tingxing Dong, Piotr Luszczek, Stanimire Tomov, Jack J. Dongarra |
| 2015 | SC | Weighted dynamic scheduling with many parallelism grains for offloading of numerical workloads to multiple varied accelerators. | Azzam Haidar, Yulu Jia, Piotr Luszczek, Stanimire Tomov, Asim YarKhan, Jack J. Dongarra |
| 2015 | SC | Performance of random sampling for computing low-rank approximations of a dense matrix on GPUs. | Tho Mary, Ichitaro Yamazaki, Jakub Kurzak, Piotr Luszczek, Stanimire Tomov, Jack J. Dongarra |
| 2015 | SC | Efficient implementation of quantum materials simulations on distributed CPU-GPU systems. | Raffaele Solc, Anton Kozhevnikov, Azzam Haidar, Stanimire Tomov, Jack J. Dongarra, Thomas C. Schulthess |
| 2015 | SC | Mixed-precision block gram Schmidt orthogonalization. | Ichitaro Yamazaki, Stanimire Tomov, Jakub Kurzak, Jack J. Dongarra, Jesse L. Barlow |
| 2014 | HPCC | LU Factorization of Small Matrices: Accelerating Batched DGETRF on the GPU. | Tingxing Dong, Azzam Haidar, Piotr Luszczek, James Austin Harris, Stanimire Tomov, Jack J. Dongarra |
| 2014 | ICPP | A Fast Batched Cholesky Factorization on a GPU. | Tingxing Dong, Azzam Haidar, Stanimire Tomov, Jack J. Dongarra |
| 2014 | SC | Performance and portability with OpenCL for throughput-oriented HPC workloads across accelerators, coprocessors, and multicore processors. | Chongxiao Cao, Mark Gates, Azzam Haidar, Piotr Luszczek, Stanimire Tomov, Ichitaro Yamazaki, Jack J. Dongarra |
| 2014 | SC | Domain Decomposition Preconditioners for Communication-Avoiding Krylov Methods on a Hybrid CPU/GPU Cluster. | Ichitaro Yamazaki, Sivasankaran Rajamanickam, Erik G. Boman, Mark Hoemmen, Michael A. Heroux, Stanimire Tomov |
| 2014 | SC | Deflation strategies to improve the convergence of communication-avoiding GMRES. | Ichitaro Yamazaki, Stanimire Tomov, Jack J. Dongarra |
| 2013 | ICS | Toward a scalable multi-GPU eigensolver via compute-intensive kernels and efficient communication. | Azzam Haidar, Mark Gates, Stanimire Tomov, Jack J. Dongarra |
| 2013 | PPAM | Portable HPC Programming on Intel Many-Integrated-Core Hardware with MAGMA Port to Xeon Phi. | Jack J. Dongarra, Mark Gates, Azzam Haidar, Yulu Jia, Khairul Kabir, Piotr Luszczek, Stanimire Tomov |
| 2012 | EuroPar | Weighted Block-Asynchronous Iteration on GPU-Accelerated Systems. | Hartwig Anzt, Stanimire Tomov, Jack J. Dongarra, Vincent Heuveline |
| 2012 | ICS | Enabling and scaling matrix computations on heterogeneous multi-core and multi-GPU systems. | Fengguang Song, Stanimire Tomov, Jack J. Dongarra |
| 2012 | SC | Abstract: Matrices Over Runtime Systems at Exascale. | Emmanuel Agullo, George Bosilca, Brenger Bramas, Cedric Castagnede, Olivier Coulaud, Eric Darve, Jack J. Dongarra, Mathieu Faverge, Nathalie Furmento, Luc Giraud, Xavier Lacoste, Julien Langou, Hatem Ltaief, Matthias Messner, Raymond Namyst, Pierre Ramet, Toru Takahashi, Samuel Thibault, Stanimire Tomov, Ichitaro Yamazaki |
| 2012 | SC | Poster: Matrices over Runtime Systems at Exascale. | Emmanuel Agullo, George Bosilca, Brenger Bramas, Cedric Castagnede, Olivier Coulaud, Eric Darve, Jack J. Dongarra, Mathieu Faverge, Nathalie Furmento, Luc Giraud, Xavier Lacoste, Julien Langou, Hatem Ltaief, Matthias Messner, Raymond Namyst, Pierre Ramet, Toru Takahashi, Samuel Thibault, Stanimire Tomov, Ichitaro Yamazaki |
| 2012 | SC | Abstract: A Novel Hybrid CPU-GPU Generalized Eigensolver for Electronic Structure Calculations Based on Fine Grained Memory Aware Tasks. | Raffaele Solc, Azzam Haidar, Stanimire Tomov, Thomas C. Schulthess, Jack J. Dongarra |
| 2012 | SC | Poster: A Novel Hybrid CPU-GPU Generalized Eigensolver for Electronic Structure Calculations Based on Fine Grained Memory Aware Tasks. | Raffaele Solc, Azzam Haidar, Stanimire Tomov, Thomas C. Schulthess, Jack J. Dongarra |
| 2011 | AICCSA | LU factorization for accelerator-based systems. | Emmanuel Agullo, Cdric Augonnet, Jack J. Dongarra, Mathieu Faverge, Julien Langou, Hatem Ltaief, Stanimire Tomov |
| 2011 | CLUSTER | Performance Portability of a GPU Enabled Factorization with the DAGuE Framework. | George Bosilca, Aurlien Bouteiller, Thomas Hrault, Pierre Lemarinier, Narapat Ohm Saengpatsa, Stanimire Tomov, Jack J. Dongarra |
| 2011 | EuroPar | A Fully Empirical Autotuned Dense QR Factorization for Multicore Architectures. | Emmanuel Agullo, Jack J. Dongarra, Rajib Nath, Stanimire Tomov |
| 2011 | EuroPar | Introduction. | Wolfgang Karl, Samuel Thibault, Stanimire Tomov, Taisuke Boku |
| 2011 | ICPP | Parallel Performance Measurement of Heterogeneous Parallel Systems with GPUs. | Allen D. Malony, Scott Biersdorff, Sameer Shende, Heike Jagode, Stanimire Tomov, Guido Juckeland, Robert Dietrich, Duncan Poole, Christopher Lamb |
| 2011 | SC | Soft error resilient QR factorization for hybrid system with GPGPU. | Peng Du, Piotr Luszczek, Stanimire Tomov, Jack J. Dongarra |
| 2011 | SC | Optimizing symmetric dense matrix-vector multiplication on GPUs. | Rajib Nath, Stanimire Tomov, Tingxing Dong, Jack J. Dongarra |
| 2009 | ICCS | A Note on Auto-tuning GEMM for GPUs. | Yinan Li, Jack J. Dongarra, Stanimire Tomov |
| 2009 | SAC | Bulk based preconditioning for quantum dot computations. | Christof Vmel, Stanimire Tomov, Osni Marques |
| 2005 | ICCS | Comparison of Nonlinear Conjugate-Gradient Methods for Computing the Electronic Properties of Nanostructure Architectures. | Stanimire Tomov, Julien Langou, Andrew Canning, Lin-Wang Wang, Jack J. Dongarra |
| 2004 | CBMS | Toward a Systems Biology Software Toolkit. | Donald J. Johann, Michael D. McGuigan, Stanimire Tomov, Eric Blum, Gordon R. Whiteley, Emanuel F. Petricoin, Lance A. Liotta |