Skip to content

Stanimire Tomov

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

53

Venues

15

Active years

2004–2023

Best venue rank

A*

Where they publish

Papers

53 indexed papers, newest first.

YearVenueTitleAuthors
2023SCGPU-based LU Factorization and Solve on Batches of Matrices with Band Structure.Ahmad Abdelfattah, Stanimire Tomov, Piotr Luszczek, Hartwig Anzt, Jack J. Dongarra
2022CLUSTERLossy all-to-all exchange for accelerating parallel 3-D FFTs on hybrid architectures with GPUs.Sbastien Cayrols, Jiali Li, George Bosilca, Stanimire Tomov, Alan Ayala, Jack J. Dongarra
2022CVPRGPU-Based Homotopy Continuation for Minimal Problems in Computer Vision.Chiang-Heng Chien, Hongyi Fan, Ahmad Abdelfattah, Elias P. Tsigaridas, Stanimire Tomov, Benjamin B. Kimia
2022SCAddressing Irregular Patterns of Matrix Computations on GPUs and Their Impact on Applications Powered by Sparse Direct Solvers.Ahmad Abdelfattah, Pieter Ghysels, Wajih Boukaram, Stanimire Tomov, Xiaoye Sherry Li, Jack J. Dongarra
2021PACTScalability Issues in FFT Computation.Alan Ayala, Stanimire Tomov, Miroslav Stoyanov, Jack J. Dongarra
2020ICCSheFFTe: Highly Efficient FFT for Exascale.Alan Ayala, Stanimire Tomov, Azzam Haidar, Jack J. Dongarra
2019SCTowards Half-Precision Computation for Complex Matrices: A Case Study for Mixed Precision Solvers on GPUs.Ahmad Abdelfattah, Stanimire Tomov, Jack J. Dongarra
2018HiPCOptimizing the Fast Fourier Transform Using Mixed Precision on Tensor Core Hardware.Anumeena Sorna, Xiaohe Cheng, Eduardo F. D'Azevedo, Kwai Wong, Stanimire Tomov
2018ICCSThe Design of Fast and Energy-Efficient Linear Solvers: On the Potential of Half-Precision Arithmetic and Iterative Refinement Techniques.Azzam Haidar, Ahmad Abdelfattah, Mawussi Zounon, Panruo Wu, Srikara Pranesh, Stanimire Tomov, Jack J. Dongarra
2018SCHarnessing GPU tensor cores for fast FP16 arithmetic to speed up mixed-precision iterative refinement solvers.Azzam Haidar, Stanimire Tomov, Jack J. Dongarra, Nicholas J. Higham
2017ICCSFactorization and Inversion of a Million Matrices using GPUs: Challenges and Countermeasures.Ahmad Abdelfattah, Azzam Haidar, Stanimire Tomov, Jack J. Dongarra
2017ICCSOptimizing the SVD Bidiagonalization Process for a Batch of Small Matrices.Tingxing Dong, Azzam Haidar, Stanimire Tomov, Jack J. Dongarra
2017ICSNovel HPC techniques to batch execution of many variable size BLAS computations on GPUs.Ahmad Abdelfattah, Azzam Haidar, Stanimire Tomov, Jack J. Dongarra
2017PPoPPHigh-performance Cholesky factorization for GPU-only execution.Azzam Haidar, Ahmad Abdelfattah, Stanimire Tomov, Jack J. Dongarra
2017SCInvestigating half precision arithmetic to accelerate dense linear system solvers.Azzam Haidar, Panruo Wu, Stanimire Tomov, Jack J. Dongarra
2016EuroParHigh-Performance Matrix-Matrix Multiplications of Very Small Matrices.Ian Masliah, Ahmad Abdelfattah, Azzam Haidar, Stanimire Tomov, Marc Baboulin, Jol Falcou, Jack J. Dongarra
2016ICCSHigh-Performance Tensor Contractions for GPUs.Ahmad Abdelfattah, Marc Baboulin, Veselin Dobrev, Jack J. Dongarra, Christopher W. Earl, Joel Falcou, Azzam Haidar, Ian Karlin, Tzanio V. Kolev, Ian Masliah, Stanimire Tomov
2016ICCSPerformance Tuning and Optimization Techniques of Fixed and Variable Size Batched Cholesky Factorization on GPUs.Ahmad Abdelfattah, Azzam Haidar, Stanimire Tomov, Jack J. Dongarra
2016SCTowards Achieving Performance Portability Using Directives for Accelerators.M. Graham Lopez, Vernica G. Vergara Larrea, Wayne Joubert, Oscar R. Hernandez, Azzam Haidar, Stanimire Tomov, Jack J. Dongarra
2015HPCCFlexible Linear Algebra Development and Scheduling with Cholesky Factorization.Azzam Haidar, Asim YarKhan, Chongxiao Cao, Piotr Luszczek, Stanimire Tomov, Jack J. Dongarra
2015ICCSPerformance Analysis and Optimisation of Two-sided Factorization Algorithms for Heterogeneous Platform.Khairul Kabir, Azzam Haidar, Stanimire Tomov, Jack J. Dongarra
2015PPAMDense Symmetric Indefinite Factorization on GPU Accelerated Architectures.Marc Baboulin, Jack J. Dongarra, Adrien Rmy, Stanimire Tomov, Ichitaro Yamazaki
2015PPoPPEnergy efficiency and performance frontiers for sparse computations on GPU supercomputers.Hartwig Anzt, Stanimire Tomov, Jack J. Dongarra
2015PPoPPTowards batched linear solvers on accelerated hardware platforms.Azzam Haidar, Tingxing Dong, Piotr Luszczek, Stanimire Tomov, Jack J. Dongarra
2015PPoPPOptimization for performance and energy for batched matrix computations on GPUs.Azzam Haidar, Tingxing Dong, Piotr Luszczek, Stanimire Tomov, Jack J. Dongarra
2015SCWeighted dynamic scheduling with many parallelism grains for offloading of numerical workloads to multiple varied accelerators.Azzam Haidar, Yulu Jia, Piotr Luszczek, Stanimire Tomov, Asim YarKhan, Jack J. Dongarra
2015SCPerformance of random sampling for computing low-rank approximations of a dense matrix on GPUs.Tho Mary, Ichitaro Yamazaki, Jakub Kurzak, Piotr Luszczek, Stanimire Tomov, Jack J. Dongarra
2015SCEfficient implementation of quantum materials simulations on distributed CPU-GPU systems.Raffaele Solc, Anton Kozhevnikov, Azzam Haidar, Stanimire Tomov, Jack J. Dongarra, Thomas C. Schulthess
2015SCMixed-precision block gram Schmidt orthogonalization.Ichitaro Yamazaki, Stanimire Tomov, Jakub Kurzak, Jack J. Dongarra, Jesse L. Barlow
2014HPCCLU Factorization of Small Matrices: Accelerating Batched DGETRF on the GPU.Tingxing Dong, Azzam Haidar, Piotr Luszczek, James Austin Harris, Stanimire Tomov, Jack J. Dongarra
2014ICPPA Fast Batched Cholesky Factorization on a GPU.Tingxing Dong, Azzam Haidar, Stanimire Tomov, Jack J. Dongarra
2014SCPerformance and portability with OpenCL for throughput-oriented HPC workloads across accelerators, coprocessors, and multicore processors.Chongxiao Cao, Mark Gates, Azzam Haidar, Piotr Luszczek, Stanimire Tomov, Ichitaro Yamazaki, Jack J. Dongarra
2014SCDomain Decomposition Preconditioners for Communication-Avoiding Krylov Methods on a Hybrid CPU/GPU Cluster.Ichitaro Yamazaki, Sivasankaran Rajamanickam, Erik G. Boman, Mark Hoemmen, Michael A. Heroux, Stanimire Tomov
2014SCDeflation strategies to improve the convergence of communication-avoiding GMRES.Ichitaro Yamazaki, Stanimire Tomov, Jack J. Dongarra
2013ICSToward a scalable multi-GPU eigensolver via compute-intensive kernels and efficient communication.Azzam Haidar, Mark Gates, Stanimire Tomov, Jack J. Dongarra
2013PPAMPortable HPC Programming on Intel Many-Integrated-Core Hardware with MAGMA Port to Xeon Phi.Jack J. Dongarra, Mark Gates, Azzam Haidar, Yulu Jia, Khairul Kabir, Piotr Luszczek, Stanimire Tomov
2012EuroParWeighted Block-Asynchronous Iteration on GPU-Accelerated Systems.Hartwig Anzt, Stanimire Tomov, Jack J. Dongarra, Vincent Heuveline
2012ICSEnabling and scaling matrix computations on heterogeneous multi-core and multi-GPU systems.Fengguang Song, Stanimire Tomov, Jack J. Dongarra
2012SCAbstract: Matrices Over Runtime Systems at Exascale.Emmanuel Agullo, George Bosilca, Brenger Bramas, Cedric Castagnede, Olivier Coulaud, Eric Darve, Jack J. Dongarra, Mathieu Faverge, Nathalie Furmento, Luc Giraud, Xavier Lacoste, Julien Langou, Hatem Ltaief, Matthias Messner, Raymond Namyst, Pierre Ramet, Toru Takahashi, Samuel Thibault, Stanimire Tomov, Ichitaro Yamazaki
2012SCPoster: Matrices over Runtime Systems at Exascale.Emmanuel Agullo, George Bosilca, Brenger Bramas, Cedric Castagnede, Olivier Coulaud, Eric Darve, Jack J. Dongarra, Mathieu Faverge, Nathalie Furmento, Luc Giraud, Xavier Lacoste, Julien Langou, Hatem Ltaief, Matthias Messner, Raymond Namyst, Pierre Ramet, Toru Takahashi, Samuel Thibault, Stanimire Tomov, Ichitaro Yamazaki
2012SCAbstract: A Novel Hybrid CPU-GPU Generalized Eigensolver for Electronic Structure Calculations Based on Fine Grained Memory Aware Tasks.Raffaele Solc, Azzam Haidar, Stanimire Tomov, Thomas C. Schulthess, Jack J. Dongarra
2012SCPoster: A Novel Hybrid CPU-GPU Generalized Eigensolver for Electronic Structure Calculations Based on Fine Grained Memory Aware Tasks.Raffaele Solc, Azzam Haidar, Stanimire Tomov, Thomas C. Schulthess, Jack J. Dongarra
2011AICCSALU factorization for accelerator-based systems.Emmanuel Agullo, Cdric Augonnet, Jack J. Dongarra, Mathieu Faverge, Julien Langou, Hatem Ltaief, Stanimire Tomov
2011CLUSTERPerformance Portability of a GPU Enabled Factorization with the DAGuE Framework.George Bosilca, Aurlien Bouteiller, Thomas Hrault, Pierre Lemarinier, Narapat Ohm Saengpatsa, Stanimire Tomov, Jack J. Dongarra
2011EuroParA Fully Empirical Autotuned Dense QR Factorization for Multicore Architectures.Emmanuel Agullo, Jack J. Dongarra, Rajib Nath, Stanimire Tomov
2011EuroParIntroduction.Wolfgang Karl, Samuel Thibault, Stanimire Tomov, Taisuke Boku
2011ICPPParallel Performance Measurement of Heterogeneous Parallel Systems with GPUs.Allen D. Malony, Scott Biersdorff, Sameer Shende, Heike Jagode, Stanimire Tomov, Guido Juckeland, Robert Dietrich, Duncan Poole, Christopher Lamb
2011SCSoft error resilient QR factorization for hybrid system with GPGPU.Peng Du, Piotr Luszczek, Stanimire Tomov, Jack J. Dongarra
2011SCOptimizing symmetric dense matrix-vector multiplication on GPUs.Rajib Nath, Stanimire Tomov, Tingxing Dong, Jack J. Dongarra
2009ICCSA Note on Auto-tuning GEMM for GPUs.Yinan Li, Jack J. Dongarra, Stanimire Tomov
2009SACBulk based preconditioning for quantum dot computations.Christof Vmel, Stanimire Tomov, Osni Marques
2005ICCSComparison of Nonlinear Conjugate-Gradient Methods for Computing the Electronic Properties of Nanostructure Architectures.Stanimire Tomov, Julien Langou, Andrew Canning, Lin-Wang Wang, Jack J. Dongarra
2004CBMSToward a Systems Biology Software Toolkit.Donald J. Johann, Michael D. McGuigan, Stanimire Tomov, Eric Blum, Gordon R. Whiteley, Emanuel F. Petricoin, Lance A. Liotta