| 2025 | CLUSTER | Parallel Tall-and-Skinny QR Factorization Based on LU-CholeskyQR Algorithm. | Yuki Uchino, Toshiyuki Imamura |
| 2025 | SC | High-Performance and Power-Efficient Emulation of Matrix Multiplication using INT8 Matrix Engines. | Yuki Uchino, Katsuhisa Ozaki, Toshiyuki Imamura |
| 2024 | SC | High-Performance Eigensolver Combining EigenExa and Iterative Refinement. | Yuki Uchino, Toshiyuki Imamura |
| 2022 | PPAM | Infinite-Precision Inner Product and Sparse Matrix-Vector Multiplication Using Ozaki Scheme with Dot2 on Manycore Processors. | Daichi Mukunoki, Katsuhisa Ozaki, Takeshi Ogita, Toshiyuki Imamura |
| 2022 | SC | GPU Optimization of Lattice Boltzmann Method with Local Ensemble Transform Kalman Filter. | Yuta Hasegawa, Toshiyuki Imamura, Takuya Ina, Naoyuki Onodera, Yuuichi Asahi, Yasuhiro Idomura |
| 2021 | ICCSA | A Rapid Euclidean Norm Calculation Algorithm that Reduces Overflow and Underflow. | Takeyuki Harayama, Shuhei Kudo, Daichi Mukunoki, Toshiyuki Imamura, Daisuke Takahashi |
| 2021 | ICPP | Accurate Matrix Multiplication on Binary128 Format Accelerated by Ozaki Scheme. | Daichi Mukunoki, Katsuhisa Ozaki, Takeshi Ogita, Toshiyuki Imamura |
| 2020 | CLUSTER | Prompt Report on Exa-Scale HPL-AI Benchmark. | Shuhei Kudo, Keigo Nitadori, Takuya Ina, Toshiyuki Imamura |
| 2020 | CLUSTER | An FPGA-based Sound Field Rendering System. | Yiyu Tan, Toshiyuki Imamura |
| 2020 | SC | Acceleration of fusion plasma turbulence simulations using the mixed-precision communication-avoiding krylov method. | Yasuhiro Idomura, Takuya Ina, Yussuf Ali, Toshiyuki Imamura |
| 2020 | SC | A 1024-member ensemble data assimilation with 3.5-km mesh global weather simulations. | Hisashi Yashiro, Koji Terasaki, Yuta Kawai, Shuhei Kudo, Takemasa Miyoshi, Toshiyuki Imamura, Kazuo Minami, Hikaru Inoue, Tatsuo Nishiki, Takayuki Saji, Masaki Satoh, Hirofumi Tomita |
| 2018 | HPDC | Performance Evaluation of a Toolkit for Sparse Tensor Decomposition. | Yiyu Tan, Toshiyuki Imamura |
| 2018 | ICCS | Performance Analysis of 2D-compatible 2.5D-PDGEMM on Knights Landing Cluster. | Daichi Mukunoki, Toshiyuki Imamura |
| 2017 | PPAM | Parallel Divide-and-Conquer Algorithm for Solving Tridiagonal Eigenvalue Problems on Manycore Systems. | Yusuke Hirota, Toshiyuki Imamura |
| 2017 | PPAM | Implementation and Performance Analysis of 2.5D-PDGEMM on the K Computer. | Daichi Mukunoki, Toshiyuki Imamura |
| 2017 | SC | Application of a communication-avoiding generalized minimal residual method to a gyrokinetic five dimensional eulerian code on many core platforms. | Yasuhiro Idomura, Takuya Ina, Akie Mayumi, Susumu Yamada, Kazuya Matsumoto, Yuuichi Asahi, Toshiyuki Imamura |
| 2016 | CLUSTER | Reduced-Precision Floating-Point Formats on GPUs for High Performance and Energy Efficient Computation. | Daichi Mukunoki, Toshiyuki Imamura |
| 2016 | SC | Left-Preconditioned Communication-Avoiding Conjugate Gradient Methods for Multiphase CFD Simulations on the K Computer. | Akie Mayumi, Yasuhiro Idomura, Takuya Ina, Susumu Yamada, Toshiyuki Imamura |
| 2015 | PDP | Fast Implementation of General Matrix-Vector Multiplication (GEMV) on Kepler GPUs. | Daichi Mukunoki, Toshiyuki Imamura, Daisuke Takahashi |
| 2015 | PPAM | Performance Analysis of the Chebyshev Basis Conjugate Gradient Method on the K Computer. | Yosuke Kumagai, Akihiro Fujii, Teruo Tanaka, Yusuke Hirota, Takeshi Fukaya, Toshiyuki Imamura, Reiji Suda |
| 2014 | EGPGV | A Study of Parallel Data Compression Using Proper Orthogonal Decomposition on the K Computer. | Chongke Bi, Kenji Ono, Kwan-Liu Ma, Haiyuan Wu, Toshiyuki Imamura |
| 2013 | PPAM | Eigen-G: GPU-Based Eigenvalue Solver for Real-Symmetric Dense Matrices. | Toshiyuki Imamura, Susumu Yamada, Masahiko Machida |
| 2012 | SC | Abstract: Communication Overlap Techniques for Improved Strong Scaling of Gyrokinetic Eulerian Code beyond 100k Cores on the K-Computer. | Yasuhiro Idomura, Motoki Nakata, Susumu Yamada, Masahiko Machida, Toshiyuki Imamura, Tomohiko Watanabe, Masanori Nunami, Hikaru Inoue, Shigenobu Tsutsumi, Ikuo Miyoshi, Naoyuki Shida |
| 2012 | SC | Poster: Communication Overlap Techniques for Improved Strong Scaling of Gyrokinetic Eulerian Code beyond 100k Cores on the K-Computer. | Yasuhiro Idomura, Motoki Nakata, Susumu Yamada, Masahiko Machida, Toshiyuki Imamura, Tomohiko Watanabe, Masanori Nunami, Hikaru Inoue, Shigenobu Tsutsumi, Ikuo Miyoshi, Naoyuki Shida |
| 2012 | SC | Abstract: Preliminary Report for a High Precision Distributed Memory Parallel Eigenvalue Solver. | Toshiyuki Imamura, Susumu Yamada, Masahiko Machida |
| 2012 | SC | Poster: Preliminary Report for a High Precision Distributed Memory Parallel Eigenvalue Solver. | Toshiyuki Imamura, Susumu Yamada, Masahiko Machida |
| 2011 | SC | Parallelization design on multi-core platforms in density matrix renormalization group toward 2-D quantum strongly-correlated systems. | Susumu Yamada, Toshiyuki Imamura, Masahiko Machida |
| 2006 | SC | Gordon Bell finalists I - High-performance computing for exact numerical approaches to quantum many-body problems on the earth simulator. | Susumu Yamada, Toshiyuki Imamura, Takuma Kano, Masahiko Machida |
| 2005 | SC | 16.447 TFlops and 159-Billion-dimensional Exact-diagonalization for Trapped Fermion-Hubbard Model on the Earth Simulator. | Susumu Yamada, Toshiyuki Imamura, Masahiko Machida |
| 2003 | ISPA | MPI-2 Support in Heterogeneous Computing Environment Using an SCore Cluster System. | Yuichi Tsujita, Toshiyuki Imamura, Nobuhiro Yamagishi, Hiroshi Takemiya |