| 2024 | EuroPar | Reduced-Precision and Reduced-Exponent Formats for Accelerating Adaptive Precision Sparse Matrix-Vector Product. | Stef Graillat, Fabienne Jzquel, Tho Mary, Romo Molina, Daichi Mukunoki |
| 2024 | SC | Performance evaluation and modelling of single-precision matrix multiplication on Cerebras CS-2. | Ryunosuke Matsuzaki, Daichi Mukunoki, Takaaki Miyajima |
| 2022 | PPAM | Infinite-Precision Inner Product and Sparse Matrix-Vector Multiplication Using Ozaki Scheme with Dot2 on Manycore Processors. | Daichi Mukunoki, Katsuhisa Ozaki, Takeshi Ogita, Toshiyuki Imamura |
| 2021 | ICCSA | A Rapid Euclidean Norm Calculation Algorithm that Reduces Overflow and Underflow. | Takeyuki Harayama, Shuhei Kudo, Daichi Mukunoki, Toshiyuki Imamura, Daisuke Takahashi |
| 2021 | ICPP | Accurate Matrix Multiplication on Binary128 Format Accelerated by Ozaki Scheme. | Daichi Mukunoki, Katsuhisa Ozaki, Takeshi Ogita, Toshiyuki Imamura |
| 2019 | PPAM | Reproducible BLAS Routines with Tunable Accuracy Using Ozaki Scheme for Many-Core Architectures. | Daichi Mukunoki, Takeshi Ogita, Katsuhisa Ozaki |
| 2018 | ICCS | Performance Analysis of 2D-compatible 2.5D-PDGEMM on Knights Landing Cluster. | Daichi Mukunoki, Toshiyuki Imamura |
| 2017 | PPAM | Implementation and Performance Analysis of 2.5D-PDGEMM on the K Computer. | Daichi Mukunoki, Toshiyuki Imamura |
| 2016 | CLUSTER | Reduced-Precision Floating-Point Formats on GPUs for High Performance and Energy Efficient Computation. | Daichi Mukunoki, Toshiyuki Imamura |
| 2015 | PDP | Fast Implementation of General Matrix-Vector Multiplication (GEMV) on Kepler GPUs. | Daichi Mukunoki, Toshiyuki Imamura, Daisuke Takahashi |
| 2013 | ICCSA | Optimization of Sparse Matrix-Vector Multiplication for CRS Format on NVIDIA Kepler Architecture GPUs. | Daichi Mukunoki, Daisuke Takahashi |
| 2013 | PPAM | Using Quadruple Precision Arithmetic to Accelerate Krylov Subspace Methods on GPUs. | Daichi Mukunoki, Daisuke Takahashi |