| 2023 | ARITH | Efficient Additions and Montgomery Reductions of Large Integers for SIMD. | Pengchang Ren, Reiji Suda, Vorapong Suppakitpaisarn |
| 2016 | ICA3PP | Efficient Parallel Algorithm for Optimal DAG Structure Search on Parallel Computer with Torus Network. | Hirokazu Honda, Yoshinori Tamada, Reiji Suda |
| 2015 | PPAM | Performance Analysis of the Chebyshev Basis Conjugate Gradient Method on the K Computer. | Yosuke Kumagai, Akihiro Fujii, Teruo Tanaka, Yusuke Hirota, Takeshi Fukaya, Toshiyuki Imamura, Reiji Suda |
| 2014 | SAC | The future of accelerator programming: abstraction, performance or can we have both? | Kamil Rocki, Martin Burtscher, Reiji Suda |
| 2013 | ICCS | A Mathematical Method for Online Autotuning of Power and Energy Consumption with Corrected Temperature Effects. | Reiji Suda, Luo Cheng, Takahiro Katagiri |
| 2013 | ICPADS | The Future of Accelerator Programming: Abstraction, Performance or Can We Have Both? | Kamil Rocki, Martin Burtscher, Reiji Suda |
| 2013 | SC | Register level sort algorithm on multi-core SIMD processors. | Tian Xiaochen, Kamil Rocki, Reiji Suda |
| 2013 | TrustCom | An Efficient Task Partitioning and Scheduling Method for Symmetric Multiple GPU Architecture. | Cheng Luo, Reiji Suda |
| 2012 | CCGRID | Accelerating 2-opt and 3-opt Local Search Using GPU in the Travelling Salesman Problem. | Kamil Rocki, Reiji Suda |
| 2012 | GECCO | An efficient GPU implementation of a multi-start TSP solver for large problem instances. | Kamil Rocki, Reiji Suda |
| 2012 | ICPADS | MSSM: An Efficient Scheduling Mechanism for CUDA Basing on Task Partition. | Cheng Luo, Reiji Suda |
| 2012 | SC | Abstract: High Performance GPU Accelerated TSP Solver. | Kamil Rocki, Reiji Suda |
| 2012 | SC | Poster: High Performance GPU Accelerated TSP Solver. | Kamil Rocki, Reiji Suda |
| 2012 | SPAA | Brief announcement: a GPU accelerated iterated local search TSP solver. | Kamil Rocki, Reiji Suda |
| 2011 | DASC | A Performance and Energy Consumption Analytical Model for GPU. | Cheng Luo, Reiji Suda |
| 2009 | ASPDAC | Aspects of GPU for general purpose high performance computing. | Reiji Suda, Takayuki Aoki, Shoichi Hirasawa, Akira Nukada, Hiroki Honda, Satoshi Matsuoka |
| 2009 | PDCAT | Accurate Measurements and Precise Modeling of Power Dissipation of CUDA Kernels toward Power Optimized High Performance CPU-GPU Computing. | Reiji Suda, Da Qi Ren |
| 2009 | PPAM | Modeling and Optimizing the Power Performance of Large Matrices Multiplication on Multi-core and GPU Platform with CUDA. | Da Qi Ren, Reiji Suda |
| 2009 | PPAM | Parallel Minimax Tree Searching on GPU. | Kamil Rocki, Reiji Suda |
| 2008 | CLUSTER | An optimized Dynamic Load Balancing method for parallel 3-D mesh refinement for finite element electromagnetics with Tetrahedra. | Da Qi Ren, Dennis Giannacopoulos, Reiji Suda |
| 2008 | CLUSTER | Divisible load scheduling with improved asymptotic optimality. | Reiji Suda |
| 2007 | HPCC | High Performance FFT on SGI Altix 3700. | Akira Nukada, Daisuke Takahashi, Reiji Suda, Akira Nishida |
| 2007 | PPAM | Cloth Simulation in the SILC Matrix Computation Framework: A Case Study. | Tamito Kajiyama, Akira Nukada, Reiji Suda, Hidehiko Hasegawa, Akira Nishida |
| 2005 | PPAM | SILC: A Flexible and Environment-Independent Interface for Matrix Computation Libraries. | Tamito Kajiyama, Akira Nukada, Hidehiko Hasegawa, Reiji Suda, Akira Nishida |
| 1998 | ASPDAC | The Ensparsed LU Decomposition Method for Large Scale Circuit Transient Analysis. | Reiji Suda, Yoshio Oyanagi |
| 1995 | ICS | Implementation of Sparta, a Highly Parallel Circuit Simulator by the Preconditioned Jacobi Method, on a Distributed Memory Machine. | Reiji Suda, Yoshio Oyanagi |