| 2026 | AAAI | On the Learning Dynamics of Two-layer Linear Networks with Label Noise SGD. | Tongcheng Zhang, Zhanpeng Zhou, Mingze Wang, Andi Han, Wei Huang, Taiji Suzuki, Junchi Yan |
| 2025 | AISTATS | Quantifying the Optimization and Generalization Advantages of Graph Neural Networks Over Multilayer Perceptrons. | Wei Huang, Yuan Cao, Haonan Wang, Xin Cao, Taiji Suzuki |
| 2025 | AISTATS | Clustered Invariant Risk Minimization. | Tomoya Murata, Atsushi Nitanda, Taiji Suzuki |
| 2025 | ICLR | Flow matching achieves almost minimax optimal convergence. | Kenji Fukumizu, Taiji Suzuki, Noboru Isobe, Kazusato Oko, Masanori Koyama |
| 2025 | ICLR | Direct Distributional Optimization for Provable Alignment of Diffusion Models. | Ryotaro Kawata, Kazusato Oko, Atsushi Nitanda, Taiji Suzuki |
| 2025 | ICLR | Optimality and Adaptivity of Deep Neural Features for Instrumental Variable Regression. | Juno Kim, Dimitri Meunier, Arthur Gretton, Taiji Suzuki, Zhu Li |
| 2025 | ICLR | Transformers Provably Solve Parity Efficiently with Chain of Thought. | Juno Kim, Taiji Suzuki |
| 2025 | ICLR | On the Optimization and Generalization of Two-layer Transformers with Sign Gradient Descent. | Bingrui Li, Wei Huang, Andi Han, Zhanpeng Zhou, Taiji Suzuki, Jun Zhu, Jianfei Chen |
| 2025 | ICLR | State Space Models are Provably Comparable to Transformers in Dynamic Token Selection. | Naoki Nishikawa, Taiji Suzuki |
| 2025 | ICLR | Weighted Point Set Embedding for Multimodal Contrastive Learning Toward Optimal Similarity Metric. | Toshimitsu Uesaka, Taiji Suzuki, Yuhta Takida, Chieh-Hsin Lai, Naoki Murata, Yuki Mitsufuji |
| 2025 | ICML | Provable In-Context Vector Arithmetic via Retrieving Task Concepts. | Dake Bu, Wei Huang, Andi Han, Atsushi Nitanda, Qingfu Zhang, Hau-San Wong, Taiji Suzuki |
| 2025 | ICML | On the Role of Label Noise in the Feature Learning Process. | Andi Han, Wei Huang, Zhanpeng Zhou, Gang Niu, Wuyang Chen, Junchi Yan, Akiko Takeda, Taiji Suzuki |
| 2025 | ICML | Direct Density Ratio Optimization: A Statistically Consistent Approach to Aligning Large Language Models. | Rei Higuchi, Taiji Suzuki |
| 2025 | ICML | Mixture of Experts Provably Detect and Learn the Latent Cluster Structure in Gradient-Based Learning. | Ryotaro Kawata, Kohsei Matsutani, Yuri Kinoshita, Naoki Nishikawa, Taiji Suzuki |
| 2025 | ICML | Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation. | Juno Kim, Denny Wu, Jason D. Lee, Taiji Suzuki |
| 2025 | ICML | Nonlinear transformers can perform inference-time feature learning. | Naoki Nishikawa, Yujin Song, Kazusato Oko, Denny Wu, Taiji Suzuki |
| 2025 | ICML | Propagation of Chaos for Mean-Field Langevin Dynamics and its Application to Model Ensemble. | Atsushi Nitanda, Anzelle Lee, Damian Tan Xing Kai, Mizuki Sakaguchi, Taiji Suzuki |
| 2025 | ICML | Quantifying Memory Utilization with Effective State-Size. | Rom N. Parnichkun, Neehal Tumma, Armin W. Thomas, Alessandro Moro, Qi An, Taiji Suzuki, Atsushi Yamashita, Michael Poli, Stefano Massaroli |
| 2024 | COLT | Learning sum of diverse features: computational hardness and efficient gradient-based training for ridge combinations. | Kazusato Oko, Yujin Song, Taiji Suzuki, Denny Wu |
| 2024 | ICLR | Koopman-based generalization bound: New aspect for full-rank weights. | Yuka Hashimoto, Sho Sonoda, Isao Ishikawa, Atsushi Nitanda, Taiji Suzuki |
| 2024 | ICLR | Understanding Convergence and Generalization in Federated Learning through Feature Learning Theory. | Wei Huang, Ye Shi, Zhongyi Cai, Taiji Suzuki |
| 2024 | ICLR | Symmetric Mean-field Langevin Dynamics for Distributional Minimax Problems. | Juno Kim, Kakei Yamamoto, Kazusato Oko, Zhuoran Yang, Taiji Suzuki |
| 2024 | ICLR | Minimax optimality of convolutional neural networks for infinite dimensional input-output problems and separation from kernel methods. | Yuto Nishimura, Taiji Suzuki |
| 2024 | ICLR | Improved statistical and computational complexity of the mean-field Langevin dynamics under structured data. | Atsushi Nitanda, Kazusato Oko, Taiji Suzuki, Denny Wu |
| 2024 | ICLR | Optimal criterion for feature learning of two-layer linear neural network in high dimensional interpolation regime. | Keita Suzuki, Taiji Suzuki |
| 2024 | ICML | Provably Neural Active Learning Succeeds via Prioritizing Perplexing Samples. | Dake Bu, Wei Huang, Taiji Suzuki, Ji Cheng, Qingfu Zhang, Zhiqiang Xu, Hau-San Wong |
| 2024 | ICML | High-Dimensional Kernel Methods under Covariate Shift: Data-Dependent Implicit Regularization. | Yihang Chen, Fanghui Liu, Taiji Suzuki, Volkan Cevher |
| 2024 | ICML | Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape. | Juno Kim, Taiji Suzuki |
| 2024 | ICML | SILVER: Single-loop variance reduction and application to federated learning. | Kazusato Oko, Shunta Akiyama, Denny Wu, Tomoya Murata, Taiji Suzuki |
| 2024 | ICML | State-Free Inference of State-Space Models: The *Transfer Function* Approach. | Rom N. Parnichkun, Stefano Massaroli, Alessandro Moro, Jimmy T. H. Smith, Ramin M. Hasani, Mathias Lechner, Qi An, Christopher R, Hajime Asama, Stefano Ermon, Taiji Suzuki, Michael Poli, Atsushi Yamashita |
| 2024 | ICML | Mechanistic Design and Scaling of Hybrid Architectures. | Michael Poli, Armin W. Thomas, Eric Nguyen, Pragaash Ponnusamy, Bjrn Deiseroth, Kristian Kersting, Taiji Suzuki, Brian L. Hie, Stefano Ermon, Christopher R, Ce Zhang, Stefano Massaroli |
| 2024 | ICML | How do Transformers Perform In-Context Autoregressive Learning ? | Michael Eli Sander, Raja Giryes, Taiji Suzuki, Mathieu Blondel, Gabriel Peyr |
| 2024 | ICML | Mean-field Analysis on Two-layer Neural Networks from a Kernel Perspective. | Shokichi Takakura, Taiji Suzuki |
| 2024 | ICML | Mean Field Langevin Actor-Critic: Faster Convergence and Global Optimality beyond Lazy Learning. | Kakei Yamamoto, Kazusato Oko, Zhuoran Yang, Taiji Suzuki |
| 2024 | ICMLA | Graph Polynomial Convolution Models for Node Classification of Non-Homophilous Graphs. | Kishan Wimalawarne, Taro Sawaki, Motokiyo Hirayama, Takanobu Kawahara, Taiji Suzuki |
| 2023 | ICLR | Excess Risk of Two-Layer ReLU Neural Networks in Teacher-Student Settings and its Superiority to Kernel Methods. | Shunta Akiyama, Taiji Suzuki |
| 2023 | ICLR | Uniform-in-time propagation of chaos for the mean-field gradient Langevin dynamics. | Taiji Suzuki, Atsushi Nitanda, Denny Wu |
| 2023 | ICML | DIFF2: Differential Private Optimization via Gradient Differences for Nonconvex Distributed Learning. | Tomoya Murata, Taiji Suzuki |
| 2023 | ICML | Primal and Dual Analysis of Entropic Fictitious Play for Finite-sum Problems. | Atsushi Nitanda, Kazusato Oko, Denny Wu, Nobuhito Takenouchi, Taiji Suzuki |
| 2023 | ICML | Diffusion Models are Minimax Optimal Distribution Estimators. | Kazusato Oko, Shunta Akiyama, Taiji Suzuki |
| 2023 | ICML | Tight and fast generalization error bound of graph embedding in metric space. | Atsushi Suzuki, Atsushi Nitanda, Taiji Suzuki, Jing Wang, Feng Tian, Kenji Yamanishi |
| 2023 | ICML | Approximation and Estimation Ability of Transformers for Sequence-to-Sequence Functions with Infinite Dimensional Input. | Shokichi Takakura, Taiji Suzuki |
| 2023 | ICMLA | Scalable Federated Learning for Clients with Different Input Image Sizes and Numbers of Output Categories. | Shuhei Nitta, Taiji Suzuki, Albert Rodrguez Mulet, Atsushi Yaguchi, Ryusuke Hirai |
| 2023 | IJCNN | Neural Network Module Decomposition and Recomposition with Superimposed Masks. | Hiroaki Kingetsu, Kenichi Kobayashi, Taiji Suzuki |
| 2022 | ACML | Layer-wise Adaptive Graph Convolution Networks Using Generalized Pagerank. | Kishan Wimalawarne, Taiji Suzuki |
| 2022 | AISTATS | Convex Analysis of the Mean Field Langevin Dynamics. | Atsushi Nitanda, Denny Wu, Taiji Suzuki |
| 2022 | COLT | Dimension-free convergence rates for gradient Langevin dynamics in RKHS. | Boris Muzellec, Kanji Sato, Mathurin Massias, Taiji Suzuki |
| 2022 | ICLR | Understanding the Variance Collapse of SVGD in High Dimensions. | Jimmy Ba, Murat A. Erdogdu, Marzyeh Ghassemi, Shengyang Sun, Taiji Suzuki, Denny Wu, Tianzong Zhang |
| 2022 | ICLR | Particle Stochastic Dual Coordinate Ascent: Exponential convergent algorithm for mean field neural network optimization. | Kazusato Oko, Taiji Suzuki, Atsushi Nitanda, Denny Wu |
| 2022 | ICLR | Learnability of convolutional neural networks for infinite dimensional input via mixed and anisotropic smoothness. | Sho Okumoto, Taiji Suzuki |
| 2022 | ICMLA | Data-Parallel Momentum Diagonal Empirical Fisher (DP-MDEF):Adaptive Gradient Method is Affected by Hessian Approximation and Multi-Class Data. | Chenyuan Xu, Kosuke Haruki, Taiji Suzuki, Masahiro Ozawa, Kazuki Uematsu, Ryuji Sakai |
| 2022 | IJCNN | MSR-DARTS: Minimum Stable Rank of Differentiable Architecture Search. | Kengo Machida, Kuniaki Uto, Koichi Shinoda, Taiji Suzuki |
| 2021 | AISTATS | Gradient Descent in RKHS with Importance Labeling. | Tomoya Murata, Taiji Suzuki |
| 2021 | AISTATS | Exponential Convergence Rates of Classification Errors on Learning with SGD and Random Features. | Shingo Yashima, Atsushi Nitanda, Taiji Suzuki |
| 2021 | ICLR | When does preconditioning help or hurt generalization? | Shun-ichi Amari, Jimmy Ba, Roger Baker Grosse, Xuechen Li, Atsushi Nitanda, Taiji Suzuki, Denny Wu, Ji Xu |
| 2021 | ICLR | Optimal Rates for Averaged Stochastic Gradient Descent under Neural Tangent Kernel Regime. | Atsushi Nitanda, Taiji Suzuki |
| 2021 | ICLR | Benefit of deep learning with non-convex noisy gradient descent: Provable excess risk bound and superiority to kernel methods. | Taiji Suzuki, Shunta Akiyama |
| 2021 | ICML | On Learnability via Gradient Method for Two-Layer ReLU Neural Networks in Teacher-Student Setting. | Shunta Akiyama, Taiji Suzuki |
| 2021 | ICML | Bias-Variance Reduced Local SGD for Less Heterogeneous Federated Learning. | Tomoya Murata, Taiji Suzuki |
| 2021 | ICML | Quantitative Understanding of VAE as a Non-linearly Scaled Isometric Embedding. | Akira Nakagawa, Keizo Kato, Taiji Suzuki |
| 2021 | IJCAI | Decomposable-Net: Scalable Low-Rank Compression for Neural Networks. | Atsushi Yaguchi, Taiji Suzuki, Shuhei Nitta, Yukinobu Sakata, Akiyuki Tanizawa |
| 2020 | AISTATS | Understanding Generalization in Deep Learning via Tensor Methods. | Jingling Li, Yanchao Sun, Jiahao Su, Taiji Suzuki, Furong Huang |
| 2020 | AISTATS | Functional Gradient Boosting for Learning Residual-like Networks with Statistical Guarantees. | Atsushi Nitanda, Taiji Suzuki |
| 2020 | BMVC | Domain Adaptation Regularization for Spectral Pruning. | Laurent Dillard, Yosuke Shinya, Taiji Suzuki |
| 2020 | ICLR | Generalization of Two-layer Neural Networks: An Asymptotic Viewpoint. | Jimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Denny Wu, Tianzong Zhang |
| 2020 | ICLR | Graph Neural Networks Exponentially Lose Expressive Power for Node Classification. | Kenta Oono, Taiji Suzuki |
| 2020 | ICLR | Compression based bound for non-compressed network: unified generalization error analysis of large compressible deep neural network. | Taiji Suzuki, Hiroshi Abe, Tomoaki Nishimura |
| 2020 | IJCAI | Spectral Pruning: Compressing Deep Neural Networks via Spectral Analysis and its Generalization Error. | Taiji Suzuki, Hiroshi Abe, Tomoya Murata, Shingo Horiuchi, Kotaro Ito, Tokuma Wachi, So Hirai, Masatoshi Yukishima, Tomoaki Nishimura |
| 2019 | AISTATS | Stochastic Gradient Descent with Exponential Convergence Rates of Expected Classification Errors. | Atsushi Nitanda, Taiji Suzuki |
| 2019 | ECIR | Cross-Domain Recommendation via Deep Domain Adaptation. | Heishiro Kanagawa, Hayato Kobayashi, Nobuyuki Shimizu, Yukihiro Tagami, Taiji Suzuki |
| 2019 | ICDM | Sharp Characterization of Optimal Minibatch Size for Stochastic Finite Sum Convex Optimization. | Atsushi Nitanda, Tomoya Murata, Taiji Suzuki |
| 2019 | ICLR | Adaptivity of deep ReLU network for learning in Besov and mixed smooth Besov spaces: optimal rate and curse of dimensionality. | Taiji Suzuki |
| 2019 | ICML | Approximation and non-parametric estimation of ResNet-type convolutional neural networks. | Kenta Oono, Taiji Suzuki |
| 2018 | AISTATS | Gradient Layer: Enhancing the Convergence of Adversarial Training for Generative Models. | Atsushi Nitanda, Taiji Suzuki |
| 2018 | AISTATS | Fast generalization error bound of deep learning from a kernel perspective. | Taiji Suzuki |
| 2018 | AISTATS | Independently Interpretable Lasso: A New Regularizer for Sparse Regression with Uncorrelated Variables. | Masaaki Takada, Taiji Suzuki, Hironori Fujisawa |
| 2018 | ICML | Functional Gradient Boosting based on Residual Network Perception. | Atsushi Nitanda, Taiji Suzuki |
| 2018 | ICMLA | Adam Induces Implicit Weight Sparsity in Rectifier Neural Networks. | Atsushi Yaguchi, Taiji Suzuki, Wataru Asano, Shuhei Nitta, Yukinobu Sakata, Akiyuki Tanizawa |
| 2017 | AISTATS | Stochastic Difference of Convex Algorithm and its Application to Training Deep Boltzmann Machines. | Atsushi Nitanda, Taiji Suzuki |
| 2016 | ICML | Gaussian process nonparametric tensor estimator and its minimax optimality. | Heishiro Kanagawa, Taiji Suzuki, Hayato Kobayashi, Nobuyuki Shimizu, Yukihiro Tagami |
| 2016 | ICML | Structure Learning of Partitioned Markov Networks. | Song Liu, Taiji Suzuki, Masashi Sugiyama, Kenji Fukumizu |
| 2015 | AAAI | Support Consistency of Direct Sparse-Change Learning in Markov Networks. | Song Liu, Taiji Suzuki, Masashi Sugiyama |
| 2015 | AISTATS | A Consistent Method for Graph Based Anomaly Localization. | Satoshi Hara, Tetsuro Morimura, Toshihiro Takahashi, Hiroki Yanagisawa, Taiji Suzuki |
| 2015 | ICML | Convergence rate of Bayesian tensor estimator and its minimax optimality. | Taiji Suzuki |
| 2014 | ICML | Stochastic Dual Coordinate Ascent with Alternating Direction Method of Multipliers. | Taiji Suzuki |
| 2013 | ICML | Dual Averaging and Proximal Gradient Descent for Online Alternating Direction Multiplier Method. | Taiji Suzuki |
| 2012 | COLT | PAC-Bayesian Bound for Gaussian Process Regression and Multiple Kernel Additive Model. | Taiji Suzuki |
| 2010 | ICML | A Fast Augmented Lagrangian Algorithm for Learning Low-Rank Matrices. | Ryota Tomioka, Taiji Suzuki, Masashi Sugiyama, Hisashi Kashima |
| 2010 | SDM | Direct Density Ratio Estimation with Dimensionality Reduction. | Masashi Sugiyama, Satoshi Hara, Paul von Bnau, Taiji Suzuki, Takafumi Kanamori, Motoaki Kawanabe |
| 2009 | IDA | Estimating Squared-Loss Mutual Information for Independent Component Analysis. | Taiji Suzuki, Masashi Sugiyama |
| 2009 | ISIT | Mutual information approximation via maximum likelihood estimation of density ratio. | Taiji Suzuki, Masashi Sugiyama, Toshiyuki Tanaka |
| 2005 | ISDA | Learning to estimate user interest utilizing the variational Bayes estimator. | Taiji Suzuki, Takamasa Koshizen, Kazuyuki Aihara, Hiroshi Tsujino |