Skip to content

Michael W. Mahoney

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

124

Venues

30

Active years

2003–2026

Best venue rank

A*

Where they publish

Papers

124 indexed papers, newest first.

YearVenueTitleAuthors
2026COLTEigen-Spike Emergence and Quadratic Equivalents for Conjugate Kernels on Nonlinearly Separable Data.Collin Cranston, Zhichao Wang, Todd Kemp, Michael W. Mahoney
2026ICDETAT: Temporal-Aligned Transformer for Multi-Horizon Peak Demand Forecasting.Zhiyuan Zhao, Sitan Yang, Kin G. Olivares, Boris N. Oreshkin, Stan Vitebsky, Michael W. Mahoney, B. Aditya Prakash, Dmitry Efimov
2025ACLSqueezed Attention: Accelerating Long Context Length LLM Inference.Coleman Richard Charles Hooper, Sehoon Kim, Hiva Mohammadzadeh, Monishwaran Maheswaran, Sebastian Zhao, June Paik, Michael W. Mahoney, Kurt Keutzer, Amir Gholami
2025AISTATSGated Recurrent Neural Networks with Weighted Time-Delay Feedback.N. Benjamin Erichson, Soon Hoe Lim, Michael W. Mahoney
2025ICLRA Statistical Framework for Ranking LLM-based Chatbots.Siavash Ameli, Siyuan Zhuang, Ion Stoica, Michael W. Mahoney
2025ICLRGradient-Free Generation for Hard-Constrained Systems.Chaoran Cheng, Boran Han, Danielle C. Maddix, Abdul Fatir Ansari, Andrew Stuart, Michael W. Mahoney, Bernie Wang
2025ICLRMitigating Memorization in Language Models.Mansi Sakarvadia, Aswathy Ajith, Arham Mushtaq Khan, Nathaniel C. Hudson, Caleb Geniesse, Kyle Chard, Yaoqing Yang, Ian T. Foster, Michael W. Mahoney
2025ICLRTuning Frequency Bias of State Space Models.Annan Yu, Dongwei Lyu, Soon Hoe Lim, Michael W. Mahoney, N. Benjamin Erichson
2025ICLRHOPE for a Robust Parameterization of Long-memory State Space Models.Annan Yu, Michael W. Mahoney, N. Benjamin Erichson
2025ICMLDeterminant Estimation under Memory Constraints and Neural Scaling Laws.Siavash Ameli, Chris van der Heide, Liam Hodgkinson, Fred Roosta, Michael W. Mahoney
2025ICMLModels of Heavy-Tailed Mechanistic Universality.Liam Hodgkinson, Zhichao Wang, Michael W. Mahoney
2025ICMLEnhancing Foundation Models for Time Series Forecasting via Wavelet-based Tokenization.Luca Masserano, Abdul Fatir Ansari, Boran Han, Xiyuan Zhang, Christos Faloutsos, Michael W. Mahoney, Andrew Gordon Wilson, Youngsuk Park, Syama Sundar Rangapuram, Danielle C. Maddix, Bernie Wang
2025ICMLFundamental Bias in Inverting Random Sampling Matrices with Application to Sub-sampled Newton.Chengmei Niu, Zhenyu Liao, Zenan Ling, Michael W. Mahoney
2025ICMLQuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache.Rishabh Tiwari, Haocheng Xi, Aditya Tomar, Coleman Richard Charles Hooper, Sehoon Kim, Maxwell Horton, Mahyar Najibi, Michael W. Mahoney, Kurt Keutzer, Amir Gholami
2024ACLLLM2LLM: Boosting LLMs with Novel Iterative Data Enhancement.Nicholas Lee, Thanakul Wattanawong, Sehoon Kim, Karttikeya Mangalam, Sheng Shen, Gopala Anumanchipalli, Michael W. Mahoney, Kurt Keutzer, Amir Gholami
2024AISTATSNoisyMix: Boosting Model Robustness to Common Corruptions.N. Benjamin Erichson, Soon Hoe Lim, Winnie Xu, Francisco Utrera, Ziang Cao, Michael W. Mahoney
2024AISTATSEquation Discovery with Bayesian Spike-and-Slab Priors and Efficient Kernels.Da Long, Wei W. Xing, Aditi S. Krishnapriyan, Robert M. Kirby, Shandian Zhe, Michael W. Mahoney
2024ICLRGenerative Modeling of Regular and Irregular Time Series Data via Koopman VAEs.Ilan Naiman, N. Benjamin Erichson, Pu Ren, Michael W. Mahoney, Omri Azencot
2024ICLRRobustifying State-space Models for Long Sequences via Approximate Diagonalization.Annan Yu, Arnur Nigmetov, Dmitriy Morozov, Michael W. Mahoney, N. Benjamin Erichson
2024ICMLSqueezeLLM: Dense-and-Sparse Quantization.Sehoon Kim, Coleman Hooper, Amir Gholami, Zhen Dong, Xiuyu Li, Sheng Shen, Michael W. Mahoney, Kurt Keutzer
2024ICMLAn LLM Compiler for Parallel Function Calling.Sehoon Kim, Suhong Moon, Ryan Tabrizi, Nicholas Lee, Michael W. Mahoney, Kurt Keutzer, Amir Gholami
2024ICMLUsing Uncertainty Quantification to Characterize and Improve Out-of-Domain Learning for PDEs.S. Chandra Mouli, Danielle C. Maddix, Shima Alizadeh, Gaurav Gupta, Andrew Stuart, Michael W. Mahoney, Bernie Wang
2024ICMLTowards Scalable and Versatile Weight Space Learning.Konstantin Schrholt, Michael W. Mahoney, Damian Borth
2024KDDRecent and Upcoming Developments in Randomized Numerical Linear Algebra for Machine Learning.Michal Derezinski, Michael W. Mahoney
2024VTSReliable edge machine learning hardware for scientific applications.Tommaso Baldi, Javier Campos, Benjamin Hawks, Jennifer Ngadiuba, Nhan Tran, Daniel Diaz, Javier M. Duarte, Ryan Kastner, Andres Meza, Melissa Quinnan, Olivia Weng, Caleb Geniesse, Amir Gholami, Michael W. Mahoney, Vladimir Loncar, Philip C. Harris, Joshua Agar, Shuyu Qin
2023AISTATSFast Feature Selection with Fairness Constraints.Francesco Quinzan, Rajiv Khanna, Moshik Hershcovitch, Sarel Cohen, Daniel G. Waddington, Tobias Friedrich, Michael W. Mahoney
2023ECAIAdaptive Self-Supervision Algorithms for Physics-Informed Neural Networks.Shashank Subramanian, Robert M. Kirby, Michael W. Mahoney, Amir Gholami
2023ICLRLearning differentiable solvers for systems with hard constraints.Geoffrey Ngiar, Michael W. Mahoney, Aditi S. Krishnapriyan
2023ICLRGradient Gating for Deep Multi-Rate Learning on Graphs.T. Konstantin Rusch, Benjamin Paul Chamberlain, Michael W. Mahoney, Michael M. Bronstein, Siddhartha Mishra
2023ICMLLearning Physical Models that Can Respect Conservation Laws.Derek Hansen, Danielle C. Maddix, Shima Alizadeh, Gaurav Gupta, Michael W. Mahoney
2023ICMLMonotonicity and Double Descent in Uncertainty Estimation with Gaussian Processes.Liam Hodgkinson, Christopher van der Heide, Fred Roosta, Michael W. Mahoney
2023ICMLConstrained Optimization via Exact Augmented Lagrangian and Randomized Iterative Sketching.Ilgee Hong, Sen Na, Michael W. Mahoney, Mladen Kolar
2023ICMLA Three-regime Model of Network Pruning.Yefan Zhou, Yaoqing Yang, Arin Chang, Michael W. Mahoney
2023KDDTest Accuracy vs. Generalization Gap: Model Selection in NLP without Accessing Training or Testing Data.Yaoqing Yang, Ryan Theisen, Liam Hodgkinson, Joseph E. Gonzalez, Kannan Ramchandran, Charles H. Martin, Michael W. Mahoney
2023SCExtensions to the SENSEI In situ Framework for Heterogeneous Architectures.Burlen Loring, E. Wes Bethel, Gunther H. Weber, Michael W. Mahoney
2022ICASSPInteger-Only Zero-Shot Quantization for Efficient Speech Recognition.Sehoon Kim, Amir Gholami, Zhewei Yao, Nicholas Lee, Patrick Wang, Aniruddha Nrusimha, Bohan Zhai, Tianren Gao, Michael W. Mahoney, Kurt Keutzer
2022ICLRDoubly Adaptive Scaled Algorithm for Machine Learning Using Second-Order Information.Majid Jahani, Sergey Rusakov, Zheng Shi, Peter Richtrik, Michael W. Mahoney, Martin Takc
2022ICLRNoisy Feature Mixup.Soon Hoe Lim, N. Benjamin Erichson, Francisco Utrera, Winnie Xu, Michael W. Mahoney
2022ICLRLong Expressive Memory for Sequence Modeling.T. Konstantin Rusch, Siddhartha Mishra, N. Benjamin Erichson, Michael W. Mahoney
2022ICMLGeneralization Bounds using Lower Tail Exponents in Stochastic Optimizers.Liam Hodgkinson, Umut Simsekli, Rajiv Khanna, Michael W. Mahoney
2022ICMLFat-Tailed Variational Inference with Anisotropic Tail Adaptive Flows.Feynman T. Liang, Michael W. Mahoney, Liam Hodgkinson
2022ICMLGACT: Activation Compressed Training for Generic Network Architectures.Xiaoxuan Liu, Lianmin Zheng, Dequan Wang, Yukuo Cen, Weize Chen, Xu Han, Jianfei Chen, Zhiyuan Liu, Jie Tang, Joey Gonzalez, Michael W. Mahoney, Alvin Cheung
2022ICMLAutoIP: A United Framework to Integrate Physics into Gaussian Processes.Da Long, Zheng Wang, Aditi S. Krishnapriyan, Robert M. Kirby, Shandian Zhe, Michael W. Mahoney
2022ICMLNeurotoxin: Durable Backdoors in Federated Learning.Zhengming Zhang, Ashwinee Panda, Linyue Song, Yaoqing Yang, Michael W. Mahoney, Prateek Mittal, Kannan Ramchandran, Joseph Gonzalez
2022WACVHessian-Aware Pruning and Optimal Neural Implant.Shixing Yu, Zhewei Yao, Amir Gholami, Zhen Dong, Sehoon Kim, Michael W. Mahoney, Kurt Keutzer
2021AAAIADAHESSIAN: An Adaptive Second Order Optimizer for Machine Learning.Zhewei Yao, Amir Gholami, Sheng Shen, Mustafa Mustafa, Kurt Keutzer, Michael W. Mahoney
2021AISTATSGood Classifiers are Abundant in the Interpolating Regime.Ryan Theisen, Jason M. Klusowski, Michael W. Mahoney
2021COLTSparse sketches with small inversion bias.Michal Derezinski, Zhenyu Liao, Edgar Dobriban, Michael W. Mahoney
2021EMNLPWhat's Hidden in a One-layer Randomly Weighted Transformer?Sheng Shen, Zhewei Yao, Douwe Kiela, Kurt Keutzer, Michael W. Mahoney
2021ICLRLipschitz Recurrent Neural Networks.N. Benjamin Erichson, Omri Azencot, Alejandro F. Queiruga, Liam Hodgkinson, Michael W. Mahoney
2021ICLRSparse Quantized Spectral Clustering.Zhenyu Liao, Romain Couillet, Michael W. Mahoney
2021ICLRAdversarially-Trained Deep Nets Transfer Better: Illustration on Image Classification.Francisco Utrera, Evan Kravitz, N. Benjamin Erichson, Rajiv Khanna, Michael W. Mahoney
2021ICMLActNN: Reducing Training Memory Footprint via 2-Bit Activation Compressed Training.Jianfei Chen, Lianmin Zheng, Zhewei Yao, Dequan Wang, Ion Stoica, Michael W. Mahoney, Joseph Gonzalez
2021ICMLMultiplicative Noise and Heavy Tails in Stochastic Optimization.Liam Hodgkinson, Michael W. Mahoney
2021ICMLI-BERT: Integer-only BERT Quantization.Sehoon Kim, Amir Gholami, Zhewei Yao, Michael W. Mahoney, Kurt Keutzer
2021ICMLHAWQ-V3: Dyadic Neural Network Quantization.Zhewei Yao, Zhen Dong, Zhangcheng Zheng, Amir Gholami, Jiali Yu, Eric Tan, Leyuan Wang, Qijing Huang, Yida Wang, Michael W. Mahoney, Kurt Keutzer
2021IJCAIImproved Guarantees and a Multiple-descent Curve for Column Subset Selection and the Nystrom Method (Extended Abstract).Michal Derezinski, Rajiv Khanna, Michael W. Mahoney
2021KDDTraining Recommender Systems at Scale: Communication-Efficient Model and Data Parallelism.Vipul Gupta, Dhruv Choudhary, Ping Tak Peter Tang, Xiaohan Wei, Xing Wang, Yuzhen Huang, Arun Kejariwal, Kannan Ramchandran, Michael W. Mahoney
2021UAILocalNewton: Reducing communication rounds for distributed learning.Vipul Gupta, Avishek Ghosh, Michal Derezinski, Rajiv Khanna, Kannan Ramchandran, Michael W. Mahoney
2021UAIStochastic continuous normalizing flows: training SDEs as ODEs.Liam Hodgkinson, Christopher van der Heide, Fred Roosta, Michael W. Mahoney
2021UAIGeometric rates of convergence for kernel-based sampling algorithms.Rajiv Khanna, Liam Hodgkinson, Michael W. Mahoney
2021SDMNoise-Response Analysis of Deep Neural Networks Quantifies Robustness and Fingerprints Structural Malware.N. Benjamin Erichson, Dane Taylor, Qixuan Wu, Michael W. Mahoney
2020AAAIInefficiency of K-FAC for Large Batch Size Training.Linjian Ma, Gabe Montague, Jiayu Ye, Zhewei Yao, Amir Gholami, Kurt Keutzer, Michael W. Mahoney
2020AAAIQ-BERT: Hessian Based Ultra Low Precision Quantization of BERT.Sheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma, Zhewei Yao, Amir Gholami, Michael W. Mahoney, Kurt Keutzer
2020AISTATSBayesian experimental design using regularized determinantal point processes.Michal Derezinski, Feynman T. Liang, Michael W. Mahoney
2020AISTATSStatistical guarantees for local graph clustering.Wooseok Ha, Kimon Fountoulakis, Michael W. Mahoney
2020AISTATSAsymptotic Analysis of Sampling Estimators for Randomized Numerical Linear Algebra Algorithms.Ping Ma, Xinlian Zhang, Xin Xing, Jingyi Ma, Michael W. Mahoney
2020CVPRZeroQ: A Novel Zero Shot Quantization Framework.Yaohui Cai, Zhewei Yao, Zhen Dong, Amir Gholami, Michael W. Mahoney, Kurt Keutzer
2020EMNLPMAF: Multimodal Alignment Framework for Weakly-Supervised Phrase Grounding.Qinxin Wang, Hao Tan, Sheng Shen, Michael W. Mahoney, Zhewei Yao
2020ICMLForecasting Sequential Data Using Consistent Koopman Autoencoders.Omri Azencot, N. Benjamin Erichson, Vanessa Lin, Michael W. Mahoney
2020ICMLError Estimation for Sketched SVD via the Bootstrap.Miles E. Lopes, N. Benjamin Erichson, Michael W. Mahoney
2020ICMLPowerNorm: Rethinking Batch Normalization in Transformers.Sheng Shen, Zhewei Yao, Amir Gholami, Michael W. Mahoney, Kurt Keutzer
2020ICPRAMJumpReLU: A Retrofit Defense Strategy for Adversarial Attacks.N. Benjamin Erichson, Zhewei Yao, Michael W. Mahoney
2020SCNewton-ADMM: a distributed GPU-accelerated optimizer for multiclass classification problems.Chih-Hao Fang, Sudhir B. Kylasa, Fred Roosta, Michael W. Mahoney, Ananth Grama
2020SDMHeavy-Tailed Universality Predicts Trends in Test Accuracies for Very Large Pre-Trained Deep Neural Networks.Charles H. Martin, Michael W. Mahoney
2020SDMSecond-order Optimization for Non-convex Machine Learning: an Empirical Study.Peng Xu, Fred Roosta, Michael W. Mahoney
2019COLTMinimax experimental design: Bridging the gap between statistical and worst-case approaches to least squares regression.Michal Derezinski, Kenneth L. Clarkson, Michael W. Mahoney, Manfred K. Warmuth
2019CVPRTrust Region Based Adversarial Attack on Neural Networks.Zhewei Yao, Amir Gholami, Peng Xu, Kurt Keutzer, Michael W. Mahoney
2019ICCVHAWQ: Hessian AWare Quantization of Neural Networks With Mixed-Precision.Zhen Dong, Zhewei Yao, Amir Gholami, Michael W. Mahoney, Kurt Keutzer
2019ICMLTraditional and Heavy Tailed Self Regularization in Neural Network Models.Michael W. Mahoney, Charles H. Martin
2019KDDStatistical Mechanics Methods for Discovering Knowledge from Modern Production Quality Neural Networks.Charles H. Martin, Michael W. Mahoney
2019SDMGPU Accelerated Sub-Sampled Newton's Method for Convex Classification Problems.Sudhir B. Kylasa, Fred (Farbod) Roosta, Michael W. Mahoney, Ananth Grama
2018AISTATSFLAG n' FLARE: Fast Linearly-Coupled Adaptive Gradient Methods.Xiang Cheng, Fred (Farbod) Roosta, Stefan Palombo, Peter L. Bartlett, Michael W. Mahoney
2018ICMLOut-of-sample extension of graph adjacency spectral embedding.Keith D. Levin, Farbod Roosta-Khorasani, Michael W. Mahoney, Carey E. Priebe
2018ICMLError Estimation for Randomized Least-Squares Algorithms via the Bootstrap.Miles E. Lopes, Shusen Wang, Michael W. Mahoney
2018KDDAccelerating Large-Scale Data Analysis by Offloading to High-Performance Computing Libraries using Alchemist.Alex Gittens, Kai Rothauge, Shusen Wang, Michael W. Mahoney, Lisa Gerhardt, Prabhat, Jey Kottalam, Michael F. Ringenburg, Kristyn J. Maschhoff
2017ACLSkip-Gram - Zipf + Uniform = Vector Additivity.Alex Gittens, Dimitris Achlioptas, Michael W. Mahoney
2017ICMLCapacity Releasing Diffusion for Speed and Locality.Di Wang, Kimon Fountoulakis, Monika Henzinger, Michael W. Mahoney, Satish Rao
2017ICMLSketched Ridge Regression: Optimization Perspective, Statistical Perspective, and Model Averaging.Shusen Wang, Alex Gittens, Michael W. Mahoney
2016ICALPApproximating the Solution to Mixed Packing and Covering LPs in Parallel O˜(epsilon^{-3}) Time.Michael W. Mahoney, Satish Rao, Di Wang, Peng Zhang
2016ICALPUnified Acceleration Method for Packing and Covering Problems via Diameter Reduction.Di Wang, Satish Rao, Michael W. Mahoney
2016ICMLA Simple and Strongly-Local Flow-Based Method for Cut Improvement.Nate Veldt, David F. Gleich, Michael W. Mahoney
2016SODAWeighted SGD forJiyan Yang, Yinlam Chow, Christopher R, Michael W. Mahoney
2015AISTATSSpectral Gap Error Bounds for Improving CUR Matrix Decomposition and the Nystrm Method.David G. Anderson, Simon S. Du, Michael W. Mahoney, Christopher Melgaard, Kunming Wu, Ming Gu
2015ICMLStatistical and Algorithmic Perspectives on Randomized Sketching for Ordinary Least-Squares.Garvesh Raskutti, Michael W. Mahoney
2015KDDUsing Local Spectral Methods to Robustify Graph-Based Learning Algorithms.David F. Gleich, Michael W. Mahoney
2014CVPRRandom Laplace Feature Maps for Semigroup Kernels on Histograms.Jiyan Yang, Vikas Sindhwani, Quanfu Fan, Haim Avron, Michael W. Mahoney
2014ICMLAnti-differentiating approximation algorithms: A case study with min-cuts, spectral, and flow.David F. Gleich, Michael W. Mahoney
2014ICMLA Statistical Perspective on Algorithmic Leveraging.Ping Ma, Michael W. Mahoney, Bin Yu
2014ICMLQuasi-Monte Carlo Feature Maps for Shift-Invariant Kernels.Jiyan Yang, Vikas Sindhwani, Haim Avron, Michael W. Mahoney
2013ICDMTree-Like Structure in Large Social and Information Networks.Aaron B. Adcock, Blair D. Sullivan, Michael W. Mahoney
2013ICMLRevisiting the Nystrom method for improved large-scale machine learning.Alex Gittens, Michael W. Mahoney
2013ICMLRobust Regression on MapReduce.Xiangrui Meng, Michael W. Mahoney
2013ICMLQuantile Regression for Large-scale Applications.Jiyan Yang, Xiangrui Meng, Michael W. Mahoney
2013SODAThe Fast Cauchy Transform and Faster Robust Linear Regression.Kenneth L. Clarkson, Petros Drineas, Malik Magdon-Ismail, Michael W. Mahoney, Xiangrui Meng, David P. Woodruff
2013STOCLow-distortion subspace embeddings in input-sparsity time and applications to robust linear regression.Xiangrui Meng, Michael W. Mahoney
2012ICMLFast approximation of matrix coherence and statistical leverage.Michael W. Mahoney, Petros Drineas, Malik Magdon-Ismail, David P. Woodruff
2012ISAACOn the Hyperbolicity of Small-World and Tree-Like Random Graphs.Wei Chen, Wenjie Fang, Guangda Hu, Michael W. Mahoney
2012PODSApproximate computation and implicit regularization for very large-scale data analysis.Michael W. Mahoney
2011ICMLImplementing regularization implicitly via approximate eigenvector computation.Michael W. Mahoney, Lorenzo Orecchia
2010WWWEmpirical comparison of algorithms for network community detection.Jure Leskovec, Kevin J. Lang, Michael W. Mahoney
2010UAIApproximating Higher-Order Distances Using Random Projections.Ping Li, Michael W. Mahoney, Yiyuan She
2009SODAAn improved approximation algorithm for the column subset selection problem.Christos Boutsidis, Michael W. Mahoney, Petros Drineas
2008KDDUnsupervised feature selection for principal components analysis.Christos Boutsidis, Michael W. Mahoney, Petros Drineas
2008WWWStatistical properties of community structure in large social and information networks.Jure Leskovec, Kevin J. Lang, Anirban Dasgupta, Michael W. Mahoney
2008SODASampling algorithms and coresets for ℓAnirban Dasgupta, Petros Drineas, Boulos Harb, Ravi Kumar, Michael W. Mahoney
2007KDDFeature selection methods for text classification.Anirban Dasgupta, Petros Drineas, Boulos Harb, Vanja Josifovski, Michael W. Mahoney
2006ESASubspace Sampling and Relative-Error Matrix Approximation: Column-Row-Based Methods.Petros Drineas, Michael W. Mahoney, S. Muthukrishnan
2006KDDTensor-CUR decompositions for tensor-based data.Michael W. Mahoney, Mauro Maggioni, Petros Drineas
2006SODASampling algorithms forPetros Drineas, Michael W. Mahoney, S. Muthukrishnan
2006VLDBRandomized Algorithms for Matrices and Massive Data Sets.Petros Drineas, Michael W. Mahoney
2005COLTApproximating a Gram Matrix for Improved Kernel-Based Learning.Petros Drineas, Michael W. Mahoney
2005STACSSampling Sub-problems of Heterogeneous Max-cut Problems and Approximation Algorithms.Petros Drineas, Ravi Kannan, Michael W. Mahoney
2003ISAACRapid Mixing of Several Markov Chains for a Hard-Core Model.Ravi Kannan, Michael W. Mahoney, Ravi Montenegro