Skip to content

Yuxiong He

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

78

Venues

36

Active years

2006–2026

Best venue rank

A*

Where they publish

Papers

78 indexed papers, newest first.

YearVenueTitleAuthors
2026ACLR³-SQL: Ranking Reward and Resampling for Text-to-SQL.Hojae Han, Yeonseok Jeong, Seung-won Hwang, Zhewei Yao, Yuxiong He
2026ACLAgentic Verification for Ambiguous Query Disambiguation.Youngwon Lee, Seung-won Hwang, Ruofan Wu, Feng Yan, Danmei Xu, Moutasem Akkad, Zhewei Yao, Yuxiong He
2026ACLGRAD: Generalizing RAG Adaptation with Decoding.Youngwon Lee, Seung-won Hwang, Zhewei Yao, Yuxiong He
2026ACLArctic-Text2SQL-R1: Simple Rewards, Strong Reasoning in Text-to-SQL.Zhewei Yao, Guoheng Sun, Lukasz Borchmann, Zheyu Shen, Minghang Deng, Bohan Zhai, Hao Zhang, Ang Li, Yuxiong He
2026ASPLOSShift Parallelism: Low-Latency, High-Throughput LLM Inference for Dynamic Workloads.Mert Hidayetoglu, Aurick Qiao, Michael Wyatt, Jeff Rasley, Yuxiong He, Samyam Rajbhandari
2026EACLTAGQuant: Token-Aware Clustering for Group-Wise Quantization.Jaeseong Lee, Seung-won Hwang, Aurick Qiao, Zhewei Yao, Yuxiong He
2026ICDCSDySkew: Dynamic Data Redistribution for Skew- Resilient Snowpark UDF Execution.Chenwei Xie, Urjeet Shrestha, Corbin, Lukas Lorimer, Gopal V, Zihao Ye, Yi Pan, Nic Crouch, Elliott Brossard, Florian Funke, Yuxiong He
2025ACLSTUN: Structured-Then-Unstructured Pruning for Scalable MoE Pruning.Jaeseong Lee, Seung-won Hwang, Aurick Qiao, Daniel F. Campos, Zhewei Yao, Yuxiong He
2025ACLOptimizing Reasoning for Text-to-SQL with Execution Feedback.Bohan Zhai, Canwen Xu, Yuxiong He, Zhewei Yao
2025EMNLPSwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation.Aurick Qiao, Zhewei Yao, Samyam Rajbhandari, Yuxiong He
2025ICLRConvCodeWorld: Benchmarking Conversational Code Generation in Reproducible Feedback Environments.Hojae Han, Seung-won Hwang, Rajhans Samdani, Yuxiong He
2025NAACLInference Scaling for Bridging Retrieval and Augmented Generation.Youngwon Lee, Seung-won Hwang, Daniel F. Campos, Filip Gralinski, Zhewei Yao, Yuxiong He
2025NAACLCORD: Balancing COnsistency and Rank Distillation for Robust Retrieval-Augmented Generation.Youngwon Lee, Seung-won Hwang, Daniel F. Campos, Filip Gralinski, Zhewei Yao, Yuxiong He
2024AAAIDeepSpeed Data Efficiency: Improving Deep Learning Model Quality and Training Efficiency via Efficient Data Sampling and Routing.Conglong Li, Zhewei Yao, Xiaoxia Wu, Minjia Zhang, Connor Holmes, Cheng Li, Yuxiong He
2024AAAIExploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation.Zhewei Yao, Xiaoxia Wu, Cheng Li, Stephen Youn, Yuxiong He
2024ICLRZeRO++: Extremely Efficient Collective Communication for Large Model Training.Guanhua Wang, Heyang Qin, Sam Ade Jacobs, Xiaoxia Wu, Connor Holmes, Zhewei Yao, Samyam Rajbhandari, Olatunji Ruwase, Feng Yan, Lei Yang, Yuxiong He
2024PODCSystem Optimizations for Enabling Training of Extreme Long Sequence Transformer Models.Sam Ade Jacobs, Masahiro Tanaka, Chengming Zhang, Minjia Zhang, Reza Yazdani Aminadabi, Shuaiwen Leon Song, Samyam Rajbhandari, Yuxiong He
2024USENIXQuant-LLM: Accelerating the Serving of Large Language Models via FP6-Centric Algorithm-System Co-Design on Modern GPUs.Haojun Xia, Zhen Zheng, Xiaoxia Wu, Shiyang Chen, Zhewei Yao, Stephen Youn, Arash Bakhtiari, Michael Wyatt, Donglin Zhuang, Zhongzhu Zhou, Olatunji Ruwase, Yuxiong He, Shuaiwen Leon Song
2023ECAIRevisiting the Efficiency-Accuracy Tradeoff in Adapting Transformer Models via Adversarial Fine-Tuning.Minjia Zhang, Uma-Naresh Niranjan, Yuxiong He
2023EMNLPScaling Vision-Language Models with Sparse Mixture of Experts.Sheng Shen, Zhewei Yao, Chunyuan Li, Trevor Darrell, Kurt Keutzer, Yuxiong He
2023ICLRMaximizing Communication Efficiency for Large-scale Training via 0/1 Adam.Yucheng Lu, Conglong Li, Minjia Zhang, Christopher De Sa, Yuxiong He
2023ICLRDySR: Adaptive Super-Resolution via Algorithm and System Co-design.Syed Zawad, Cheng Li, Zhewei Yao, Elton Zheng, Yuxiong He, Feng Yan
2023ICMLUnderstanding Int4 Quantization for Language Models: Latency Speedup, Composability, and Failure Cases.Xiaoxia Wu, Cheng Li, Reza Yazdani Aminabadi, Zhewei Yao, Yuxiong He
2023ICSHEAT: A Highly Efficient and Affordable Training System for Collaborative Filtering Based Recommendation on CPUs.Chengming Zhang, Shaden Smith, Baixi Sun, Jiannan Tian, Jonathan Soifer, Xiaodong Yu, Shuaiwen Leon Song, Yuxiong He, Dingwen Tao
2023ICSA Hybrid Tensor-Expert-Data Parallelism Approach to Optimize Mixture-of-Experts Training.Siddharth Singh, Olatunji Ruwase, Ammar Ahmad Awan, Samyam Rajbhandari, Yuxiong He, Abhinav Bhatele
2022AAAIAdversarial Data Augmentation for Task-Specific Knowledge Distillation of Pre-trained Transformers.Minjia Zhang, Uma-Naresh Niranjan, Yuxiong He
2022HiPC1-bit LAMB: Communication Efficient Large-Scale Large-Batch Training with LAMB's Convergence Speed.Conglong Li, Ammar Ahmad Awan, Hanlin Tang, Samyam Rajbhandari, Yuxiong He
2022ICMLDeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale.Samyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang, Reza Yazdani Aminabadi, Ammar Ahmad Awan, Jeff Rasley, Yuxiong He
2022SCDeepSpeed- Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale.Reza Yazdani Aminabadi, Samyam Rajbhandari, Ammar Ahmad Awan, Cheng Li, Du Li, Elton Zheng, Olatunji Ruwase, Shaden Smith, Minjia Zhang, Jeff Rasley, Yuxiong He
2022WSDMGraSP: Optimizing Graph-based Nearest Neighbor Search with Subgraph Sampling and Pruning.Minjia Zhang, Wenhan Wang, Yuxiong He
2021ICML1-bit Adam: Communication Efficient Large-Scale Training with Adam's Convergence Speed.Hanlin Tang, Shaoduo Gan, Ammar Ahmad Awan, Samyam Rajbhandari, Conglong Li, Xiangru Lian, Ji Liu, Ce Zhang, Yuxiong He
2021SCZeRO-infinity: breaking the GPU memory wall for extreme scale deep learning.Samyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith, Yuxiong He
2021USENIXZeRO-Offload: Democratizing Billion-Scale Model Training.Jie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, Yuxiong He
2020KDDDeepSpeed: System Optimizations Enable Training Deep Learning Models with Over 100 Billion Parameters.Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, Yuxiong He
2020SCZeRO: memory optimizations toward training trillion parameter models.Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong He
2020SIGMODImproving Approximate Nearest Neighbor Search through Learned Adaptive Early Termination.Conglong Li, Minjia Zhang, David G. Andersen, Yuxiong He
2019CIKMGRIP: Multi-Store Capacity-Optimized High-Performance Nearest Neighbor Search for Vector Search Engine.Minjia Zhang, Yuxiong He
2019EuroSysGRNN: Low-Latency and Scalable RNN Inference on GPUs.Connor Holmes, Daniel Mawhirter, Yuxiong He, Feng Yan, Bo Wu
2019ICDMFast LSTM Inference by Dynamic Decomposition on Cloud Systems.Yang You, Yuxiong He, Samyam Rajbhandari, Wenhan Wang, Cho-Jui Hsieh, Kurt Keutzer, James Demmel
2018ICLRLearning Intrinsic Sparse Structures within Long Short-Term Memory.Wei Wen, Yuxiong He, Samyam Rajbhandari, Minjia Zhang, Wenhan Wang, Fang Liu, Bin Hu, Yiran Chen, Hai Li
2018WWWBetter Caching in Search Advertising Systems with Rapid Refresh Predictions.Conglong Li, David G. Andersen, Qiang Fu, Sameh Elnikety, Yuxiong He
2018USENIXDeepCPU: Serving RNN-based Deep Learning Models 10x Faster.Minjia Zhang, Samyam Rajbhandari, Wenhan Wang, Yuxiong He
2017ASPLOSOptimizing CNNs on Multicores for Scalability, Performance and Goodput.Samyam Rajbhandari, Yuxiong He, Olatunji Ruwase, Michael Carbin, Trishul M. Chilimbi
2017CLOUDWorkload analysis and caching strategies for search advertising systems.Conglong Li, David G. Andersen, Qiang Fu, Sameh Elnikety, Yuxiong He
2017MICROExploiting heterogeneity for tail latency and energy efficiency.Md. Enamul Haque, Yuxiong He, Sameh Elnikety, Thu D. Nguyen, Ricardo Bianchini, Kathryn S. McKinley
2017MiddlewareSwayam: distributed autoscaling to meet SLAs of machine learning inference services with resource efficiency.Arpan Gujarati, Sameh Elnikety, Yuxiong He, Kathryn S. McKinley, Bjrn B. Brandenburg
2017MiddlewareHyperDrive: exploring hyperparameters with POP scheduling.Jeff Rasley, Yuxiong He, Feng Yan, Olatunji Ruwase, Rodrigo Fonseca
2017SIGIRBitFunnel: Revisiting Signatures for Search.Bob Goodwin, Michael Hopcroft, Dan Luu, Alex Clemmer, Mihaela Curmei, Sameh Elnikety, Yuxiong He
2017SPAAOptimal Reissue Policies for Reducing Tail Latency.Tim Kaler, Yuxiong He, Sameh Elnikety
2016ASPLOSTPC: Target-Driven Parallelism Combining Prediction and Correction to Reduce Tail Latency in Interactive Services.Myeongjae Jeon, Yuxiong He, Hwanju Kim, Sameh Elnikety, Scott Rixner, Alan L. Cox
2016PPoPPWork stealing for interactive services to meet target latency.Jing Li, Kunal Agrawal, Sameh Elnikety, Yuxiong He, I-Ting Angelina Lee, Chenyang Lu, Kathryn S. McKinley
2016SCSERF: efficient scheduling for fast deep neural network serving via judicious parallelism.Feng Yan, Yuxiong He, Olatunji Ruwase, Evgenia Smirni
2015ASPLOSFew-to-Many: Incremental Parallelism for Reducing Tail Latency in Interactive Services.Md. Enamul Haque, Yong Hun Eom, Yuxiong He, Sameh Elnikety, Ricardo Bianchini, Kathryn S. McKinley
2015KDDPerformance Modeling and Scalability Optimization of Distributed Deep Learning Systems.Feng Yan, Olatunji Ruwase, Yuxiong He, Trishul M. Chilimbi
2015MASCOTSBATS: Budget-Constrained Autoscaling for Cloud Performance Optimization.A. Hasan Mahmud, Yuxiong He, Shaolei Ren
2015SIGIROptimal Aggregation Policy for Reducing Tail Latency of Web Search.Jeong-Min Yun, Yuxiong He, Sameh Elnikety, Shaolei Ren
2015WSDMDelayed-Dynamic-Selective (DDS) Prediction for Reducing Extreme Tail Latency in Web Search.Saehoon Kim, Yuxiong He, Seung-won Hwang, Sameh Elnikety, Seungjin Choi
2014ICDEMars: Real-time spatio-temporal queries on microblogs.Amr Magdy, Ahmed M. Aly, Mohamed F. Mokbel, Sameh Elnikety, Yuxiong He, Suman Nath
2014ICDEMercury: A memory-constrained spatio-temporal real-time search on microblogs.Amr Magdy, Mohamed F. Mokbel, Sameh Elnikety, Suman Nath, Yuxiong He
2014SIGIRPredictive parallelization: taming tail latencies in web search.Myeongjae Jeon, Saehoon Kim, Seung-won Hwang, Yuxiong He, Sameh Elnikety, Alan L. Cox, Scott Rixner
2014SIGMETRICSBATS: budget-constrained autoscaling for cloud performance optimization.A. Hasan Mahmud, Yuxiong He, Shaolei Ren
2013EuroParTopic 3: Scheduling and Load Balancing - (Introduction).Zhihui Du, Ramin Yahyapour, Yuxiong He, Nectarios Koziris, Bilha Mendelson, Veronika Sonigo, Achim Streit, Andrei Tchernykh
2013EuroSysAdaptive parallelism for web search.Myeongjae Jeon, Yuxiong He, Sameh Elnikety, Alan L. Cox, Scott Rixner
2013IMPower-effiicent resource allocation in MapReduce clusters.Kaiqi Xiong, Yuxiong He
2013SCCOCA: online distributed resource management for cost minimization and carbon neutrality in data centers.Shaolei Ren, Yuxiong He
2013SPIRESolving Graph Isomorphism Using Parameterized Matching.Juan Mendivelso, Sunghwan Kim, Sameh Elnikety, Yuxiong He, Seung-won Hwang, Yoan J. Pinzn
2012CIKMG-SPARQL: a hybrid engine for querying large attributed graphs.Sherif Sakr, Sameh Elnikety, Yuxiong He
2012CLOUDZeta: scheduling interactive services with partial execution.Yuxiong He, Sameh Elnikety, James R. Larus, Chenyu Yan
2012ICDCSProvably-Efficient Job Scheduling for Energy and Fairness in Geographically Distributed Data Centers.Shaolei Ren, Yuxiong He, Fei Xu
2012ICDEHorton: Online Query Execution Engine for Large Distributed Graphs.Mohamed Sarwat, Sameh Elnikety, Yuxiong He, Gabriel Kliot
2011AAAIPosition Paper: Embracing Heterogeneity - Improving Energy Efficiency for Interactive Services on Heterogeneous Data Center Hardware.Yuxiong He, Sameh Elnikety
2011ICDCSTians Scheduling: Using Partial Processing in Best-Effort Applications.Yuxiong He, Sameh Elnikety, Hongyang Sun
2010SPAAThe Cilkview scalability analyzer.Yuxiong He, Charles E. Leiserson, William M. Leiserson
2007ICPPAdaptive Scheduling of Parallel Jobs on Functionally Heterogeneous Resources.Yuxiong He, Hongyang Sun, Wen-Jing Hsu
2007PPoPPAdaptive work stealing with parallelism feedback.Kunal Agrawal, Yuxiong He, Charles E. Leiserson
2006ICDCSAn Empirical Evaluation ofWork Stealing with Parallelism Feedback.Kunal Agrawal, Yuxiong He, Charles E. Leiserson
2006JSSPPProvably Efficient Two-Level Adaptive Scheduling.Yuxiong He, Wen-Jing Hsu, Charles E. Leiserson
2006PPoPPAdaptive scheduling with parallelism feedback.Kunal Agrawal, Yuxiong He, Wen-Jing Hsu, Charles E. Leiserson