| 2026 | HPCA | PASCAL: A Phase-Aware Scheduling Algorithm for Serving Reasoning-based Large Language Models. | Eunyeong Cho, Jehyeon Bang, Ranggi Hwang, Minsoo Rhu |
| 2026 | HPCA | The Cost of Dynamic Reasoning: Demystifying AI Agents and Test-Time Scaling from an AI Infrastructure Perspective. | Jiin Kim, Byeongjun Shin, Jinha Chung, Minsoo Rhu |
| 2026 | HPCA | PIM-Malloc: A Fast and Scalable Dynamic Memory Allocator for Processing-In-Memory (PIM) Architectures. | Dongjae Lee, Bongjoon Hyun, Youngjin Kwon, Minsoo Rhu |
| 2026 | Mobisys | Agent-X: Full Pipeline Acceleration of On-device AI Agents. | Jinha Chung, Byeongjun Shin, Jiin Kim, Minsoo Rhu |
| 2025 | ICCAD | Mamba-X: An End-to-End Vision Mamba Accelerator for Edge Computing Devices. | Dongho Yoon, Gungyu Lee, Jaewon Chang, Yunjae Lee, Dongjae Lee, Minsoo Rhu |
| 2025 | ISCA | Debunking the CUDA Myth Towards GPU-based AI Systems: Evaluation of the Performance and Programmability of Intel's Gaudi NPU for AI Model Serving. | Yunjae Lee, Juntaek Lim, Jehyeon Bang, Eunyeong Cho, Huijong Jeong, Taesu Kim, Hyungjun Kim, Joonhyung Lee, Jinseop Im, Ranggi Hwang, Se Jung Kwon, Dongsoo Lee, Minsoo Rhu |
| 2024 | ASPLOS | GPU-based Private Information Retrieval for On-Device Machine Learning Inference. | Maximilian Lam, Jeff Johnson, Wenjie Xiong, Kiwan Maeng, Udit Gupta, Yang Li, Liangzhen Lai, Ilias Leontiadis, Minsoo Rhu, Hsien-Hsin S. Lee, Vijay Janapa Reddi, Gu-Yeon Wei, David Brooks, G. Edward Suh |
| 2024 | ASPLOS | LazyDP: Co-Designing Algorithm-Software for Scalable Training of Differentially Private Recommendation Models. | Juntaek Lim, Youngeun Kwon, Ranggi Hwang, Kiwan Maeng, G. Edward Suh, Minsoo Rhu |
| 2024 | HPCA | Pathfinding Future PIM Architectures by Demystifying a Commercial PIM Technology. | Bongjoon Hyun, Taehun Kim, Dongjae Lee, Minsoo Rhu |
| 2024 | ISCA | ElasticRec: A Microservice-based Model Serving Architecture Enabling Elastic Resource Scaling for Recommendation Models. | Yujeong Choi, Jiin Kim, Minsoo Rhu |
| 2024 | ISCA | PreSto: An In-Storage Data Preprocessing System for Training Recommendation Models. | Yunjae Lee, Hyeseong Kim, Minsoo Rhu |
| 2024 | MICRO | vTrain: A Simulation Framework for Evaluating Cost-Effective and Compute-Optimal Large Language Model Training. | Jehyeon Bang, Yujeong Choi, Myeongwoo Kim, Yongdeok Kim, Minsoo Rhu |
| 2024 | MICRO | Uncovering Real GPU NoC Characteristics: Implications on Interconnect Architecture. | Zhixian Jin, Christopher Rocca, Jiho Kim, Hans Kasan, Minsoo Rhu, Ali Bakhoda, Tor M. Aamodt, John Kim |
| 2024 | MICRO | PIM-MMU: A Memory Management Unit for Accelerating Data Transfers in Commercial PIM Systems. | Dongjae Lee, Bongjoon Hyun, Taehun Kim, Minsoo Rhu |
| 2023 | HPCA | GROW: A Row-Stationary Sparse-Dense GEMM Accelerator for Memory-Efficient Graph Convolutional Neural Networks. | Ranggi Hwang, Minhoo Kang, Jiwon Lee, Dongyun Kam, Youngjoo Lee, Minsoo Rhu |
| 2022 | DAC | PARIS and ELSA: an elastic scheduling algorithm for reconfigurable multi-GPU inference servers. | Yunseong Kim, Yujeong Choi, Minsoo Rhu |
| 2022 | ISCA | BTS: an accelerator for bootstrappable fully homomorphic encryption. | Sangpyo Kim, Jongmin Kim, Michael Jaemin Kim, Wonkyung Jung, John Kim, Minsoo Rhu, Jung Ho Ahn |
| 2022 | ISCA | Training personalized recommendation systems from (GPU) scratch: look forward not backwards. | Youngeun Kwon, Minsoo Rhu |
| 2022 | ISCA | SmartSAGE: training large-scale graph neural networks using in-storage processing architectures. | Yunjae Lee, Jinha Chung, Minsoo Rhu |
| 2022 | MICRO | ARK: Fully Homomorphic Encryption Accelerator with Runtime Data Generation and Inter-Operation Key Reuse. | Jongmin Kim, Gwangho Lee, Sangpyo Kim, Gina Sohn, Minsoo Rhu, John Kim, Jung Ho Ahn |
| 2022 | MICRO | DiVa: An Accelerator for Differentially Private Machine Learning. | Beomsik Park, Ranggi Hwang, Dongho Yoon, Yoonhyuk Choi, Minsoo Rhu |
| 2021 | HPCA | Trident: A Hybrid Correlation-Collision GPU Cache Timing Attack for AES Key Recovery. | Jaeguk Ahn, Cheolgyu Jin, Jiho Kim, Minsoo Rhu, Yunsi Fei, David R. Kaeli, John Kim |
| 2021 | HPCA | Lazy Batching: An SLA-aware Batching System for Cloud Machine Learning Inference. | Yujeong Choi, Yunseong Kim, Minsoo Rhu |
| 2021 | HPCA | Tensor Casting: Co-Designing Algorithm-Architecture for Personalized Recommendation Training. | Youngeun Kwon, Yunjae Lee, Minsoo Rhu |
| 2021 | MICRO | TRiM: Enhancing Processor-Memory Interfaces with Scalable Tensor Reduction in Memory. | Jaehyun Park, Byeongho Kim, Sungmin Yun, Eojin Lee, Minsoo Rhu, Jung Ho Ahn |
| 2020 | ASPLOS | NeuMMU: Architectural Support for Efficient Address Translations in Neural Processing Units. | Bongjoon Hyun, Youngeun Kwon, Yujeong Choi, John Kim, Minsoo Rhu |
| 2020 | HPCA | PREMA: A Predictive Multi-Task Scheduling Algorithm For Preemptible Neural Processing Units. | Yujeong Choi, Minsoo Rhu |
| 2020 | ISCA | Centaur: A Chiplet-based, Hybrid Sparse-Dense Accelerator for Personalized Recommendations. | Ranggi Hwang, Taehun Kim, Youngeun Kwon, Minsoo Rhu |
| 2019 | MICRO | TensorDIMM: A Practical Near-Memory Processing Architecture for Embeddings and Tensor Operations in Deep Learning. | Youngeun Kwon, Yunjae Lee, Minsoo Rhu |
| 2018 | ASPDAC | Accelerator-centric deep learning systems for enhanced scalability, energy-efficiency, and programmability. | Minsoo Rhu |
| 2018 | HPCA | Compressing DMA Engine: Leveraging Activation Sparsity for Training Deep Neural Networks. | Minsoo Rhu, Mike O'Connor, Niladrish Chatterjee, Jeff Pool, Youngeun Kwon, Stephen W. Keckler |
| 2018 | MICRO | Beyond the Memory Wall: A Case for Memory-Centric HPC System for Deep Learning. | Youngeun Kwon, Minsoo Rhu |
| 2017 | HPCA | Architecting an Energy-Efficient DRAM System for GPUs. | Niladrish Chatterjee, Mike O'Connor, Donghyuk Lee, Daniel R. Johnson, Stephen W. Keckler, Minsoo Rhu, William J. Dally |
| 2017 | ISCA | SCNN: An Accelerator for Compressed-sparse Convolutional Neural Networks. | Angshuman Parashar, Minsoo Rhu, Anurag Mukkara, Antonio Puglielli, Rangharajan Venkatesan, Brucek Khailany, Joel S. Emer, Stephen W. Keckler, William J. Dally |
| 2017 | MICRO | GPUpd: a fast and scalable multi-GPU architecture using cooperative projection and distribution. | Youngsok Kim, Jae-Eon Jo, Hanhwi Jang, Minsoo Rhu, Hanjun Kim, Jangwoo Kim |
| 2016 | MICRO | vDNN: Virtualized deep neural networks for scalable, memory-efficient neural network design. | Minsoo Rhu, Natalia Gimelshein, Jason Clemons, Arslan Zulfiqar, Stephen W. Keckler |
| 2015 | HPCA | Priority-based cache allocation in throughput processors. | Dong Li, Minsoo Rhu, Daniel R. Johnson, Mike O'Connor, Mattan Erez, Doug Burger, Donald S. Fussell, Stephen W. Redder |
| 2015 | MICRO | CLEAN-ECC: high reliability ECC for adaptive granularity memory system. | Seong-Lyong Gong, Minsoo Rhu, Jungrae Kim, Jinsuk Chung, Mattan Erez |
| 2014 | ISLPED | GPUVolt: modeling and characterizing voltage noise in GPU architectures. | Jingwen Leng, Yazhou Zu, Minsoo Rhu, Meeta Sharma Gupta, Vijay Janapa Reddi |
| 2013 | HPCA | The dual-path execution model for efficient GPU control flow. | Minsoo Rhu, Mattan Erez |
| 2013 | ISCA | Maximizing SIMD resource utilization in GPGPUs with SIMD lane permutation. | Minsoo Rhu, Mattan Erez |
| 2013 | MICRO | A locality-aware memory hierarchy for energy-efficient GPU architectures. | Minsoo Rhu, Michael B. Sullivan, Jingwen Leng, Mattan Erez |
| 2012 | ISCA | CAPRI: Prediction of compaction-adequacy for handling control-divergence in GPGPU architectures. | Minsoo Rhu, Mattan Erez |
| 2009 | ICIP | Architecture design of a high-performance dual-symbol binary arithmetic coder for JPEG2000. | Minsoo Rhu, In-Cheol Park |
| 2009 | ICIP | Memory-less bit-plane coder architecture for JPEG2000 with concurrent column-stripe coding. | Minsoo Rhu, In-Cheol Park |