| 2020 | DATE | Tango: An Optimizing Compiler for Just-In-Time RTL Simulation. | Blaise-Pascal Tine, Sudhakar Yalamanchili, Hyesoon Kim |
| 2020 | HPCA | ALRESCHA: A Lightweight Reconfigurable Sparse-Computation Accelerator. | Bahar Asgari, Ramyad Hadidi, Tushar Krishna, Hyesoon Kim, Sudhakar Yalamanchili |
| 2019 | DAC | LODESTAR: Creating Locally-Dense CNNs for Efficient Inference on Systolic Arrays. | Bahar Asgari, Ramyad Hadidi, Hyesoon Kim, Sudhakar Yalamanchili |
| 2018 | ASPLOS | Slim NoC: A Low-Diameter On-Chip Network Topology for High Energy Efficiency and Scalability. | Maciej Besta, Syed Minhaj Hassan, Sudhakar Yalamanchili, Rachata Ausavarungnirun, Onur Mutlu, Torsten Hoefler |
| 2018 | ICCAD | A ferroelectric FET based power-efficient architecture for data-intensive computing. | Yun Long, Taesik Na, Prakshi Rastogi, Karthik Rao, Asif Islam Khan, Sudhakar Yalamanchili, Saibal Mukhopadhyay |
| 2017 | HPCA | Application-Specific Performance-Aware Energy Optimization on Android Mobile Devices. | Karthik Rao, Jun Wang, Sudhakar Yalamanchili, Yorai Wardi, Handong Ye |
| 2016 | HPCA | Amdahl's law for lifetime reliability scaling in heterogeneous multicore processors. | William J. Song, Saibal Mukhopadhyay, Sudhakar Yalamanchili |
| 2016 | ISCA | Neurocube: A Programmable Digital Neuromorphic Architecture with High-Density 3D Memory. | Duckhwan Kim, Jaeha Kung, Sek M. Chai, Sudhakar Yalamanchili, Saibal Mukhopadhyay |
| 2016 | ISCA | LaPerm: Locality Aware Scheduler for Dynamic Parallelism on GPUs. | Jin Wang, Norm Rubin, Albert Sidelnik, Sudhakar Yalamanchili |
| 2016 | PPoPP | General-purpose join algorithms for large graph triangle listing on heterogeneous systems. | Daniel Zinn, Haicheng Wu, Jin Wang, Molham Aref, Sudhakar Yalamanchili |
| 2016 | SC | Power-Constrained Performance Scheduling of Data Parallel Tasks. | Eric Anger, Jeremiah J. Wilke, Sudhakar Yalamanchili |
| 2015 | HiPC | Throughput Regulation in Shared Memory Multicore Processors. | X. Chen, H. Xiao, Yorai Wardi, Sudhakar Yalamanchili |
| 2015 | HPCC | Application Modeling for Scalable Simulation of Massively Parallel Systems. | Eric Anger, Damian Dechev, Gilbert Hendry, Jeremiah J. Wilke, Sudhakar Yalamanchili |
| 2015 | ISCA | Harmonia: balancing compute and memory power in high-performance GPUs. | Indrani Paul, Wei Huang, Manish Arora, Sudhakar Yalamanchili |
| 2015 | ISCA | Dynamic thread block launch: a lightweight execution mechanism to support irregular applications on GPUs. | Jin Wang, Norm Rubin, Albert Sidelnik, Sudhakar Yalamanchili |
| 2014 | ASPLOS | Efficient Instrumentation of GPGPU Applications Using Information Flow Analysis and Symbolic Execution. | Naila Farooqui, Karsten Schwan, Sudhakar Yalamanchili |
| 2014 | ASPLOS | ParallelJS: An Execution Framework for JavaScript on Heterogeneous Systems. | Jin Wang, Norman Rubin, Sudhakar Yalamanchili |
| 2014 | CGO | Red Fox: An Execution Environment for Relational Query Processing on GPUs. | Haicheng Wu, Gregory F. Diamos, Tim Sheard, Molham Aref, Sean Baxter, Michael Garland, Sudhakar Yalamanchili |
| 2014 | FCCM | Harmonica: An FPGA-Based Data Parallel Soft Core. | Chad D. Kersey, Sudhakar Yalamanchili, Hyojong Kim, Nimit Nigania, Hyesoon Kim |
| 2014 | ISPASS | Energy Introspector: A parallel, composable framework for integrated power-reliability-thermal modeling for multicore architectures. | William J. Song, Saibal Mukhopadhyay, Sudhakar Yalamanchili |
| 2014 | ISPASS | Manifold: A parallel simulation framework for multicore systems. | Jun Wang, Jesse G. Beu, Rishiraj A. Bheda, Tom Conte, Zhenjiang Dong, Chad D. Kersey, Mitchelle Rasquinha, George F. Riley, William J. Song, He Xiao, Peng Xu, Sudhakar Yalamanchili |
| 2014 | SC | Architecture-independent modeling of intra-node data movement. | Eric Anger, Sudhakar Yalamanchili, Scott Pakin, Patrick S. McCormick |
| 2014 | VLDB | Multipredicate Join Algorithms for Accelerating Relational Graph Processing on GPUs. | Haicheng Wu, Daniel Zinn, Molham Aref, Sudhakar Yalamanchili |
| 2013 | ASPLOS | Accelerating simulation of agent-based models on heterogeneous architectures. | Jin Wang, Norman Rubin, Haicheng Wu, Sudhakar Yalamanchili |
| 2013 | CLUSTER | Oncilla: A GAS runtime for efficient resource allocation and data movement in accelerated clusters. | Jeffrey S. Young, Se Hoon Shon, Sudhakar Yalamanchili, Alex Merritt, Karsten Schwan, Holger Frning |
| 2013 | ISCA | Cooperative boosting: needy versus greedy power management. | Indrani Paul, Srilatha Manne, Manish Arora, William Lloyd Bircher, Sudhakar Yalamanchili |
| 2013 | MASCOTS | A Study of the Effect of Partitioning on Parallel Simulation of Multicore Systems. | Zhenjiang Dong, Jun Wang, George F. Riley, Sudhakar Yalamanchili |
| 2013 | PADS | Optimizing parallel simulation of multicore systems using domain-specific knowledge. | Jun Wang, Zhenjiang Dong, Sudhakar Yalamanchili, George F. Riley |
| 2013 | PPoPP | Relational algorithms for multi-bulk-synchronous processors. | Gregory Frederick Diamos, Haicheng Wu, Jin Wang, Ashwin Sanjay Lele, Sudhakar Yalamanchili |
| 2013 | SC | Coordinated energy management in heterogeneous processors. | Indrani Paul, Vignesh T. Ravi, Srilatha Manne, Manish Arora, Sudhakar Yalamanchili |
| 2012 | CGO | Dynamic compilation of data-parallel kernels for vector processors. | Andrew Kerr, Gregory Frederick Diamos, Sudhakar Yalamanchili |
| 2012 | HiPC | Eiger: A framework for the automated synthesis of statistical performance models. | Andrew Kerr, Eric Anger, Gilbert Hendry, Sudhakar Yalamanchili |
| 2012 | HPCC | Commodity Converged Fabrics for Global Address Spaces in Accelerator Clouds. | Jeffrey S. Young, Sudhakar Yalamanchili |
| 2012 | IPCCC | Performance impact of virtual machine placement in a datacenter. | Indrani Paul, Sudhakar Yalamanchili, Lizy K. John |
| 2012 | ISPASS | Lynx: A dynamic instrumentation system for data-parallel applications on GPGPU architectures. | Naila Farooqui, Andrew Kerr, Greg Eisenhauer, Karsten Schwan, Sudhakar Yalamanchili |
| 2012 | MICRO | Kernel Weaver: Automatically Fusing Database Primitives for Efficient GPU Computation. | Haicheng Wu, Gregory Frederick Diamos, Srihari Cadambi, Sudhakar Yalamanchili |
| 2012 | SC | Designing Configurable, Modifiable and Reusable Components for Simulation of Multicore Systems. | Jun Wang, Jesse G. Beu, Sudhakar Yalamanchili, Tom Conte |
| 2012 | SC | Satisfying Data-Intensive Queries Using GPU Clusters. | Jeffrey S. Young, Haicheng Wu, Sudhakar Yalamanchili |
| 2011 | ASPLOS | A framework for dynamically instrumenting GPU compute applications within GPU Ocelot. | Naila Farooqui, Andrew Kerr, Gregory Frederick Diamos, Sudhakar Yalamanchili, Karsten Schwan |
| 2011 | MICRO | SIMD re-convergence at thread frontiers. | Gregory Frederick Diamos, Benjamin Ashbaugh, Subramaniam Maiyuran, Andrew Kerr, Haicheng Wu, Sudhakar Yalamanchili |
| 2010 | ASPLOS | Modeling GPU-CPU workloads and systems. | Andrew Kerr, Gregory F. Diamos, Sudhakar Yalamanchili |
| 2010 | ISLPED | An energy efficient cache design using spin torque transfer (STT) RAM. | Mitchelle Rasquinha, Dhruv Choudhary, Subho Chatterjee, Saibal Mukhopadhyay, Sudhakar Yalamanchili |
| 2009 | ICCAD | A methodology for robust, energy efficient design of Spin-Torque-Transfer RAM arrays at scaled technologies. | Subho Chatterjee, Mitchelle Rasquinha, Sudhakar Yalamanchili, Saibal Mukhopadhyay |
| 2008 | FCCM | ShareStreams-V: A Virtualized QoS Packet Scheduling Accelerator. | Kangtao Kendall Chuang, Sudhakar Yalamanchili, Ada Gavrilovska, Karsten Schwan |
| 2008 | HiPC | An Utilization Driven Framework for Energy Efficient Caches. | Subramanian Ramaswamy, Sudhakar Yalamanchili |
| 2008 | HPDC | Harmony: an execution model and runtime for heterogeneous many core systems. | Gregory F. Diamos, Sudhakar Yalamanchili |
| 2007 | ICCD | Improving cache efficiency via resizing + remapping. | Subramanian Ramaswamy, Sudhakar Yalamanchili |
| 2006 | ICCD | Customizable Fault Tolerant Caches for Embedded Processors. | Subramanian Ramaswamy, Sudhakar Yalamanchili |
| 2004 | FCCM | ShareStreams: A Scalable Architecture and Hardware Support for High-Speed QoS Packet Schedulers. | Raj Krishnamurthy, Sudhakar Yalamanchili, Karsten Schwan, Richard West |
| 2003 | PDPTA | A Hardware Approach to QoS Support in Cluster Environments: The Multimedia Router MMR. | Mara Blanca Caminero, Carmen Carrin, Francisco J. Quiles, Jos Duato, Sudhakar Yalamanchili |
| 2002 | HiPC | Algorithms for Switch-Scheduling in the Multimedia Router for LANs. | Indrani Paul, Sudhakar Yalamanchili, Jos Duato |
| 2002 | HiPC | The Customization Landscape for Embedded Systems. | Sudhakar Yalamanchili |
| 2002 | HOTI | Architecture and Hardware for Scheduling Gigabit Packet Streams. | Raj Krishnamurthy, Sudhakar Yalamanchili, Karsten Schwan, Richard West |
| 2002 | PDPTA | A Tunable Communications Library for Data Injection. | Craig D. Ulmer, Sudhakar Yalamanchili |
| 2000 | PDPTA | An Extensible Message Layer for High-Performance Clusters. | Craig D. Ulmer, Sudhakar Yalamanchili |
| 1999 | HCW | QUIC: A Quality of Service Network Interface Layer for Communication in NOWs. | Richard West, Raj Krishnamurthy, W. K. Norton, Karsten Schwan, Sudhakar Yalamanchili, Marcel-Catalin Rosu, V. Sarat |
| 1999 | HPCA | MMR: A High-Performance Multimedia Router - Architecture and Design Trade-Offs. | Jos Duato, Sudhakar Yalamanchili, Mara Blanca Caminero, Damon S. Love, Francisco J. Quiles |
| 1998 | RTAS | FARA - A Framework for Adaptive Resource Allocation in Complex Real-Time Systems. | Daniela Rosu, Karsten Schwan, Sudhakar Yalamanchili |
| 1997 | HPCA | Architectural Support for Reducing Communication Overhead in Multiprocessor Interconnection Networks. | Binh Vien Dao, Sudhakar Yalamanchili, Jos Duato |
| 1997 | ICCD | Power Constrained Design of Multiprocessor Interconnection Networks. | Chirag S. Patel, Sek M. Chai, Sudhakar Yalamanchili, David E. Schimmel |
| 1997 | RTSS | On adaptive resource allocation for complex real-time application. | Daniela Rosu, Karsten Schwan, Sudhakar Yalamanchili, Rakesh Jha |
| 1996 | DATE | Incorporating Multi-Chip Module Packaging Constraints into System Design. | Vivek Garg, Steve Lacy, David E. Schimmel, Darrell Stogner, Craig D. Ulmer, D. Scott Wills, Sudhakar Yalamanchili |
| 1996 | HiPC | Adaptive resource allocation for embedded parallel applications. | Rakesh Jha, Mustafa Muhammad, Sudhakar Yalamanchili, Karsten Schwan, Daniela Ivan-Rosu, Chris deCastro |
| 1996 | ICPP | A High Performance Router Architecture for Interconnection Networks. | Jos Duato, Pedro Lpez, Federico Silla, Sudhakar Yalamanchili |
| 1996 | WCAE | Using rapid prototyping in computer architecture design laboratories. | James O. Hamblen, Henry Owen, Sudhakar Yalamanchili, Binh Vien Dao |
| 1995 | ICPP | Software Based Fault-Tolerant Oblivious Routing in Pipelined Networks. | Young-Joo Suh, Binh Vien Dao, Jos Duato, Sudhakar Yalamanchili |
| 1995 | ISCA | Configurable Flow Control Mechanisms for Fault-Tolerant Routing. | Binh Vien Dao, Jos Duato, Sudhakar Yalamanchili |
| 1994 | ICPADS | Scouting: Fully Adaptive, Deadlock-Free Routing in Faulty Pipelined Networks. | Jos Duato, V. B. Dao, Patrick T. Gaughan, Sudhakar Yalamanchili |
| 1994 | ISCA | Ariadne - An Adaptive Router for Fault-Tolerant Multicomputers. | James D. Allen, Patrick T. Gaughan, David E. Schimmel, Sudhakar Yalamanchili |
| 1994 | MASCOTS | Simulation of Marked Graphs on SIMD Architectures Using Efficient Memory Management. | Hatem Sellami, James D. Allen, David E. Schimmel, Sudhakar Yalamanchili |
| 1984 | ICDE | Algebraic Properties of some Parallel Processor Interconnection Networks. | Sudhakar Yalamanchili, Jake K. Aggarwal |