| 2022 | PPoPP | The SuperCodelet architecture. | Jose Manuel Monsalve Diaz, Kevin Harms, Rafael A. Herrera Guaitero, Diego A. Roa Perdomo, Kalyan Kumaran, Guang R. Gao |
| 2021 | SC | E.T.: re-thinking self-attention for transformer models on GPUs. | Shiyang Chen, Shaoyi Huang, Santosh Pandey, Bingbing Li, Guang R. Gao, Long Zheng, Caiwen Ding, Hang Liu |
| 2020 | HiPC | On the Marriage of Asynchronous Many Task Runtimes and Big Data: A Glance. | Joshua Suetterlein, Joseph B. Manzano, Andres Marquez, Guang R. Gao |
| 2020 | JSSPP | PDAWL: Profile-Based Iterative Dynamic Adaptive WorkLoad Balance on Heterogeneous Architectures. | Tongsheng Geng, Marcos Amaris, Stphane Zuckerman, Alfredo Goldman, Guang R. Gao, Jean-Luc Gaudiot |
| 2020 | SC | CODIR: Towards an MLIR Codelet Model Dialect. | Ryan Kabrick, Diego A. Roa Perdomo, Siddhisanket Raskar, Jose Manuel Monsalve Diaz, Dawson Fox, Guang R. Gao |
| 2020 | SC | DEMAC: A Modular Platform for HW-SW Co-Design. | Diego A. Roa Perdomo, Ryan Kabrick, Jose Manuel Monsalve Diaz, Siddhisanket Raskar, Dawson Fox, Guang R. Gao |
| 2019 | HPCC | swFLOW: A Dataflow Deep Learning Framework on Sunway TaihuLight Supercomputer. | Han Lin, Zeng Lin, Jose Monsalve Diaz, Mingfan Li, Hong An, Guang R. Gao |
| 2017 | DAC | Leveraging Compiler Optimizations to Reduce Runtime Fault Recovery Overhead. | Fateme S. Hosseini, Pouya Fotouhi, Chengmo Yang, Guang R. Gao |
| 2017 | DATE | Leveraging access port positions to accelerate page table walk in DWM-based main memory. | Hoda Aghaei Khouzani, Pouya Fotouhi, Chengmo Yang, Guang R. Gao |
| 2017 | SC | Verification of the Extended Roofline Model for Asynchronous Many Task Runtimes. | Joshua Suetterlein, Joshua Landwehr, Andres Marquez, Joseph B. Manzano, Kevin J. Barker, Guang R. Gao |
| 2016 | CLUSTER | Extending the Roofline Model for Asynchronous Many-Task Runtimes. | Joshua D. Suetterlein, Joshua Landwehr, Andrs Mrquez, Joseph B. Manzano, Guang R. Gao |
| 2016 | NPC | Toward a Parallel Turing Machine Model. | Peng Qu, Jin Yan, Guang R. Gao |
| 2015 | CGO | Locality aware concurrent start for stencil applications. | Sunil Shrestha, Guang R. Gao, Joseph B. Manzano, Andrs Mrquez, John Feo |
| 2015 | HPCC | Gregarious Data Re-structuring in a Many Core Architecture. | Sunil Shrestha, Joseph B. Manzano, Andrs Mrquez, Stphane Zuckerman, Shuaiwen Song, Guang R. Gao |
| 2015 | ICCS | FreshBreeze: A Data Flow Approach for Meeting DDDAS Challenges. | Xiaoming Li, Jack B. Dennis, Guang R. Gao, Willie Y.-P. Lim, Haitao Wei, Chao Yang, Robert S. Pavel |
| 2015 | PPoPP | Design and evaluation of a novel dataflow based bigdata solution. | Yao Wu, Long Zheng, Brian Heilig, Guang R. Gao |
| 2014 | ICCS | A Dataflow Programming Language and its Compiler for Streaming Systems. | Haitao Wei, Stphane Zuckerman, Xiaoming Li, Guang R. Gao |
| 2014 | ICPADS | ACDT: Architected Composite Data Types trading-in unfettered data access for improved execution. | Andres Marquez, Joseph B. Manzano, Shuaiwen Leon Song, Benot Meister, Sunil Shrestha, Thomas St. John, Guang R. Gao |
| 2013 | DSD | The TERAFLUX Project: Exploiting the DataFlow Paradigm in Next Generation Teradevices. | Marco Solinas, Rosa M. Badia, Franois Bodin, Albert Cohen, Paraskevas Evripidou, Paolo Faraboschi, Bernhard Fechner, Guang R. Gao, Arne Garbade, Sylvain Girbal, Daniel Goodman, Behram Khan, Souad Koliai, Feng Li, Mikel Lujn, Laurent Morin, Avi Mendelson, Nacho Navarro, Antoniu Pop, Pedro Trancoso, Theo Ungerer, Mateo Valero, Sebastian Weis, Ian Watson, Stphane Zuckerman, Roberto Giorgi |
| 2013 | EuroPar | Toward a Self-aware System for Exascale Architectures. | Aaron Myles Landwehr, Stphane Zuckerman, Guang R. Gao |
| 2013 | EuroPar | An Implementation of the Codelet Model. | Joshua Suetterlein, Stphane Zuckerman, Guang R. Gao |
| 2013 | HiPC | A dynamic schema to increase performance in many-core architectures through percolation operations. | Elkin Garcia, Daniel A. Orozco, Rishi Khan, Ioannis E. Venetis, Kelly Livingston, Guang R. Gao |
| 2013 | TrustCom | Automatic Locality Exploitation in the Codelet Model. | Chen Chen, Yao Wu, Joshua Suetterlein, Long Zheng, Minyi Guo, Guang R. Gao |
| 2012 | NPC | Demystifying Performance Predictions of Distributed FFT3D Implementations. | Daniel A. Orozco, Elkin Garcia, Robert S. Pavel, Orlando Ayala, Lian-Ping Wang, Guang R. Gao |
| 2012 | PPoPP | Massively parallel breadth first search using a tree-structured memory model. | Tom St. John, Jack B. Dennis, Guang R. Gao |
| 2011 | CLUSTER | Exploring Fine-Grained Task-Based Execution on Multi-GPU Systems. | Long Chen, Oreste Villa, Guang R. Gao |
| 2011 | EuroPar | Hardware and Software Tradeoffs for Task Synchronization on Manycore Architectures. | Yonghong Yan, Sanjay Chatterjee, Daniel A. Orozco, Elkin Garcia, Zoran Budimlic, Jun Shirako, Robert S. Pavel, Guang R. Gao, Vivek Sarkar |
| 2011 | FPGA | DEEP: an iterative fpga-based many-core emulation system for chip verification and architecture research. | Juergen Ributzka, Yuhei Hayashi, Fei Chen, Guang R. Gao |
| 2011 | ICPADS | Source Code Partitioning in Program Optimization. | Murat Bolat, Kirk Kelsey, Xiaoming Li, Guang R. Gao |
| 2011 | ICS | The elephant and the mice: the role of non-strict fine-grain synchronization for modern many-core architectures. | Juergen Ributzka, Yuhei Hayashi, Joseph B. Manzano, Guang R. Gao |
| 2010 | CGO | Minimizing communication in rate-optimal software pipelining for stream programs. | Haitao Wei, Junqing Yu, Huafei Yu, Guang R. Gao |
| 2010 | EuroPar | A Study of a Software Cache Implementation of the OpenMP Memory Model for Multicore and Manycore Architectures. | Chen Chen, Joseph B. Manzano, Ge Gan, Guang R. Gao, Vivek Sarkar |
| 2010 | EuroPar | Optimized Dense Matrix Multiplication on a Many-Core Architecture. | Elkin Garcia, Ioannis E. Venetis, Rishi Khan, Guang R. Gao |
| 2009 | EuroPar | Tile Percolation: An OpenMP Tile Aware Parallelization Technique for the Cyclops-64 Multicore Processor. | Ge Gan, Xu Wang, Joseph B. Manzano, Guang R. Gao |
| 2009 | ICPP | Mapping the FDTD Application to Many-Core Chip Architectures. | Daniel A. Orozco, Guang R. Gao |
| 2009 | IPCCC | Iterative layer-based raytracing on CUDA. | Alejandro Segovia, Xiaoming Li, Guang R. Gao |
| 2008 | PPoPP | Experience on optimizing irregular computation for memory hierarchy in manycore architecture. | Guangming Tan, Dongrui Fan, Junchao Zhang, Andrew Russo, Guang R. Gao |
| 2007 | ISCA | Synchronization state buffer: supporting efficient fine-grain synchronization on many-core architectures. | Weirong Zhu, Vugranam C. Sreedhar, Ziang Hu, Guang R. Gao |
| 2007 | NPC | On Parallel Models of Computation. | Guang R. Gao |
| 2007 | PPoPP | Optimized lock assignment and allocation: a method for exploiting concurrency among critical sections. | Yuan Zhang, Vugranam C. Sreedhar, Weirong Zhu, Vivek Sarkar, Guang R. Gao |
| 2007 | SC | Implementation of the Smith-Waterman algorithm on a reconfigurable supercomputing platform. | Peiheng Zhang, Guangming Tan, Guang R. Gao |
| 2007 | SPAA | A parallel dynamic programming algorithm on a multi-core architecture. | Guangming Tan, Ninghui Sun, Guang R. Gao |
| 2006 | EuroPar | Multi-dimensional Kernel Generation for Loop Nest Software Pipelining. | Alban Douillet, Hongbo Rong, Guang R. Gao |
| 2006 | EuroPar | Optimization of Dense Matrix Multiplication on IBM Cyclops-64: Challenges and Experiences. | Ziang Hu, Juan del Cuvillo, Weirong Zhu, Guang R. Gao |
| 2006 | ISPA | Exploring Financial Applications on Many-Core-on-a-Chip Architecture: A First Experiment. | Weirong Zhu, Parimala Thulasiraman, Ruppa K. Thulasiram, Guang R. Gao |
| 2005 | ICMLA | Discriminating transmembrane proteins from signal peptides using SVM-Fisher approach. | Li Liao, Robel Y. Kahsay, Guang R. Gao |
| 2005 | ISLPED | An energy efficient TLB design methodology. | Dongrui Fan, Zhimin Tang, Hailin Huang, Guang R. Gao |
| 2005 | NPC | Performance Modelling and Optimization of Memory Access on Cellular Computer Architecture Cyclops64. | Yanwei Niu, Ziang Hu, Kenneth E. Barner, Guang R. Gao |
| 2005 | PLDI | Register allocation for software pipelined multi-dimensional loops. | Hongbo Rong, Alban Douillet, Guang R. Gao |
| 2004 | CGO | Code Generation for Single-Dimension Software Pipelining of Multi-Dimensional Loops. | Hongbo Rong, Alban Douillet, Ramaswamy Govindarajan, Guang R. Gao |
| 2004 | CGO | Single-Dimension Software Pipelining for Multi-Dimensional Loops. | Hongbo Rong, Zhizhong Tang, Ramaswamy Govindarajan, Alban Douillet, Guang R. Gao |
| 2004 | CLUSTER | Implementing parallel conjugate gradient on the EARTH multithreaded architecture. | Fei Chen, Kevin B. Theobald, Guang R. Gao |
| 2004 | EuroPar | If-Conversion in SSA Form. | Arthur Stoutchinin, Guang R. Gao |
| 2004 | ICTAI | An Improved Hidden Markov Model for Transmembrane Topology Prediction. | Robel Y. Kahsay, Li Liao, Guang R. Gao |
| 2003 | CLUSTER | A Cluster-Based Solution for High Performance Hmmpfam Using EARTH Execution Model. | Weirong Zhu, Yanwei Niu, Jizhu Lu, Chuan Shen, Guang R. Gao |
| 2003 | IJCNN | Biologically motivated computational models. | Mitra Basu, Jon Timmis, Dipankar Dasgupta, Daniel D. Lee, Guang R. Gao, Kwabena A. Boahen |
| 2003 | ICS | Inter-procedural stacked register allocation for itanium® like architecture. | Liu Yang, Sun Chan, Guang R. Gao, Roy Ju, Guei-Yuan Lueh, Zhaoqing Zhang |
| 2002 | CASES | On achieving balanced power consumption in software pipelined loops. | Hongbo Yang, Guang R. Gao, Clement Leung |
| 2002 | ICCD | Power-Performance Trade-Offs for Energy-Efficient Architectures: A Quantitative Study. | Hongbo Yang, Ramaswamy Govindarajan, Guang R. Gao, Kevin B. Theobald |
| 2001 | CC | Speculative Prefetching of Induction Pointers. | Artour Stoutchinin, Jos Nelson Amaral, Guang R. Gao, James C. Dehnert, Suneel Jain, Alban Douillet |
| 2001 | EuroPar | Topic 08+13: Instruction-Level Parallelism and Computer Architecture. | Eduard Ayguad, Fredrik Dahlgren, Christine Eisenbeis, Roger Espasa, Guang R. Gao, Henk L. Muller, Rizos Sakellariou, Andr Seznec |
| 2001 | PSB | A Multithreaded Parallel Implementation of a Dynamic Programming Algorithm for Sequence Comparison. | Wellington Santos Martins, Juan del Cuvillo, F. J. Useche, Kevin B. Theobald, Guang R. Gao |
| 2000 | EuroPar | Developing a Communication Intensive Application on the EARTH Multithreaded Architecture (Distinguished Paper). | Kevin B. Theobald, Rishi Kumar, Gagan Agrawal, Gerd Heber, Ruppa K. Thulasiram, Guang R. Gao |
| 2000 | ICS | Automatic compiler techniques for thread coarsening for multithreaded architectures. | Gary M. Zoppetti, Gagan Agrawal, Lori L. Pollock, Jos Nelson Amaral, Xinan Tang, Guang R. Gao |
| 2000 | PDPTA | Recursive and Iterative Multithreaded Algorithms for Pricing American Securities. | Ruppa K. Thulasiram, Christopher T. Downing, Guang R. Gao |
| 2000 | SC | Landing CG on EARTH: A Case Study of Fine-Grained Multithreading on an Evolutionary Path. | Kevin B. Theobald, Gagan Agrawal, Rishi Kumar, Gerd Heber, Guang R. Gao, Paul Stodghill, Keshav Pingali |
| 2000 | SPAA | Multithreaded algorithms for the fast Fourier transform. | Parimala Thulasiraman, Kevin B. Theobald, Ashfaq A. Khokhar, Guang R. Gao |
| 1999 | CC | Efficient State-Diagram Construction Methods for Software Pipelining. | Chihong Zhang, Ramaswamy Govindarajan, Sean Ryan, Guang R. Gao |
| 1999 | HPCA | Multithreaded Execution Architecture and Compilation. | Dean M. Tullsen, Guang R. Gao |
| 1998 | CC | A New Fast Algorithm for Optimal Register Allocation in Modulo Scheduled Loops. | Sylvain Lelait, Guang R. Gao, Christine Eisenbeis |
| 1998 | HPCA | Partial Sampling with Reverse State Reconstruction: A New Technique for Branch Predictor Performance Estimation. | Darren Erik Vengroff, Guang R. Gao |
| 1998 | ICPADS | Automatically Partitioning Threads Based on Remote Paths. | Xinan Tang, Guang R. Gao |
| 1998 | SPAA | How "Hard" is Thread Partitioning and How "Bad" is a List Scheduling Based Partitioning Algorithm? | Xinan Tang, Guang R. Gao |
| 1997 | ICCD | Elastic History Buffer: A Low-Cost Method to Improve Branch Prediction Accuracy. | Maria-Dana Tarlescu, Kevin B. Theobald, Guang R. Gao |
| 1997 | PPoPP | Experiences with Non-numeric Applications on Multithreaded Architectures. | Angela C. Sodan, Guang R. Gao, Olivier Maquelin, Jens-Uwe Schultz, Xinmin Tian |
| 1997 | SPAA | Thread Partitioning and Scheduling Based on Cost Model. | Xinan Tang, Jing Wang, Kevin B. Theobald, Guang R. Gao |
| 1996 | CC | Pipelining-Dovetailing: A Transformation to Enhance Software Pipelining for Nested Loops. | Jian Wang, Guang R. Gao |
| 1996 | EuroPar | Optimal Software Pipelining Through Enumeration of Schedules. | Erik R. Altman, Guang R. Gao |
| 1996 | HiPC | Multithreading implementation of a distributed shortest path algorithm on EARTH multiprocessor. | Parimala Thulasiraman, Xinmin Tian, Guang R. Gao |
| 1996 | HiPC | Quantitive studies of data-locality sensitivity on the EARTH multithreaded architecture: preliminary results. | Xinmin Tian, Shashank S. Nemawarkar, Guang R. Gao, Herbert H. J. Hum, Olivier Maquelin, Angela C. Sodan, Kevin B. Theobald |
| 1996 | HPCA | Co-Scheduling Hardware and Software Pipelines. | Ramaswamy Govindarajan, Erik R. Altman, Guang R. Gao |
| 1996 | ISCA | Polling Watchdog: Combining Polling and Interrupts for Efficient Message Handling. | Olivier Maquelin, Guang R. Gao, Herbert H. J. Hum, Kevin B. Theobald, Xinmin Tian |
| 1996 | MASCOTS | Measurement and Modeling of EARTH-MANNA Multithreaded Architecture. | Shashank S. Nemawarkar, Guang R. Gao |
| 1996 | PLDI | Software Pipelining Showdown: Optimal vs. Heuristic Methods in a Production Compiler. | John C. Ruttenberg, Guang R. Gao, Woody Lichtenstein, Artour Stoutchinin |
| 1996 | PLDI | A New Framework for Exhaustive and Incremental Data Flow Analysis Using DJ Graphs. | Vugranam C. Sreedhar, Guang R. Gao, Yong-Fong Lee |
| 1995 | EuroPar | Costs and Benefits of Multithreading with Off-the-Shelf RISC Processors. | Olivier Maquelin, Herbert H. J. Hum, Guang R. Gao |
| 1995 | HPCA | A Design Frame for Hybrid Access Caches. | Kevin B. Theobald, Herbert H. J. Hum, Guang R. Gao |
| 1995 | ICPP | Location Consistency: Stepping Beyond the Memory Coherence Barrier. | Guang R. Gao, Vivek Sarkar |
| 1995 | ICS | The Threaded Communication Library: Preliminary Experiences on a Multiprocessor with Dual-Processor Nodes. | Nasser Elmasri, Herbert H. J. Hum, Guang R. Gao |
| 1995 | MICRO | Exploiting short-lived variables in superscalar processors. | Luis A. Lozano, Guang R. Gao |
| 1995 | PLDI | Scheduling and Mapping: Software Pipelining in the Presence of Structural Hazards. | Erik R. Altman, Ramaswamy Govindarajan, Guang R. Gao |
| 1995 | POPL | A Linear Time Algorithm for Placing phi-nodes. | Vugranam C. Sreedhar, Guang R. Gao |
| 1994 | MICRO | Minimizing register requirements under resource-constrained rate-optimal software pipelining. | Ramaswamy Govindarajan, Erik R. Altman, Guang R. Gao |
| 1993 | ICS | Speculative Execution and Branch Prediction on Parallel Machines. | Kevin B. Theobald, Guang R. Gao, Laurie J. Hendren |
| 1993 | POPL | A Novel Framework of Register Allocation for Software Pipelining. | Qi Ning, Guang R. Gao |
| 1992 | CC | A Register Allocation Framework Based on Hierarchical Cyclic Interval Graphs. | Laurie J. Hendren, Guang R. Gao, Erik R. Altman, Chandrika Mukerji |
| 1992 | ICASSP | Well-behaved dataflow programs for DSP computation. | Guang R. Gao, R. Govindarajan, Prakash Panangaden |
| 1992 | ICCI | Performance Evaluation of Latency Tolerant Architectures. | Shashank S. Nemawarkar, Ramaswamy Govindarajan, Guang R. Gao, Vinod K. Agarwal |
| 1992 | ICPP | Efficient Interprocessor Synchronization/Communication on a Dataflow Multiprocessor Architecture. | Jean-Marc Monti, Guang R. Gao |
| 1992 | MICRO | On the limits of program parallelism and its smoothability. | Kevin B. Theobald, Guang R. Gao, Laurie J. Hendren |
| 1991 | ICS | Optimization of array accesses by collective loop transformations. | Vivek Sarkar, Guang R. Gao |
| 1991 | PLDI | A Timed Petri-Net Model for Fine-Grain Loop Scheduling. | Guang R. Gao, Yue-Bong Wong, Qi Ning |
| 1991 | SC | An efficient parallel algorithm for all pairs examination. | Kevin B. Theobald, Guang R. Gao |
| 1990 | ICS | Towards efficient fine-grain software pipelining. | Guang R. Gao, Herbert H. J. Hum, Yue-Bong Wong |
| 1988 | SC | An efficient pipelined dataflow processor architecture. | Jack B. Dennis, Guang R. Gao |
| 1986 | ICPP | A Pipelined Solution Method of Tridiagonal Linear Equation Systems. | Guang R. Gao |
| 1983 | ICPP | Maximum Pipelining of Array Operations on Static Data Flow Machine. | Jack B. Dennis, Guang R. Gao |