Skip to content

Guang R. Gao

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

107

Venues

38

Active years

1983–2022

Best venue rank

A*

Where they publish

Papers

107 indexed papers, newest first.

YearVenueTitleAuthors
2022PPoPPThe SuperCodelet architecture.Jose Manuel Monsalve Diaz, Kevin Harms, Rafael A. Herrera Guaitero, Diego A. Roa Perdomo, Kalyan Kumaran, Guang R. Gao
2021SCE.T.: re-thinking self-attention for transformer models on GPUs.Shiyang Chen, Shaoyi Huang, Santosh Pandey, Bingbing Li, Guang R. Gao, Long Zheng, Caiwen Ding, Hang Liu
2020HiPCOn the Marriage of Asynchronous Many Task Runtimes and Big Data: A Glance.Joshua Suetterlein, Joseph B. Manzano, Andres Marquez, Guang R. Gao
2020JSSPPPDAWL: Profile-Based Iterative Dynamic Adaptive WorkLoad Balance on Heterogeneous Architectures.Tongsheng Geng, Marcos Amaris, Stphane Zuckerman, Alfredo Goldman, Guang R. Gao, Jean-Luc Gaudiot
2020SCCODIR: Towards an MLIR Codelet Model Dialect.Ryan Kabrick, Diego A. Roa Perdomo, Siddhisanket Raskar, Jose Manuel Monsalve Diaz, Dawson Fox, Guang R. Gao
2020SCDEMAC: A Modular Platform for HW-SW Co-Design.Diego A. Roa Perdomo, Ryan Kabrick, Jose Manuel Monsalve Diaz, Siddhisanket Raskar, Dawson Fox, Guang R. Gao
2019HPCCswFLOW: A Dataflow Deep Learning Framework on Sunway TaihuLight Supercomputer.Han Lin, Zeng Lin, Jose Monsalve Diaz, Mingfan Li, Hong An, Guang R. Gao
2017DACLeveraging Compiler Optimizations to Reduce Runtime Fault Recovery Overhead.Fateme S. Hosseini, Pouya Fotouhi, Chengmo Yang, Guang R. Gao
2017DATELeveraging access port positions to accelerate page table walk in DWM-based main memory.Hoda Aghaei Khouzani, Pouya Fotouhi, Chengmo Yang, Guang R. Gao
2017SCVerification of the Extended Roofline Model for Asynchronous Many Task Runtimes.Joshua Suetterlein, Joshua Landwehr, Andres Marquez, Joseph B. Manzano, Kevin J. Barker, Guang R. Gao
2016CLUSTERExtending the Roofline Model for Asynchronous Many-Task Runtimes.Joshua D. Suetterlein, Joshua Landwehr, Andrs Mrquez, Joseph B. Manzano, Guang R. Gao
2016NPCToward a Parallel Turing Machine Model.Peng Qu, Jin Yan, Guang R. Gao
2015CGOLocality aware concurrent start for stencil applications.Sunil Shrestha, Guang R. Gao, Joseph B. Manzano, Andrs Mrquez, John Feo
2015HPCCGregarious Data Re-structuring in a Many Core Architecture.Sunil Shrestha, Joseph B. Manzano, Andrs Mrquez, Stphane Zuckerman, Shuaiwen Song, Guang R. Gao
2015ICCSFreshBreeze: A Data Flow Approach for Meeting DDDAS Challenges.Xiaoming Li, Jack B. Dennis, Guang R. Gao, Willie Y.-P. Lim, Haitao Wei, Chao Yang, Robert S. Pavel
2015PPoPPDesign and evaluation of a novel dataflow based bigdata solution.Yao Wu, Long Zheng, Brian Heilig, Guang R. Gao
2014ICCSA Dataflow Programming Language and its Compiler for Streaming Systems.Haitao Wei, Stphane Zuckerman, Xiaoming Li, Guang R. Gao
2014ICPADSACDT: Architected Composite Data Types trading-in unfettered data access for improved execution.Andres Marquez, Joseph B. Manzano, Shuaiwen Leon Song, Benot Meister, Sunil Shrestha, Thomas St. John, Guang R. Gao
2013DSDThe TERAFLUX Project: Exploiting the DataFlow Paradigm in Next Generation Teradevices.Marco Solinas, Rosa M. Badia, Franois Bodin, Albert Cohen, Paraskevas Evripidou, Paolo Faraboschi, Bernhard Fechner, Guang R. Gao, Arne Garbade, Sylvain Girbal, Daniel Goodman, Behram Khan, Souad Koliai, Feng Li, Mikel Lujn, Laurent Morin, Avi Mendelson, Nacho Navarro, Antoniu Pop, Pedro Trancoso, Theo Ungerer, Mateo Valero, Sebastian Weis, Ian Watson, Stphane Zuckerman, Roberto Giorgi
2013EuroParToward a Self-aware System for Exascale Architectures.Aaron Myles Landwehr, Stphane Zuckerman, Guang R. Gao
2013EuroParAn Implementation of the Codelet Model.Joshua Suetterlein, Stphane Zuckerman, Guang R. Gao
2013HiPCA dynamic schema to increase performance in many-core architectures through percolation operations.Elkin Garcia, Daniel A. Orozco, Rishi Khan, Ioannis E. Venetis, Kelly Livingston, Guang R. Gao
2013TrustComAutomatic Locality Exploitation in the Codelet Model.Chen Chen, Yao Wu, Joshua Suetterlein, Long Zheng, Minyi Guo, Guang R. Gao
2012NPCDemystifying Performance Predictions of Distributed FFT3D Implementations.Daniel A. Orozco, Elkin Garcia, Robert S. Pavel, Orlando Ayala, Lian-Ping Wang, Guang R. Gao
2012PPoPPMassively parallel breadth first search using a tree-structured memory model.Tom St. John, Jack B. Dennis, Guang R. Gao
2011CLUSTERExploring Fine-Grained Task-Based Execution on Multi-GPU Systems.Long Chen, Oreste Villa, Guang R. Gao
2011EuroParHardware and Software Tradeoffs for Task Synchronization on Manycore Architectures.Yonghong Yan, Sanjay Chatterjee, Daniel A. Orozco, Elkin Garcia, Zoran Budimlic, Jun Shirako, Robert S. Pavel, Guang R. Gao, Vivek Sarkar
2011FPGADEEP: an iterative fpga-based many-core emulation system for chip verification and architecture research.Juergen Ributzka, Yuhei Hayashi, Fei Chen, Guang R. Gao
2011ICPADSSource Code Partitioning in Program Optimization.Murat Bolat, Kirk Kelsey, Xiaoming Li, Guang R. Gao
2011ICSThe elephant and the mice: the role of non-strict fine-grain synchronization for modern many-core architectures.Juergen Ributzka, Yuhei Hayashi, Joseph B. Manzano, Guang R. Gao
2010CGOMinimizing communication in rate-optimal software pipelining for stream programs.Haitao Wei, Junqing Yu, Huafei Yu, Guang R. Gao
2010EuroParA Study of a Software Cache Implementation of the OpenMP Memory Model for Multicore and Manycore Architectures.Chen Chen, Joseph B. Manzano, Ge Gan, Guang R. Gao, Vivek Sarkar
2010EuroParOptimized Dense Matrix Multiplication on a Many-Core Architecture.Elkin Garcia, Ioannis E. Venetis, Rishi Khan, Guang R. Gao
2009EuroParTile Percolation: An OpenMP Tile Aware Parallelization Technique for the Cyclops-64 Multicore Processor.Ge Gan, Xu Wang, Joseph B. Manzano, Guang R. Gao
2009ICPPMapping the FDTD Application to Many-Core Chip Architectures.Daniel A. Orozco, Guang R. Gao
2009IPCCCIterative layer-based raytracing on CUDA.Alejandro Segovia, Xiaoming Li, Guang R. Gao
2008PPoPPExperience on optimizing irregular computation for memory hierarchy in manycore architecture.Guangming Tan, Dongrui Fan, Junchao Zhang, Andrew Russo, Guang R. Gao
2007ISCASynchronization state buffer: supporting efficient fine-grain synchronization on many-core architectures.Weirong Zhu, Vugranam C. Sreedhar, Ziang Hu, Guang R. Gao
2007NPCOn Parallel Models of Computation.Guang R. Gao
2007PPoPPOptimized lock assignment and allocation: a method for exploiting concurrency among critical sections.Yuan Zhang, Vugranam C. Sreedhar, Weirong Zhu, Vivek Sarkar, Guang R. Gao
2007SCImplementation of the Smith-Waterman algorithm on a reconfigurable supercomputing platform.Peiheng Zhang, Guangming Tan, Guang R. Gao
2007SPAAA parallel dynamic programming algorithm on a multi-core architecture.Guangming Tan, Ninghui Sun, Guang R. Gao
2006EuroParMulti-dimensional Kernel Generation for Loop Nest Software Pipelining.Alban Douillet, Hongbo Rong, Guang R. Gao
2006EuroParOptimization of Dense Matrix Multiplication on IBM Cyclops-64: Challenges and Experiences.Ziang Hu, Juan del Cuvillo, Weirong Zhu, Guang R. Gao
2006ISPAExploring Financial Applications on Many-Core-on-a-Chip Architecture: A First Experiment.Weirong Zhu, Parimala Thulasiraman, Ruppa K. Thulasiram, Guang R. Gao
2005ICMLADiscriminating transmembrane proteins from signal peptides using SVM-Fisher approach.Li Liao, Robel Y. Kahsay, Guang R. Gao
2005ISLPEDAn energy efficient TLB design methodology.Dongrui Fan, Zhimin Tang, Hailin Huang, Guang R. Gao
2005NPCPerformance Modelling and Optimization of Memory Access on Cellular Computer Architecture Cyclops64.Yanwei Niu, Ziang Hu, Kenneth E. Barner, Guang R. Gao
2005PLDIRegister allocation for software pipelined multi-dimensional loops.Hongbo Rong, Alban Douillet, Guang R. Gao
2004CGOCode Generation for Single-Dimension Software Pipelining of Multi-Dimensional Loops.Hongbo Rong, Alban Douillet, Ramaswamy Govindarajan, Guang R. Gao
2004CGOSingle-Dimension Software Pipelining for Multi-Dimensional Loops.Hongbo Rong, Zhizhong Tang, Ramaswamy Govindarajan, Alban Douillet, Guang R. Gao
2004CLUSTERImplementing parallel conjugate gradient on the EARTH multithreaded architecture.Fei Chen, Kevin B. Theobald, Guang R. Gao
2004EuroParIf-Conversion in SSA Form.Arthur Stoutchinin, Guang R. Gao
2004ICTAIAn Improved Hidden Markov Model for Transmembrane Topology Prediction.Robel Y. Kahsay, Li Liao, Guang R. Gao
2003CLUSTERA Cluster-Based Solution for High Performance Hmmpfam Using EARTH Execution Model.Weirong Zhu, Yanwei Niu, Jizhu Lu, Chuan Shen, Guang R. Gao
2003IJCNNBiologically motivated computational models.Mitra Basu, Jon Timmis, Dipankar Dasgupta, Daniel D. Lee, Guang R. Gao, Kwabena A. Boahen
2003ICSInter-procedural stacked register allocation for itanium® like architecture.Liu Yang, Sun Chan, Guang R. Gao, Roy Ju, Guei-Yuan Lueh, Zhaoqing Zhang
2002CASESOn achieving balanced power consumption in software pipelined loops.Hongbo Yang, Guang R. Gao, Clement Leung
2002ICCDPower-Performance Trade-Offs for Energy-Efficient Architectures: A Quantitative Study.Hongbo Yang, Ramaswamy Govindarajan, Guang R. Gao, Kevin B. Theobald
2001CCSpeculative Prefetching of Induction Pointers.Artour Stoutchinin, Jos Nelson Amaral, Guang R. Gao, James C. Dehnert, Suneel Jain, Alban Douillet
2001EuroParTopic 08+13: Instruction-Level Parallelism and Computer Architecture.Eduard Ayguad, Fredrik Dahlgren, Christine Eisenbeis, Roger Espasa, Guang R. Gao, Henk L. Muller, Rizos Sakellariou, Andr Seznec
2001PSBA Multithreaded Parallel Implementation of a Dynamic Programming Algorithm for Sequence Comparison.Wellington Santos Martins, Juan del Cuvillo, F. J. Useche, Kevin B. Theobald, Guang R. Gao
2000EuroParDeveloping a Communication Intensive Application on the EARTH Multithreaded Architecture (Distinguished Paper).Kevin B. Theobald, Rishi Kumar, Gagan Agrawal, Gerd Heber, Ruppa K. Thulasiram, Guang R. Gao
2000ICSAutomatic compiler techniques for thread coarsening for multithreaded architectures.Gary M. Zoppetti, Gagan Agrawal, Lori L. Pollock, Jos Nelson Amaral, Xinan Tang, Guang R. Gao
2000PDPTARecursive and Iterative Multithreaded Algorithms for Pricing American Securities.Ruppa K. Thulasiram, Christopher T. Downing, Guang R. Gao
2000SCLanding CG on EARTH: A Case Study of Fine-Grained Multithreading on an Evolutionary Path.Kevin B. Theobald, Gagan Agrawal, Rishi Kumar, Gerd Heber, Guang R. Gao, Paul Stodghill, Keshav Pingali
2000SPAAMultithreaded algorithms for the fast Fourier transform.Parimala Thulasiraman, Kevin B. Theobald, Ashfaq A. Khokhar, Guang R. Gao
1999CCEfficient State-Diagram Construction Methods for Software Pipelining.Chihong Zhang, Ramaswamy Govindarajan, Sean Ryan, Guang R. Gao
1999HPCAMultithreaded Execution Architecture and Compilation.Dean M. Tullsen, Guang R. Gao
1998CCA New Fast Algorithm for Optimal Register Allocation in Modulo Scheduled Loops.Sylvain Lelait, Guang R. Gao, Christine Eisenbeis
1998HPCAPartial Sampling with Reverse State Reconstruction: A New Technique for Branch Predictor Performance Estimation.Darren Erik Vengroff, Guang R. Gao
1998ICPADSAutomatically Partitioning Threads Based on Remote Paths.Xinan Tang, Guang R. Gao
1998SPAAHow "Hard" is Thread Partitioning and How "Bad" is a List Scheduling Based Partitioning Algorithm?Xinan Tang, Guang R. Gao
1997ICCDElastic History Buffer: A Low-Cost Method to Improve Branch Prediction Accuracy.Maria-Dana Tarlescu, Kevin B. Theobald, Guang R. Gao
1997PPoPPExperiences with Non-numeric Applications on Multithreaded Architectures.Angela C. Sodan, Guang R. Gao, Olivier Maquelin, Jens-Uwe Schultz, Xinmin Tian
1997SPAAThread Partitioning and Scheduling Based on Cost Model.Xinan Tang, Jing Wang, Kevin B. Theobald, Guang R. Gao
1996CCPipelining-Dovetailing: A Transformation to Enhance Software Pipelining for Nested Loops.Jian Wang, Guang R. Gao
1996EuroParOptimal Software Pipelining Through Enumeration of Schedules.Erik R. Altman, Guang R. Gao
1996HiPCMultithreading implementation of a distributed shortest path algorithm on EARTH multiprocessor.Parimala Thulasiraman, Xinmin Tian, Guang R. Gao
1996HiPCQuantitive studies of data-locality sensitivity on the EARTH multithreaded architecture: preliminary results.Xinmin Tian, Shashank S. Nemawarkar, Guang R. Gao, Herbert H. J. Hum, Olivier Maquelin, Angela C. Sodan, Kevin B. Theobald
1996HPCACo-Scheduling Hardware and Software Pipelines.Ramaswamy Govindarajan, Erik R. Altman, Guang R. Gao
1996ISCAPolling Watchdog: Combining Polling and Interrupts for Efficient Message Handling.Olivier Maquelin, Guang R. Gao, Herbert H. J. Hum, Kevin B. Theobald, Xinmin Tian
1996MASCOTSMeasurement and Modeling of EARTH-MANNA Multithreaded Architecture.Shashank S. Nemawarkar, Guang R. Gao
1996PLDISoftware Pipelining Showdown: Optimal vs. Heuristic Methods in a Production Compiler.John C. Ruttenberg, Guang R. Gao, Woody Lichtenstein, Artour Stoutchinin
1996PLDIA New Framework for Exhaustive and Incremental Data Flow Analysis Using DJ Graphs.Vugranam C. Sreedhar, Guang R. Gao, Yong-Fong Lee
1995EuroParCosts and Benefits of Multithreading with Off-the-Shelf RISC Processors.Olivier Maquelin, Herbert H. J. Hum, Guang R. Gao
1995HPCAA Design Frame for Hybrid Access Caches.Kevin B. Theobald, Herbert H. J. Hum, Guang R. Gao
1995ICPPLocation Consistency: Stepping Beyond the Memory Coherence Barrier.Guang R. Gao, Vivek Sarkar
1995ICSThe Threaded Communication Library: Preliminary Experiences on a Multiprocessor with Dual-Processor Nodes.Nasser Elmasri, Herbert H. J. Hum, Guang R. Gao
1995MICROExploiting short-lived variables in superscalar processors.Luis A. Lozano, Guang R. Gao
1995PLDIScheduling and Mapping: Software Pipelining in the Presence of Structural Hazards.Erik R. Altman, Ramaswamy Govindarajan, Guang R. Gao
1995POPLA Linear Time Algorithm for Placing phi-nodes.Vugranam C. Sreedhar, Guang R. Gao
1994MICROMinimizing register requirements under resource-constrained rate-optimal software pipelining.Ramaswamy Govindarajan, Erik R. Altman, Guang R. Gao
1993ICSSpeculative Execution and Branch Prediction on Parallel Machines.Kevin B. Theobald, Guang R. Gao, Laurie J. Hendren
1993POPLA Novel Framework of Register Allocation for Software Pipelining.Qi Ning, Guang R. Gao
1992CCA Register Allocation Framework Based on Hierarchical Cyclic Interval Graphs.Laurie J. Hendren, Guang R. Gao, Erik R. Altman, Chandrika Mukerji
1992ICASSPWell-behaved dataflow programs for DSP computation.Guang R. Gao, R. Govindarajan, Prakash Panangaden
1992ICCIPerformance Evaluation of Latency Tolerant Architectures.Shashank S. Nemawarkar, Ramaswamy Govindarajan, Guang R. Gao, Vinod K. Agarwal
1992ICPPEfficient Interprocessor Synchronization/Communication on a Dataflow Multiprocessor Architecture.Jean-Marc Monti, Guang R. Gao
1992MICROOn the limits of program parallelism and its smoothability.Kevin B. Theobald, Guang R. Gao, Laurie J. Hendren
1991ICSOptimization of array accesses by collective loop transformations.Vivek Sarkar, Guang R. Gao
1991PLDIA Timed Petri-Net Model for Fine-Grain Loop Scheduling.Guang R. Gao, Yue-Bong Wong, Qi Ning
1991SCAn efficient parallel algorithm for all pairs examination.Kevin B. Theobald, Guang R. Gao
1990ICSTowards efficient fine-grain software pipelining.Guang R. Gao, Herbert H. J. Hum, Yue-Bong Wong
1988SCAn efficient pipelined dataflow processor architecture.Jack B. Dennis, Guang R. Gao
1986ICPPA Pipelined Solution Method of Tridiagonal Linear Equation Systems.Guang R. Gao
1983ICPPMaximum Pipelining of Array Operations on Static Data Flow Machine.Jack B. Dennis, Guang R. Gao