Skip to content

Wen-mei W. Hwu

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

118

Venues

27

Active years

1985–2020

Best venue rank

A*

Where they publish

Papers

118 indexed papers, newest first.

YearVenueTitleAuthors
2020DACEDD: Efficient Differentiable DNN Architecture and Implementation Co-search for Embedded AI Solutions.Yuhong Li, Cong Hao, Xiaofan Zhang, Xinheng Liu, Yao Chen, Jinjun Xiong, Wen-mei W. Hwu, Deming Chen
2020SCPetascale XCT: 3D image reconstruction with hierarchical communications on multi-GPU nodes.Mert Hidayetoglu, Tekin Bicer, Simon Garcia De Gonzalo, Bin Ren, Vincent De Andrade, Doga Grsoy, Raj Kettimuthu, Ian T. Foster, Wen-mei W. Hwu
2019ASPLOSPUMA: A Programmable Ultra-efficient Memristor-based Accelerator for Machine Learning Inference.Aayush Ankit, Izzat El Hajj, Sai Rahul Chalamalasetti, Geoffrey Ndu, Martin Foltin, R. Stanley Williams, Paolo Faraboschi, Wen-mei W. Hwu, John Paul Strachan, Kaushik Roy, Dejan S. Milojicic
2019SCMemXCT: memory-centric X-ray CT reconstruction with massive parallelization.Mert Hidayetoglu, Tekin Bier, Simon Garcia De Gonzalo, Bin Ren, Doga Grsoy, Rajkumar Kettimuthu, Ian T. Foster, Wen-mei W. Hwu
2017FPGAHardware Acceleration of the Pair-HMM Algorithm for DNA Variant Calling.Sitao Huang, Gowthami Jayashri Manikandan, Anand Ramachandran, Kyle Rupnow, Wen-mei W. Hwu, Deming Chen
2017ISPASSChai: Collaborative heterogeneous applications for integrated-architectures.Juan Gmez-Luna, Izzat El Hajj, Li-Wen Chang, Victor Garcia-Flores, Simon Garcia De Gonzalo, Thomas B. Jablin, Antonio J. Pea, Wen-mei W. Hwu
2016ASPLOSDySel: Lightweight Dynamic Selection for Kernel-based Data-parallel Programming Model.Li-Wen Chang, Hee-Seok Kim, Wen-mei W. Hwu
2016ASPLOSSpaceJMP: Programming with Multiple Virtual Address Spaces.Izzat El Hajj, Alexander Merritt, Gerd Zellweger, Dejan S. Milojicic, Reto Achermann, Paolo Faraboschi, Wen-mei W. Hwu, Timothy Roscoe, Karsten Schwan
2016FCCMAcceleration of the Pair-HMM Algorithm for DNA Variant Calling.Gowthami Jayashri Manikandan, Sitao Huang, Kyle Rupnow, Wen-mei W. Hwu, Deming Chen
2016HPDCEfficient and Scalable Workflows for Genomic Analyses.Subho S. Banerjee, Arjun P. Athreya, Liudmila S. Mainzer, C. Victor Jongeneel, Wen-mei W. Hwu, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
2016MICROEfficient kernel synthesis for performance portable programming.Li-Wen Chang, Izzat El Hajj, Christopher I. Rodrigues, Juan Gmez-Luna, Wen-mei W. Hwu
2016MICROKLAP: Kernel launch aggregation and promotion for optimizing dynamic parallelism.Izzat El Hajj, Juan Gmez-Luna, Cheng Li, Li-Wen Chang, Dejan S. Milojicic, Wen-mei W. Hwu
2016PPoPPA programming system for future proofing performance critical libraries.Li-Wen Chang, Izzat El Hajj, Hee-Seok Kim, Juan Gmez-Luna, Abdul Dakkak, Wen-mei W. Hwu
2015CGOLocality-centric thread scheduling for bulk-synchronous programming models on CPU architectures.Hee-Seok Kim, Izzat El Hajj, John A. Stratton, Steven S. Lumetta, Wen-mei W. Hwu
2015DATEFPGA accelerated DNA error correction.Anand Ramachandran, Yun Heo, Wen-mei W. Hwu, Jian Ma, Deming Chen
2015ICPPIn-Place Data Sliding Algorithms for Many-Core Architectures.Juan Gmez-Luna, Li-Wen Chang, I-Jui Sung, Wen-mei W. Hwu, Nicols Guil
2015ICSAutomatic Parallelization of Kernels in Shared-Memory Multi-GPU Nodes.Javier Cabezas, Llus Vilanova, Isaac Gelado, Thomas B. Jablin, Nacho Navarro, Wen-mei W. Hwu
2015PPoPPGPU-SM: shared memory multi-GPU programming.Javier Cabezas, Marc Jord, Isaac Gelado, Nacho Navarro, Wen-mei W. Hwu
2015UCCEnhancing the Usability and Utilization of Accelerated Architectures via Docker.Nicholas Haydel, Sandra Gesing, Ian J. Taylor, Gregory R. Madey, Abdul Dakkak, Simon Garcia De Gonzalo, Wen-mei W. Hwu
2014ISCAAdaptive Cache Bypass and Insertion for Many-core Accelerators.Xuhao Chen, Shengzhao Wu, Li-Wen Chang, Wei-Sheng Huang, Carl Pearson, Zhiying Wang, Wen-mei W. Hwu
2014MICROAdaptive Cache Management for Energy-Efficient GPU Computing.Xuhao Chen, Li-Wen Chang, Christopher I. Rodrigues, Jie Lv, Zhiying Wang, Wen-mei W. Hwu
2014PPoPPTriolet: a programming system that unifies algorithmic skeleton interfaces for high-performance cluster computing.Christopher I. Rodrigues, Thomas B. Jablin, Abdul Dakkak, Wen-mei W. Hwu
2014PPoPPIn-place transposition of rectangular matrices on accelerators.I-Jui Sung, Juan Gmez-Luna, Jos Mara Gonzlez-Linares, Nicols Guil, Wen-mei W. Hwu
2014SCSPEC ACCEL: A Standard Application Suite for Measuring Hardware Accelerator Performance.Guido Juckeland, William C. Brantley, Sunita Chandrasekaran, Barbara M. Chapman, Shuai Che, Mathew E. Colgrove, Huiyu Feng, Alexander Grund, Robert Henschel, Wen-mei W. Hwu, Huian Li, Matthias S. Mller, Wolfgang E. Nagel, Maxim Perminov, Pavel Shelepugin, Kevin Skadron, John A. Stratton, Alexey Titov, Ke Wang, G. Matthijs van Waveren, Brian Whitney, Sandra Wienke, Rengan Xu, Kalyan Kumaran
2013ASPLOSComparison based sorting for systems with multiple GPUs.Ivan Tanasic, Llus Vilanova, Marc Jord, Javier Cabezas, Isaac Gelado, Nacho Navarro, Wen-mei W. Hwu
2013DACThroughput-oriented kernel porting onto FPGAs.Alexandros Papakonstantinou, Deming Chen, Wen-mei W. Hwu, Jason Cong, Yun Liang
2012ICDMEfficient Pattern-Based Time Series Classification on GPU.Kai-Wei Chang, Biplab Deka, Wen-mei W. Hwu, Dan Roth
2012PPoPPEfficient performance evaluation of memory hierarchy for highly multithreaded graphics processors.Sara S. Baghsorkhi, Isaac Gelado, Matthieu Delahaye, Wen-mei W. Hwu
2012SCA scalable, numerically stable, high-performance tridiagonal solver using GPUs.Li-Wen Chang, John A. Stratton, Hee-Seok Kim, Wen-mei W. Hwu
2011FCCMMultilevel Granularity Parallelism Synthesis on FPGAs.Alexandros Papakonstantinou, Yun Liang, John A. Stratton, Karthik Gururaj, Deming Chen, Wen-mei W. Hwu, Jason Cong
2011ICASSPParallel implementation of Multi-dimensional Ensemble Empirical Mode Decomposition.Li-Wen Chang, Men-Tzung Lo, Nasser Anssari, Ke-Hsin Hsu, Norden E. Huang, Wen-mei W. Hwu
2011ICPPA Scalable Tridiagonal Solver for GPUs.Hee-Seok Kim, Shengzhao Wu, Li-Wen Chang, Wen-mei W. Hwu
2010ASPLOSAn asymmetric distributed shared memory model for heterogeneous parallel systems.Isaac Gelado, Javier Cabezas, Nacho Navarro, John E. Stone, Sanjay J. Patel, Wen-mei W. Hwu
2010CGOEfficient compilation of fine-grained SPMD-threaded programs for multicore CPUs.John A. Stratton, Vinod Grover, Jaydeep Marathe, Bastiaan Aarts, Mike Murphy, Ziang Hu, Wen-mei W. Hwu
2010DACAn effective GPU implementation of breadth-first search.Lijuan Luo, Martin D. F. Wong, Wen-mei W. Hwu
2010ISCAImplementing a GPU Programming Model on a Non-GPU Accelerator Architecture.Stephen M. Kofsky, Daniel R. Johnson, John A. Stratton, Wen-mei W. Hwu, Sanjay J. Patel, Steven S. Lumetta
2010PPoPPAn adaptive performance modeling tool for GPU architectures.Sara S. Baghsorkhi, Matthieu Delahaye, Sanjay J. Patel, William D. Gropp, Wen-mei W. Hwu
2009ASPLOSOptimization of tele-immersion codes.Albert Sidelnik, I-Jui Sung, Wanmin Wu, Mara Jess Garzarn, Wen-mei W. Hwu, Klara Nahrstedt, David A. Padua, Sanjay J. Patel
2009ASPLOSHigh performance computation and interactive display of molecular orbitals on GPUs and multi-core CPUs.John E. Stone, Jan Saam, David J. Hardy, Kirby L. Vandivort, Wen-mei W. Hwu, Klaus Schulten
2009CLUSTERGPU clusters for high-performance computing.Volodymyr V. Kindratenko, Jeremy Enos, Guochun Shi, Michael T. Showerman, Galen Wesley Arnold, John E. Stone, James C. Phillips, Wen-mei W. Hwu
2009ICSHigh-performance CUDA kernel execution on FPGAs.Alexandros Papakonstantinou, Karthik Gururaj, John A. Stratton, Deming Chen, Jason Cong, Wen-mei W. Hwu
2008CGOProgram optimization space pruning for a multithreaded gpu.Shane Ryoo, Christopher I. Rodrigues, Sam S. Stone, Sara S. Baghsorkhi, Sain-Zee Ueng, John A. Stratton, Wen-mei W. Hwu
2008ICSCUBA: an architecture for efficient CPU/co-processor data communication.Isaac Gelado, John H. Kelm, Shane Ryoo, Steven S. Lumetta, Nacho Navarro, Wen-mei W. Hwu
2008PPoPPOptimization principles and application performance evaluation of a multithreaded GPU using CUDA.Shane Ryoo, Christopher I. Rodrigues, Sara S. Baghsorkhi, Sam S. Stone, David Blair Kirk, Wen-mei W. Hwu
2007DACImplicitly Parallel Programming Models for Thousand-Core Microprocessors.Wen-mei W. Hwu, Shane Ryoo, Sain-Zee Ueng, John H. Kelm, Isaac Gelado, Sam S. Stone, Robert E. Kidd, Sara S. Baghsorkhi, Aqeel Mahesri, Stephanie C. Tsao, Nacho Navarro, Steven S. Lumetta, Matthew I. Frank, Sanjay J. Patel
2007DACCorezilla: Build and Tame the Multicore Beast?Lauren Sarno, Wen-mei W. Hwu, Craig Lund, Markus Levy, James R. Larus, James Reinders, Gordon Cameron, Chris Lennard, Takashi Yoshimori
2005HPCAThe Future of Computer Architecture Research: An Industrial Perspective.Wen-mei W. Hwu, Sanjay J. Patel
2005MICRO"Flea-flicker" Multipass Pipelining: An Alternative to the High-Power Out-of-Order Offense.Ronald D. Barnes, Shane Ryoo, Wen-mei W. Hwu
2004ISCAField-testing IMPACT EPIC research results in Itanium 2.John W. Sias, Sain-Zee Ueng, Geoff A. Kent, Ian M. Steiner, Erik M. Nystrom, Wen-mei W. Hwu
2004SASBottom-Up and Top-Down Context-Sensitive Summary-Based Pointer Analysis.Erik M. Nystrom, Hong-Seok Kim, Wen-mei W. Hwu
2003MICROBeating in-order stalls with "flea-flicker" two-pass pipelining.Ronald D. Barnes, Erik M. Nystrom, John W. Sias, Sanjay J. Patel, Nacho Navarro, Wen-mei W. Hwu
2002CASESCode coverage and input variability: effects on architecture and compiler research.Hillery C. Hunter, Wen-mei W. Hwu
2002MICROVacuum packing: extracting hardware-detected program phases for post-link optimization.Ronald D. Barnes, Erik M. Nystrom, Matthew C. Merten, Wen-mei W. Hwu
2001INFOCOMA Power Controlled Multiple Access Protocol for Wireless Packet Networks.Jeffrey P. Monks, Vaduvur Bharghavan, Wen-mei W. Hwu
2001LCNA Study of the Energy Saving and Capacity Improvement Potential of Power Control in Multi-Hop Wireless Networks.Jeffrey P. Monks, Jean-Pierre Ebert, Adam Wolisz, Wen-mei W. Hwu
2001MICROModulo schedule buffers.Matthew C. Merten, Wen-mei W. Hwu
2001MICROEnhancing loop buffering of media and telecommunications applications using low-overhead predication.John W. Sias, Hillery C. Hunter, Wen-mei W. Hwu
2000ASPLOSHardware Support for Dynamic Management of Compiler-Directed Computation Reuse.Daniel A. Connors, Hillery C. Hunter, Ben-Chung Cheng, Wen-mei W. Hwu
2000ISCAA hardware mechanism for dynamic extraction and relayout of program hot spots.Matthew C. Merten, Andrew R. Trick, Erik M. Nystrom, Ronald D. Barnes, Wen-mei W. Hwu
2000LCNTransmission Power Control for Multiple Access Wireless Packet Networks.Jeffrey P. Monks, Vaduvur Bharghavan, Wen-mei W. Hwu
2000MICROAccurate and efficient predicate analysis with binary decision diagrams.John W. Sias, Wen-mei W. Hwu, David I. August
2000PLDIModular interprocedural pointer analysis using access paths: design, implementation, and evaluation.Ben-Chung Cheng, Wen-mei W. Hwu
1999EuroParAn Architecture Framework for Introducing Predicated Execution into Embedded Microprocessors.Daniel A. Connors, Jean-Michel Puiatti, David I. August, Kevin M. Crozier, Wen-mei W. Hwu
1999ISCAThe Program Decision Logic Approach to Predicated Execution.David I. August, John W. Sias, Jean-Michel Puiatti, Scott A. Mahlke, Daniel A. Connors, Kevin M. Crozier, Wen-mei W. Hwu
1999ISCAA Hardware-Driven Profiling Scheme for Identifying Program Hot Spots to Support Runtime Optimization.Matthew C. Merten, Andrew R. Trick, Christopher N. George, John C. Gyllenhaal, Wen-mei W. Hwu
1999MICROCompiler-Directed Dynamic Computation Reuse: Rationale and Initial Results.Daniel A. Connors, Wen-mei W. Hwu
1999PLDIA New Framework for Debugging Globally Optimized Code.Le-Chun Wu, Rajiv Mirani, Harish Patil, Bruce Olsen, Wen-mei W. Hwu
1998ISCAIntegrated Predicated and Speculative Execution in the IMPACT EPIC Architecture.David I. August, Daniel A. Connors, Scott A. Mahlke, John W. Sias, Kevin M. Crozier, Ben-Chung Cheng, Patrick R. Eaton, Qudus B. Olaniran, Wen-mei W. Hwu
1998ISCAIMPACT: An Architectural Framework for Multiple-Instruction-Issue Processors.Pohua P. Chang, Scott A. Mahlke, William Y. Chen, Nancy J. Warter, Wen-mei W. Hwu
1998ISCARetrospective: IMPACT: An Architectural Framework for Multiple-Instruction Issue.Wen-mei W. Hwu
1998ISCARetrospective: HPSm, a High Performance Restricted Data Flow Architecture Having Minimal Functionality.Wen-mei W. Hwu, Yale N. Patt
1998ISCAHPSm, a High Performance Restricted Data Flow Architecture Having Minimal Functionality.Wen-mei W. Hwu, Yale N. Patt
1998MICROCompiler-Directed Early Load-Address Generation.Ben-Chung Cheng, Daniel A. Connors, Wen-mei W. Hwu
1997HPCAArchitectural Support for Compiler-Synthesized Dynamic Branch Prediction Strategies: Rationale and Initial Results.David I. August, Daniel A. Connors, John C. Gyllenhaal, Wen-mei W. Hwu
1997ISCARun-Time Adaptive Cache Hierarchy Management via Reference Analysis.Teresa L. Johnson, Wen-mei W. Hwu
1997MICROA Framework for Balancing Control Flow and Predication.David I. August, Wen-mei W. Hwu, Scott A. Mahlke
1997MICRORun-Time Spatial Locality Detection and Optimization.Teresa L. Johnson, Matthew C. Merten, Wen-mei W. Hwu
1996MICROSpeculative Hedge: Regulating Compile-time Speculation Against Profile Variations.Brian L. Deitrich, Wen-mei W. Hwu
1996MICROOptimization of Machine Descriptions for Efficient Use.John C. Gyllenhaal, Wen-mei W. Hwu, B. Ramakrishna Rau
1996MICROJava Bytecode to Native Code Translation: The Caffeine Prototype and Preliminary Results.Cheng-Hsueh A. Hsieh, John C. Gyllenhaal, Wen-mei W. Hwu
1996MICROModulo Scheduling of Loops in Control-intensive Non-numeric Programs.Daniel M. Lavery, Wen-mei W. Hwu
1995ISCAA Comparison of Full and Partial Predicated Execution Support for ILP Processors.Scott A. Mahlke, Richard E. Hank, James E. McCormick, David I. August, Wen-mei W. Hwu
1995MICRORegion-based compilation: an introduction and motivation.Richard E. Hank, Wen-mei W. Hwu, B. Ramakrishna Rau
1995MICROUnrolling-based optimizations for modulo scheduling.Daniel M. Lavery, Wen-mei W. Hwu
1994ASPLOSDynamic Memory Disambiguation Using the Memory Conflict Buffer.David M. Gallagher, William Y. Chen, Scott A. Mahlke, John C. Gyllenhaal, Wen-mei W. Hwu
1994ICPPAn Analytical Approach to Scheduling Code for Superscalar and VLIW Architectures.Shyh-Kwei Chen, W. Kent Fuchs, Wen-mei W. Hwu
1994MICROCharacterizing the impact of predicated execution on branch prediction.Scott A. Mahlke, Richard E. Hank, Roger A. Bringmann, John C. Gyllenhaal, David M. Gallagher, Wen-mei W. Hwu
1994MICROData relocation and prefetching for programs with large data sets.Yoji Yamada, John C. Gyllenhaal, Grant E. Haab, Wen-mei W. Hwu
1993ISCARegister Connection: A New Approach to Adding Registers into Instruction Set Architectures.Tokuzo Kiyohara, Scott A. Mahlke, William Y. Chen, Roger A. Bringmann, Richard E. Hank, Sadun Anik, Wen-mei W. Hwu
1993MICROSpeculative execution exception recovery using write-back suppression.Roger A. Bringmann, Scott A. Mahlke, Richard E. Hank, John C. Gyllenhaal, Wen-mei W. Hwu
1993MICROSuperblock formation using static program analysis.Richard E. Hank, Scott A. Mahlke, Roger A. Bringmann, John C. Gyllenhaal, Wen-mei W. Hwu
1993PLDIReverse If-Conversion.Nancy J. Warter, Scott A. Mahlke, Wen-mei W. Hwu, B. Ramakrishna Rau
1992ASPLOSSentinel Scheduling for VLIW and Superscalar Processors.Scott A. Mahlke, William Y. Chen, Wen-mei W. Hwu, B. Ramakrishna Rau, Michael S. Schlansker
1992ICPPExecuting Nested Parallel Loops on Shared-Memory Multiprocessors.Sadun Anik, Wen-mei W. Hwu
1992ICPPTolerating First Level Memory Access Latency in High-Performance Systems.William Y. Chen, Scott A. Mahlke, Wen-mei W. Hwu
1992ICSTolerating data access latency with register preloading.William Y. Chen, Scott A. Mahlke, Wen-mei W. Hwu, Tokuzo Kiyohara, Pohua P. Chang
1992SCCompiler Code Transformations for Superscalar-Based High Performance Systems.Scott A. Mahlke, William Y. Chen, John C. Gyllenhaal, Wen-mei W. Hwu
1992SIGMETRICSXprof: Profiling the Execution of X Window Programs.Aloke Gupta, Wen-mei W. Hwu
1992RSPSystematic prototyping of superscalar computer architectures.Thomas M. Conte, Wen-mei W. Hwu
1991ICPPThe Effect of Compiler Optimizations on Available Parallelism in Scalar Programs.Scott A. Mahlke, Nancy J. Warter, William Y. Chen, Pohua P. Chang, Wen-mei W. Hwu
1991ISCAIMPACT: An Architectural Framework for Multiple-Instruction-Issue Processors.Pohua P. Chang, Scott A. Mahlke, William Y. Chen, Nancy J. Warter, Wen-mei W. Hwu
1991MICROComparing Static and Dynamic Code Scheduling for Multiple-Instruction-Issue Processors.Pohua P. Chang, William Y. Chen, Scott A. Mahlke, Wen-mei W. Hwu
1991MICROData Access Microarchitectures for Superscalar Processors with Compiler-Assisted Data Prefetching.William Y. Chen, Scott A. Mahlke, Pohua P. Chang, Wen-mei W. Hwu
1989ISCAAchieving High Instruction Cache Performance with an Optimizing Compiler.Wen-mei W. Hwu, Pohua P. Chang
1989ISCAComparing Software and Hardware Schemes For Reducing the Cost of Branches.Wen-mei W. Hwu, Thomas M. Conte, Pohua P. Chang
1989ICSControl flow optimization for supercomputer scalar processing.Pohua P. Chang, Wen-mei W. Hwu
1989MICROForward semantic: a compiler-assisted instruction fetch method for heavily pipelined processors.P.-H. Chang, Wen-mei W. Hwu
1989PLDIInline Function Expansion for Compiling C Programs.Wen-mei W. Hwu, Pohua P. Chang
1989SIGMETRICSA Simulation Study of Simultaneous Vector Prefetch Performance in Multiprocessor Memory Subsystems (Extended Abstract).Wen-mei W. Hwu, Thomas M. Conte
1988ISCAExploiting Parallel Microprocessor Microarchitectures With a Compiler Code Generator.Wen-mei W. Hwu, Pohua P. Chang
1988MICROTrace selection for compiling large C application programs to microcode.Pohua P. Chang, Wen-mei W. Hwu
1987ISCACheckpoint Repair for Out-of-order Execution Machines.Wen-mei W. Hwu, Yale N. Patt
1987MICROExploiting horizontal and vertical concurrency via the HPSm microprocessor.Wen-mei W. Hwu, Yale N. Patt
1987MICROOn tuning the microarchitecture of an HPS implementation of the VAX.James E. Wilson, Stephen W. Melvin, Michael Shebanow, Wen-mei W. Hwu, Yale N. Patt
1986ISCAHPSm, a High Performance Restricted Data Flow Architecture Having Minimal Functionality.Wen-mei W. Hwu, Yale N. Patt
1986MICRORun-time generation of HPS microinstructions from a VAX instruction stream.Yale N. Patt, Stephen W. Melvin, Wen-mei W. Hwu, Michael Shebanow, Chein Chen
1985MICROHPS, a new microarchitecture: rationale and introduction.Yale N. Patt, Wen-mei W. Hwu, Michael Shebanow
1985MICROCritical issues regarding HPS, a high performance microarchitecture.Yale N. Patt, Stephen W. Melvin, Wen-mei W. Hwu, Michael Shebanow