Scott A. Mahlke
Publication record assembled from the DBLP archive of ranked conferences.
Papers indexed
136
Venues
23
Active years
1991–2025
Best venue rank
A*
Where they publish
- A*MICRO36 papers
- A*ISCA16 papers
- ACGO15 papers
- Journal PublishedCASES13 papers
- A*ASPLOS12 papers
- A*HPCA11 papers
- A*PLDI7 papers
- A*DAC3 papers
- ADSN3 papers
- CICCD2 papers
- AICS2 papers
- ASC2 papers
- A*OSDI2 papers
- AISLPED2 papers
- BICPP2 papers
- BCC1 paper
- A*SIGMETRICS1 paper
- BISPASS1 paper
- AICCAD1 paper
- AMobisys1 paper
- CISCAS1 paper
- A*POPL1 paper
- MulticonferenceICASSP1 paper
Papers
136 indexed papers, newest first.
| Year | Venue | Title | Authors |
|---|---|---|---|
| 2025 | HPCA | Multi-Dimensional Vector ISA Extension for Mobile In-Cache Computing. | Alireza Khadem, Daichi Fujiki, Hilbert Chen, Yufeng Gu, Nishil Talati, Scott A. Mahlke, Reetuparna Das |
| 2025 | ISCA | DX100: Programmable Data Access Accelerator for Indirection. | Alireza Khadem, Kamalavasan Kamalakkannan, Zhenyan Zhu, Akash Poptani, Yufeng Gu, Jered Benjamin Dominguez-Trujillo, Nishil Talati, Daichi Fujiki, Scott A. Mahlke, Galen M. Shipman, Reetuparna Das |
| 2024 | ASPLOS | SlimSLAM: An Adaptive Runtime for Visual-Inertial Simultaneous Localization and Mapping. | Armand Behroozi, Yuxiang Chen, Vlad Fruchter, Lavanya Subramanian, Sriseshan Srikanth, Scott A. Mahlke |
| 2022 | CC | Loner: utilizing the CPU vector datapath to process scalar integer data. | Armand Behroozi, Sunghyun Park, Scott A. Mahlke |
| 2022 | CGO | SRTuner: Effective Compiler Optimization Customization by Exposing Synergistic Relations. | Sunghyun Park, Salar Latifi, Yongjun Park, Armand Behroozi, Byungsoo Jeon, Scott A. Mahlke |
| 2022 | ICCD | SoftFusion: A Low-Cost Approach to Enhance Reliability of Object Detection Applications. | Salar Latifi, Babak Zamirai, Scott A. Mahlke |
| 2022 | MICRO | Multi-Layer In-Memory Processing. | Daichi Fujiki, Alireza Khadem, Scott A. Mahlke, Reetuparna Das |
| 2021 | HPCA | Prodigy: Improving the Memory Latency of Data-Indirect Irregular Workloads Using Hardware-Software Co-Design. | Nishil Talati, Kyle May, Armand Behroozi, Yichen Yang, Kuba Kaszyk, Christos Vasiladiotis, Tarunesh Verma, Lu Li, Brandon Nguyen, Jiawen Sun, John Magnus Morton, Agreen Ahmadi, Todd M. Austin, Michael F. P. O'Boyle, Scott A. Mahlke, Trevor N. Mudge, Ronald G. Dreslinski |
| 2021 | SIGMETRICS | A Systematic Framework to Identify Violations of Scenario-dependent Driving Rules in Autonomous Vehicle Software. | Qingzhao Zhang, David Ke Hong, Ze Zhang, Qi Alfred Chen, Scott A. Mahlke, Z. Morley Mao |
| 2020 | CGO | Low-cost prediction-based fault protection strategy. | Sunghyun Park, Shikai Li, Ze Zhang, Scott A. Mahlke |
| 2020 | DAC | SIEVE: Speculative Inference on the Edge with Versatile Exportation. | Babak Zamirai, Salar Latifi, Pedram Zamirai, Scott A. Mahlke |
| 2020 | DSN | PolygraphMR: Enhancing the Reliability and Dependability of CNNs. | Salar Latifi, Babak Zamirai, Scott A. Mahlke |
| 2019 | ISCA | Duality cache for data parallel acceleration. | Daichi Fujiki, Scott A. Mahlke, Reetuparna Das |
| 2019 | ISPASS | Characterization of Unnecessary Computations in Web Applications. | Hossein Golestani, Scott A. Mahlke, Satish Narayanasamy |
| 2018 | ASPLOS | In-Memory Data Parallel Processor. | Daichi Fujiki, Scott A. Mahlke, Reetuparna Das |
| 2018 | DSN | Low Cost Transient Fault Protection Using Loop Output Prediction. | Sunghyun Park, Shikai Li, Scott A. Mahlke |
| 2018 | ICS | Sculptor: Flexible Approximation with Selective Dynamic Loop Perforation. | Shikai Li, Sunghyun Park, Scott A. Mahlke |
| 2017 | ASPLOS | Dynamic Resource Management for Efficient Utilization of Multitasking GPUs. | Jason Jong Kyu Park, Yongjun Park, Scott A. Mahlke |
| 2017 | ISCA | Scalpel: Customizing DNN Pruning to the Underlying Hardware Parallelism. | Jiecao Yu, Andrew Lukefahr, David J. Palframan, Ganesh S. Dasika, Reetuparna Das, Scott A. Mahlke |
| 2017 | MICRO | DeftNN: addressing bottlenecks for DNN execution on GPUs via synapse vector elimination and near-compute data fission. | Parker Hill, Animesh Jain, Mason Hill, Babak Zamirai, Chang-Hong Hsu, Michael A. Laurenzano, Scott A. Mahlke, Lingjia Tang, Jason Mars |
| 2017 | MICRO | Regless: just-in-time operand staging for GPUs. | John Kloosterman, Jonathan Beaumont, Davoud Anoushe Jamshidi, Jonathan Bailey, Trevor N. Mudge, Scott A. Mahlke |
| 2017 | MICRO | Mirage cores: the illusion of many out-of-order cores using in-order hardware. | Shruti Padmanabha, Andrew Lukefahr, Reetuparna Das, Scott A. Mahlke |
| 2016 | ICCAD | BugMD: automatic mismatch diagnosis for bug triaging. | Biruk Mammo, Milind Furia, Valeria Bertacco, Scott A. Mahlke, Daya Shanker Khudia |
| 2016 | MICRO | Concise loads and stores: The case for an asymmetric compute-memory architecture for approximation. | Animesh Jain, Parker Hill, Shih-Chieh Lin, Muneeb Khan, Md. Enamul Haque, Michael A. Laurenzano, Scott A. Mahlke, Lingjia Tang, Jason Mars |
| 2016 | PLDI | Input responsiveness: using canary inputs to dynamically steer approximation. | Michael A. Laurenzano, Parker Hill, Mehrzad Samadi, Scott A. Mahlke, Jason Mars, Lingjia Tang |
| 2015 | ASPLOS | Chimera: Collaborative Preemption for Multitasking on a Shared GPU. | Jason Jong Kyu Park, Yongjun Park, Scott A. Mahlke |
| 2015 | HPCA | Mascar: Speeding up GPU warps by reducing memory pitstops. | Ankit Sethia, Davoud Anoushe Jamshidi, Scott A. Mahlke |
| 2015 | ISCA | Accelerating asynchronous programs through event sneak peek. | Gaurav Chadha, Scott A. Mahlke, Satish Narayanasamy |
| 2015 | ISCA | Rumba: an online quality management system for approximate computing. | Daya Shanker Khudia, Babak Zamirai, Mehrzad Samadi, Scott A. Mahlke |
| 2015 | MICRO | WarpPool: sharing requests with inter-warp coalescing for throughput processors. | John Kloosterman, Jonathan Beaumont, Mick Wollman, Ankit Sethia, Ronald G. Dreslinski, Trevor N. Mudge, Scott A. Mahlke |
| 2015 | MICRO | DynaMOS: dynamic schedule migration for heterogeneous cores. | Shruti Padmanabha, Andrew Lukefahr, Reetuparna Das, Scott A. Mahlke |
| 2015 | Mobisys | Accelerating Mobile Applications through Flip-Flop Replication. | Mark S. Gordon, David Ke Hong, Peter M. Chen, Jason Flinn, Scott A. Mahlke, Zhuoqing Morley Mao |
| 2015 | SC | ELF: maximizing memory-level parallelism for GPUs with coordinated warp and fetch scheduling. | Jason Jong Kyu Park, Yongjun Park, Scott A. Mahlke |
| 2014 | ASPLOS | Paraprox: pattern-based approximation for data parallel applications. | Mehrzad Samadi, Davoud Anoushe Jamshidi, Janghaeng Lee, Scott A. Mahlke |
| 2014 | MICRO | Harnessing Soft Computations for Low-Budget Fault Tolerance. | Daya Shanker Khudia, Scott A. Mahlke |
| 2014 | MICRO | Equalizer: Dynamic Tuning of GPU Resources for Efficient Execution. | Ankit Sethia, Scott A. Mahlke |
| 2013 | CGO | Practical lock/unlock pairing for concurrent programs. | Hyoun Kyu Cho, Terence Kelly, Yin Wang, Stphane Lafortune, Hongwei Liao, Scott A. Mahlke |
| 2013 | CGO | Instant profiling: Instrumentation sampling for profiling datacenter applications. | Hyoun Kyu Cho, Tipp Moseley, Richard E. Hank, Derek Bruening, Scott A. Mahlke |
| 2013 | HPCA | Illusionist: Transforming lightweight cores into aggressive cores on demand. | Amin Ansari, Shuguang Feng, Shantanu Gupta, Josep Torrellas, Scott A. Mahlke |
| 2013 | ISCAS | Parallelization techniques for implementing trellis algorithms on graphics processors. | Qi Zheng, Yen-Po Chen, Ronald G. Dreslinski, Chaitali Chakrabarti, Achilleas Anastasopoulos, Scott A. Mahlke, Trevor N. Mudge |
| 2013 | MICRO | Trace based phase prediction for tightly-coupled heterogeneous cores. | Shruti Padmanabha, Andrew Lukefahr, Reetuparna Das, Scott A. Mahlke |
| 2013 | MICRO | SAGE: self-tuning approximation for graphics engines. | Mehrzad Samadi, Janghaeng Lee, Davoud Anoushe Jamshidi, Amir Hormati, Scott A. Mahlke |
| 2012 | ASPLOS | SIMD defragmenter: efficient ILP realization on data-parallel architectures. | Yongjun Park, Sangwon Seo, Hyunchul Park, Hyoun Kyu Cho, Scott A. Mahlke |
| 2012 | ASPLOS | Paragon: collaborative speculative loop execution on GPU and CPU. | Mehrzad Samadi, Amir Hormati, Janghaeng Lee, Scott A. Mahlke |
| 2012 | CASES | When less is more (LIMO): controlled parallelism forimproved efficiency. | Gaurav Chadha, Scott A. Mahlke, Satish Narayanasamy |
| 2012 | CGO | Automatic speculative DOALL for clusters. | Hanjun Kim, Nick P. Johnson, Jae W. Lee, Scott A. Mahlke, David I. August |
| 2012 | CGO | Runtime asynchronous fault tolerance via speculation. | Yun Zhang, Soumyadeep Ghosh, Jialu Huang, Jae W. Lee, Scott A. Mahlke, David I. August |
| 2012 | DAC | Process variation in near-threshold wide SIMD architectures. | Sangwon Seo, Ronald G. Dreslinski, Mark Woh, Yongjun Park, Chaitali Chakrabarti, Scott A. Mahlke, David T. Blaauw, Trevor N. Mudge |
| 2012 | MICRO | Dynamic acceleration of multithreaded program critical paths in near-threshold systems. | Hyoun Kyu Cho, Scott A. Mahlke |
| 2012 | MICRO | Composite Cores: Pushing Heterogeneity Into a Core. | Andrew Lukefahr, Shruti Padmanabha, Reetuparna Das, Faissal M. Sleiman, Ronald G. Dreslinski, Thomas F. Wenisch, Scott A. Mahlke |
| 2012 | MICRO | Libra: Tailoring SIMD Execution Using Heterogeneous Hardware and Dynamic Configurability. | Yongjun Park, Jason Jong Kyu Park, Hyunchul Park, Scott A. Mahlke |
| 2012 | OSDI | COMET: Code Offload by Migrating Execution Transparently. | Mark S. Gordon, Davoud Anoushe Jamshidi, Scott A. Mahlke, Zhuoqing Morley Mao, Xu Chen |
| 2012 | PLDI | Adaptive input-aware compilation for graphics engines. | Mehrzad Samadi, Amir Hormati, Mojtaba Mehrara, Janghaeng Lee, Scott A. Mahlke |
| 2011 | ASPLOS | Sponge: portable stream programming on graphics engines. | Amir Hormati, Mehrzad Samadi, Mark Woh, Trevor N. Mudge, Scott A. Mahlke |
| 2011 | CGO | Dynamically accelerating client-side web applications through decoupled execution. | Mojtaba Mehrara, Scott A. Mahlke |
| 2011 | HPCA | Archipelago: A polymorphic cache design for enabling robust near-threshold operation. | Amin Ansari, Shuguang Feng, Shantanu Gupta, Scott A. Mahlke |
| 2011 | HPCA | Dynamic parallelization of JavaScript applications using an ultra-lightweight speculation mechanism. | Mojtaba Mehrara, Po-Chun Hsu, Mehrzad Samadi, Scott A. Mahlke |
| 2011 | MICRO | Encore: low-cost, fine-grained transient fault recovery. | Shuguang Feng, Shantanu Gupta, Amin Ansari, Scott A. Mahlke, David I. August |
| 2011 | MICRO | Bundled execution of recurring traces for energy-efficient general purpose processing. | Shantanu Gupta, Shuguang Feng, Amin Ansari, Scott A. Mahlke, David I. August |
| 2010 | ASPLOS | Shoestring: probabilistic soft error reliability on the cheap. | Shuguang Feng, Shantanu Gupta, Amin Ansari, Scott A. Mahlke |
| 2010 | ASPLOS | MacroSS: macro-SIMDization of streaming applications. | Amir Hormati, Yoonseo Choi, Mark Woh, Manjunath Kudlur, Rodric M. Rabbah, Trevor N. Mudge, Scott A. Mahlke |
| 2010 | CASES | Mighty-morphing power-SIMD. | Ganesh S. Dasika, Mark Woh, Sangwon Seo, Nathan Clark, Trevor N. Mudge, Scott A. Mahlke |
| 2010 | CASES | Resource recycling: putting idle resources to work on a composable accelerator. | Yongjun Park, Hyunchul Park, Scott A. Mahlke, Sukjin Kim |
| 2010 | DSN | StageWeb: Interweaving pipeline stages into a wearout and variation tolerant CMP fabric. | Shantanu Gupta, Amin Ansari, Shuguang Feng, Scott A. Mahlke |
| 2010 | ISCA | Necromancer: enhancing system throughput by animating dead cores. | Amin Ansari, Shuguang Feng, Shantanu Gupta, Scott A. Mahlke |
| 2010 | ISLPED | Diet SODA: a power-efficient processor for digital cameras. | Sangwon Seo, Ronald G. Dreslinski, Mark Woh, Chaitali Chakrabarti, Scott A. Mahlke, Trevor N. Mudge |
| 2010 | MICRO | Erasing Core Boundaries for Robust and Configurable Performance. | Shantanu Gupta, Shuguang Feng, Amin Ansari, Scott A. Mahlke |
| 2009 | CASES | CGRA express: accelerating execution using dynamic operation fusion. | Yongjun Park, Hyunchul Park, Scott A. Mahlke |
| 2009 | CGO | Stream Compilation for Real-Time Embedded Multicore Systems. | Yoonseo Choi, Yuan Lin, Nathan Chong, Scott A. Mahlke, Trevor N. Mudge |
| 2009 | HPCA | Bridging the computation gap between programmable processors and hardwired accelerators. | Kevin Fan, Manjunath Kudlur, Ganesh S. Dasika, Scott A. Mahlke |
| 2009 | ICCD | Adaptive online testing for efficient hard fault detection. | Shantanu Gupta, Amin Ansari, Shuguang Feng, Scott A. Mahlke |
| 2009 | ISCA | AnySP: anytime anywhere anyway signal processing. | Mark Woh, Sangwon Seo, Scott A. Mahlke, Trevor N. Mudge, Chaitali Chakrabarti, Krisztin Flautner |
| 2009 | ISLPED | Enabling ultra low voltage system operation by tolerating on-chip cache failures. | Amin Ansari, Shuguang Feng, Shantanu Gupta, Scott A. Mahlke |
| 2009 | MICRO | ZerehCache: armoring cache architectures in high defect density technologies. | Amin Ansari, Shantanu Gupta, Shuguang Feng, Scott A. Mahlke |
| 2009 | MICRO | Polymorphic pipeline array: a flexible multicore accelerator with virtualized execution for mobile multimedia applications. | Hyunchul Park, Yongjun Park, Scott A. Mahlke |
| 2009 | PLDI | Parallelizing sequential applications on commodity hardware using a low-cost software transactional memory. | Mojtaba Mehrara, Jeff Hao, Po-Chun Hsu, Scott A. Mahlke |
| 2009 | POPL | The theory of deadlock avoidance via discrete control. | Yin Wang, Stphane Lafortune, Terence Kelly, Manjunath Kudlur, Scott A. Mahlke |
| 2008 | CASES | StageNetSlice: a reconfigurable microarchitecture building block for resilient CMP systems. | Shantanu Gupta, Shuguang Feng, Amin Ansari, Jason A. Blome, Scott A. Mahlke |
| 2008 | CASES | Optimus: efficient realization of streaming applications on FPGAs. | Amir Hormati, Manjunath Kudlur, Scott A. Mahlke, David F. Bacon, Rodric M. Rabbah |
| 2008 | CGO | Modulo scheduling for highly customized datapaths to increase hardware reusability. | Kevin Fan, Hyunchul Park, Manjunath Kudlur, Scott A. Mahlke |
| 2008 | DAC | DVFS in loop accelerators using BLADES. | Ganesh S. Dasika, Shidhartha Das, Kevin Fan, Scott A. Mahlke, David M. Bull |
| 2008 | HPCA | Uncovering hidden loop level parallelism in sequential applications. | Hongtao Zhong, Mojtaba Mehrara, Steven A. Lieberman, Scott A. Mahlke |
| 2008 | ICASSP | Analyzing the scalability of SIMD for the next generation software defined radio. | Mark Woh, Yuan Lin, Sangwon Seo, Trevor N. Mudge, Scott A. Mahlke |
| 2008 | ISCA | VEAL: Virtualized Execution Accelerator for Loops. | Nathan Clark, Amir Hormati, Scott A. Mahlke |
| 2008 | MICRO | The StageNet fabric for constructing resilient multicore systems. | Shantanu Gupta, Shuguang Feng, Amin Ansari, Jason A. Blome, Scott A. Mahlke |
| 2008 | MICRO | From SODA to scotch: The evolution of a wireless baseband processor. | Mark Woh, Yuan Lin, Sangwon Seo, Scott A. Mahlke, Trevor N. Mudge, Chaitali Chakrabarti, Richard Bruce, Danny Kershaw, Alastair Reid, Mladen Wilder, Krisztin Flautner |
| 2008 | OSDI | Gadara: Dynamic Deadlock Avoidance for Multithreaded Programs. | Yin Wang, Terence Kelly, Manjunath Kudlur, Stphane Lafortune, Scott A. Mahlke |
| 2008 | PLDI | Orchestrating the execution of stream programs on multicore platforms. | Manjunath Kudlur, Scott A. Mahlke |
| 2007 | CASES | Hierarchical coarse-grained stream compilation for software defined radio. | Yuan Lin, Manjunath Kudlur, Scott A. Mahlke, Trevor N. Mudge |
| 2007 | CGO | Exploiting Narrow Accelerators with Data-Centric Subgraph Mapping. | Amir Hormati, Nathan Clark, Scott A. Mahlke |
| 2007 | HPCA | Liquid SIMD: Abstracting SIMD Hardware using Lightweight Dynamic Mapping. | Nathan Clark, Amir Hormati, Sami Yehia, Scott A. Mahlke, Krisztin Flautner |
| 2007 | HPCA | Extending Multicore Architectures to Exploit Hybrid Parallelism in Single-thread Applications. | Hongtao Zhong, Steven A. Lieberman, Scott A. Mahlke |
| 2007 | MICRO | Self-calibrating Online Wearout Detection. | Jason A. Blome, Shuguang Feng, Shantanu Gupta, Scott A. Mahlke |
| 2007 | MICRO | Data Access Partitioning for Fine-grain Parallelism on Multicore Architectures. | Michael L. Chu, Rajiv A. Ravindran, Scott A. Mahlke |
| 2006 | CASES | Cost-efficient soft error protection for embedded microprocessors. | Jason A. Blome, Shantanu Gupta, Shuguang Feng, Scott A. Mahlke |
| 2006 | CASES | Scalable subgraph mapping for acyclic computation accelerators. | Nathan Clark, Amir Hormati, Scott A. Mahlke, Sami Yehia |
| 2006 | CASES | Modulo graph embedding: mapping applications onto coarse-grained reconfigurable architectures. | Hyunchul Park, Kevin Fan, Manjunath Kudlur, Scott A. Mahlke |
| 2006 | CGO | Compiler-directed Data Partitioning for Multicluster Processors. | Michael L. Chu, Scott A. Mahlke |
| 2006 | HPCA | BulletProof: a defect-tolerant CMP switch architecture. | Kypros Constantinides, Stephen Plaza, Jason A. Blome, Bin Zhang, Valeria Bertacco, Scott A. Mahlke, Todd M. Austin, Michael Orshansky |
| 2006 | ISCA | SODA: A Low-power Architecture For Software Radio. | Yuan Lin, Hyunseok Lee, Mark Woh, Yoav Harel, Scott A. Mahlke, Trevor N. Mudge, Chaitali Chakrabarti, Krisztin Flautner |
| 2005 | CASES | Exploring the design space of LUT-based transparent accelerators. | Sami Yehia, Nathan Clark, Scott A. Mahlke, Krisztin Flautner |
| 2005 | CGO | Compiler Managed Dynamic Instruction Placement in a Low-Power Code Cache. | Rajiv A. Ravindran, Pracheeti D. Nagarkar, Ganesh S. Dasika, Eric D. Marsman, Robert M. Senger, Scott A. Mahlke, Richard B. Brown |
| 2005 | ISCA | An Architecture Framework for Transparent Instruction Set Customization in Embedded Processors. | Nathan Clark, Jason A. Blome, Michael L. Chu, Scott A. Mahlke, Stuart Biles, Krisztin Flautner |
| 2005 | MICRO | Cost Sensitive Modulo Scheduling in a Loop Accelerator Synthesis System. | Kevin Fan, Manjunath Kudlur, Hyunchul Park, Scott A. Mahlke |
| 2004 | CGO | FLASH: Foresighted Latency-Aware Scheduling Heuristic for Processors with Customized Datapaths. | Manjunath Kudlur, Kevin Fan, Michael L. Chu, Rajiv A. Ravindran, Nathan Clark, Scott A. Mahlke |
| 2004 | CGO | Probabilistic Predicate-Aware Modulo Scheduling. | Mikhail Smelyanskiy, Scott A. Mahlke, Edward S. Davidson |
| 2004 | MICRO | Application-Specific Processing on a General-Purpose Core via Transparent Instruction Set Customization. | Nathan Clark, Manjunath Kudlur, Hyunchul Park, Scott A. Mahlke, Krisztin Flautner |
| 2003 | CASES | Architectural optimizations for low-power, real-time speech recognition. | Rajeev Krishna, Scott A. Mahlke, Todd M. Austin |
| 2003 | CASES | Increasing the number of effective registers in a low-power processor using a windowed register file. | Rajiv A. Ravindran, Robert M. Senger, Eric D. Marsman, Ganesh S. Dasika, Matthew R. Guthaus, Scott A. Mahlke, Richard B. Brown |
| 2003 | CGO | Predicate-Aware Scheduling: A Technique for Reducing Resource Constraints. | Mikhail Smelyanskiy, Scott A. Mahlke, Edward S. Davidson, Hsien-Hsin S. Lee |
| 2003 | MICRO | Processor Acceleration Through Automated Instruction Set Customization. | Nathan Clark, Hongtao Zhong, Scott A. Mahlke |
| 2003 | PLDI | Region-based hierarchical operation partitioning for multicluster processors. | Michael L. Chu, Kevin Fan, Scott A. Mahlke |
| 1999 | ISCA | The Program Decision Logic Approach to Predicated Execution. | David I. August, John W. Sias, Jean-Michel Puiatti, Scott A. Mahlke, Daniel A. Connors, Kevin M. Crozier, Wen-mei W. Hwu |
| 1999 | MICRO | Automatic and Efficient Evaluation of Memory Hierarchies for Embedded Systems. | Santosh G. Abraham, Scott A. Mahlke |
| 1999 | PLDI | Control CPR: A Branch Height Reduction Optimization for EPIC Architectures. | Michael S. Schlansker, Scott A. Mahlke, Richard Johnson |
| 1998 | ISCA | Integrated Predicated and Speculative Execution in the IMPACT EPIC Architecture. | David I. August, Daniel A. Connors, Scott A. Mahlke, John W. Sias, Kevin M. Crozier, Ben-Chung Cheng, Patrick R. Eaton, Qudus B. Olaniran, Wen-mei W. Hwu |
| 1998 | ISCA | IMPACT: An Architectural Framework for Multiple-Instruction-Issue Processors. | Pohua P. Chang, Scott A. Mahlke, William Y. Chen, Nancy J. Warter, Wen-mei W. Hwu |
| 1997 | MICRO | A Framework for Balancing Control Flow and Predication. | David I. August, Wen-mei W. Hwu, Scott A. Mahlke |
| 1996 | MICRO | Compiler Synthesized Dynamic Branch Prediction. | Scott A. Mahlke, Balas K. Natarajan |
| 1995 | ISCA | A Comparison of Full and Partial Predicated Execution Support for ILP Processors. | Scott A. Mahlke, Richard E. Hank, James E. McCormick, David I. August, Wen-mei W. Hwu |
| 1994 | ASPLOS | Dynamic Memory Disambiguation Using the Memory Conflict Buffer. | David M. Gallagher, William Y. Chen, Scott A. Mahlke, John C. Gyllenhaal, Wen-mei W. Hwu |
| 1994 | MICRO | Characterizing the impact of predicated execution on branch prediction. | Scott A. Mahlke, Richard E. Hank, Roger A. Bringmann, John C. Gyllenhaal, David M. Gallagher, Wen-mei W. Hwu |
| 1993 | ISCA | Register Connection: A New Approach to Adding Registers into Instruction Set Architectures. | Tokuzo Kiyohara, Scott A. Mahlke, William Y. Chen, Roger A. Bringmann, Richard E. Hank, Sadun Anik, Wen-mei W. Hwu |
| 1993 | MICRO | Speculative execution exception recovery using write-back suppression. | Roger A. Bringmann, Scott A. Mahlke, Richard E. Hank, John C. Gyllenhaal, Wen-mei W. Hwu |
| 1993 | MICRO | Superblock formation using static program analysis. | Richard E. Hank, Scott A. Mahlke, Roger A. Bringmann, John C. Gyllenhaal, Wen-mei W. Hwu |
| 1993 | PLDI | Reverse If-Conversion. | Nancy J. Warter, Scott A. Mahlke, Wen-mei W. Hwu, B. Ramakrishna Rau |
| 1992 | ASPLOS | Sentinel Scheduling for VLIW and Superscalar Processors. | Scott A. Mahlke, William Y. Chen, Wen-mei W. Hwu, B. Ramakrishna Rau, Michael S. Schlansker |
| 1992 | ICPP | Tolerating First Level Memory Access Latency in High-Performance Systems. | William Y. Chen, Scott A. Mahlke, Wen-mei W. Hwu |
| 1992 | ICS | Tolerating data access latency with register preloading. | William Y. Chen, Scott A. Mahlke, Wen-mei W. Hwu, Tokuzo Kiyohara, Pohua P. Chang |
| 1992 | MICRO | An efficient architecture for loop based data preloading. | William Y. Chen, Roger A. Bringmann, Scott A. Mahlke, Richard E. Hank, James E. Sicolo |
| 1992 | MICRO | Effective compiler support for predicated execution using the hyperblock. | Scott A. Mahlke, David C. Lin, William Y. Chen, Richard E. Hank, Roger A. Bringmann |
| 1992 | SC | Compiler Code Transformations for Superscalar-Based High Performance Systems. | Scott A. Mahlke, William Y. Chen, John C. Gyllenhaal, Wen-mei W. Hwu |
| 1991 | ICPP | The Effect of Compiler Optimizations on Available Parallelism in Scalar Programs. | Scott A. Mahlke, Nancy J. Warter, William Y. Chen, Pohua P. Chang, Wen-mei W. Hwu |
| 1991 | ISCA | IMPACT: An Architectural Framework for Multiple-Instruction-Issue Processors. | Pohua P. Chang, Scott A. Mahlke, William Y. Chen, Nancy J. Warter, Wen-mei W. Hwu |
| 1991 | MICRO | Comparing Static and Dynamic Code Scheduling for Multiple-Instruction-Issue Processors. | Pohua P. Chang, William Y. Chen, Scott A. Mahlke, Wen-mei W. Hwu |
| 1991 | MICRO | Data Access Microarchitectures for Superscalar Processors with Compiler-Assisted Data Prefetching. | William Y. Chen, Scott A. Mahlke, Pohua P. Chang, Wen-mei W. Hwu |