| 2017 | Sharing the instruction cache among lean cores on an asymmetric CMP for HPC applications. | Ugljesa Milic, Alejandro Rico, Paul M. Carpenter, Alex Ramrez |
| 2017 | Exploring GPU performance, power and energy-efficiency bounds with Cache-aware Roofline Modeling. | Andre Lopes, Frederico Pratas, Leonel Sousa, Aleksandar Ilic |
| 2017 | StressRight: Finding the right stress for accurate in-development system evaluation. | Jaewon Lee, Hanhwi Jang, Jae-Eon Jo, Gyu-hyeon Lee, Jangwoo Kim |
| 2017 | Fast IPC estimation for performance projections using proxy suites and decision trees. | Kanishka Lahiri, Subhash Kunnoth |
| 2017 | OpenSMART: Single-cycle multi-hop NoC generator in BSV and Chisel. | Hyoukjun Kwon, Tushar Krishna |
| 2017 | DARTS: Performance-counter driven sampling using binary translators. | Rajesh Kumar, Suchita Pati, Kanishka Lahiri |
| 2017 | HW/SW co-designed processors: Challenges, design choices and a simulation infrastructure for evaluation. | Rakesh Kumar, Jos Cano, Aleksandar Brankovic, Demos Pavlou, Kyriakos Stavrou, Enric Gibert, Alejandro Martnez, Antonio Gonzalez |
| 2017 | Performance analysis of CNN frameworks for GPUs. | Heehoon Kim, Hyoungwook Nam, Wookeun Jung, Jaejin Lee |
| 2017 | Machine learning for performance and power modeling/prediction. | Lizy Kurian John |
| 2017 | Treelogy: A benchmark suite for tree traversals. | Nikhil Hegde, Jianqiao Liu, Kirshanthan Sundararajah, Milind Kulkarni |
| 2017 | SASSIFI: An architecture-level fault injection tool for GPU application resilience evaluation. | Siva Kumar Sastry Hari, Timothy Tsai, Mark Stephenson, Stephen W. Keckler, Joel S. Emer |
| 2017 | Multi2Sim Kepler: A detailed architectural GPU simulator. | Xun Gong, Rafael Ubal, David R. Kaeli |
| 2017 | Chai: Collaborative heterogeneous applications for integrated-architectures. | Juan Gmez-Luna, Izzat El Hajj, Li-Wen Chang, Victor Garcia-Flores, Simon Garcia De Gonzalo, Thomas B. Jablin, Antonio J. Pea, Wen-mei W. Hwu |
| 2017 | Crossing the architectural barrier: Evaluating representative regions of parallel HPC applications. | Alexandra Ferreron, Radhika Jagtap, Sascha Bischoff, Roxana Rusitoru |
| 2017 | Predicting memory page stability and its application to memory deduplication and live migration. | Karim Elghamrawy, Diana Franklin, Frederic T. Chong |
| 2017 | Evaluating and mitigating bandwidth bottlenecks across the memory hierarchy in GPUs. | Saumay Dublish, Vijay Nagarajan, Nigel P. Topham |
| 2017 | GaaS workload characterization under NUMA architecture for virtualized GPU. | Huixiang Chen, Meng Wang, Yang Hu, Mingcong Song, Tao Li |
| 2017 | A taxonomy of out-of-order instruction commit. | Mehdi Alipour, Trevor E. Carlson, Stefanos Kaxiras |
| 2017 | dist-gem5: Distributed simulation of computer clusters. | Mohammad Alian, Umur Darbaz, Gbor Dzsa, Stephan Diestelhorst, Daehoon Kim, Nam Sung Kim |
| 2016 | Observations and opportunities in architecting shared virtual memory for heterogeneous systems. | Jn Vesel, Arkaprava Basu, Mark Oskin, Gabriel H. Loh, Abhishek Bhattacharjee |
| 2016 | GUFI: A framework for GPUs reliability assessment. | Sotiris Tselonis, Dimitris Gizopoulos |
| 2016 | RTHpower: Accurate fine-grained power models for predicting race-to-halt effect on ultra-low power embedded systems. | Vi Ngoc-Nha Tran, Brendan Barry, Phuong Hoai Ha |
| 2016 | EmerGPU: Understanding and mitigating resonance-induced voltage noise in GPU architectures. | Renji Thomas, Naser Sedaghati, Radu Teodorescu |
| 2016 | Performance analysis of a hardware accelerator of dependence management for task-based dataflow programming models. | Xubin Tan, Jaume Bosch, Daniel Jimnez-Gonzlez, Carlos lvarez-Martnez, Eduard Ayguad, Mateo Valero |
| 2016 | Analysis of PARSEC workload scalability. | Gabriel Southern, Jose Renau |