| 2016 | Low-Cost Inter-Linked Subarrays (LISA): Enabling fast inter-subarray data movement in DRAM. | Kevin K. Chang, Prashant J. Nair, Donghyuk Lee, Saugata Ghose, Moinuddin K. Qureshi, Onur Mutlu |
| 2016 | Memristive Boltzmann machine: A hardware accelerator for combinatorial optimization and deep learning. | Mahdi Nazm Bojnordi, Engin Ipek |
| 2016 | Modeling cache performance beyond LRU. | Nathan Beckmann, Daniel Snchez |
| 2016 | Selective GPU caches to eliminate CPU-GPU HW cache coherence. | Neha Agarwal, David W. Nellans, Eiman Ebrahimi, Thomas F. Wenisch, John Danskin, Stephen W. Keckler |
| 2016 | Approximating warps with intra-warp operand value similarity. | Daniel Wong, Nam Sung Kim, Murali Annavaram |
| 2015 | Event-based scheduling for energy-efficient QoS (eQoS) in mobile Web applications. | Yuhao Zhu, Matthew Halpern, Vijay Janapa Reddi |
| 2015 | Studying the impact of multicore processor scaling on directory techniques via reuse distance analysis. | Minshu Zhao, Donald Yeung |
| 2015 | Overcoming the challenges of crossbar resistive memory architectures. | Cong Xu, Dimin Niu, Naveen Muralimanohar, Rajeev Balasubramonian, Tao Zhang, Shimeng Yu, Yuan Xie |
| 2015 | Quantifying sources of error in McPAT and potential impacts on architectural studies. | Sam Likun Xi, Hans M. Jacobson, Pradip Bose, Gu-Yeon Wei, David M. Brooks |
| 2015 | Coordinated static and dynamic cache bypassing for GPUs. | Xiaolong Xie, Yun Liang, Yu Wang, Guangyu Sun, Tao Wang |
| 2015 | GPGPU performance and power estimation using machine learning. | Gene Y. Wu, Joseph L. Greathouse, Alexander Lyashevsky, Nuwan Jayasena, Derek Chiou |
| 2015 | Overcoming far-end congestion in large-scale networks. | Jongmin Won, Gwangsun Kim, John Kim, Ted Jiang, Mike Parker, Steve Scott |
| 2015 | Alloy: Parallel-serial memory channel architecture for single-chip heterogeneous processor systems. | Hao Wang, Chang-Jae Park, Gyungsu Byun, Jung Ho Ahn, Nam Sung Kim |
| 2015 | XChange: A market-based approach to scalable dynamic multi-resource allocation in multicore architectures. | Xiaodong Wang, Jos F. Martnez |
| 2015 | Understanding GPU errors on large-scale HPC systems and the implications for system design and operation. | Devesh Tiwari, Saurabh Gupta, James H. Rogers, Don Maxwell, Paolo Rech, Sudharshan S. Vazhkudai, Daniel Oliveira, Dave Londo, Nathan DeBardeleben, Philippe Olivier Alexandre Navaux, Luigi Carro, Arthur S. Bland |
| 2015 | CiDRA: A cache-inspired DRAM resilience architecture. | Young Hoon Son, Sukhan Lee, Seongil O, Sanghyuk Kwon, Nam Sung Kim, Jung Ho Ahn |
| 2015 | Mascar: Speeding up GPU warps by reducing memory pitstops. | Ankit Sethia, Davoud Anoushe Jamshidi, Scott A. Mahlke |
| 2015 | Hierarchical private/shared classification: The key to simple and efficient coherence for clustered cache hierarchies. | Alberto Ros, Mahdad Davari, Stefanos Kaxiras |
| 2015 | Octopus-Man: QoS-driven task management for heterogeneous multicores in warehouse-scale computers. | Vinicius Petrucci, Michael A. Laurenzano, John Doherty, Yunqi Zhang, Daniel Moss, Jason Mars, Lingjia Tang |
| 2015 | BeBoP: A cost effective predictor infrastructure for superscalar value prediction. | Arthur Perais, Andr Seznec |
| 2015 | Exploiting compressed block size as an indicator of future reuse. | Gennady Pekhimenko, Tyler Huberty, Rui Cai, Onur Mutlu, Phillip B. Gibbons, Michael A. Kozuch, Todd C. Mowry |
| 2015 | Prediction-based superpage-friendly TLB designs. | Misel-Myrto Papadopoulou, Xin Tong, Andr Seznec, Andreas Moshovos |
| 2015 | iPatch: Intelligent fault patching to improve energy efficiency. | David J. Palframan, Nam Sung Kim, Mikko H. Lipasti |
| 2015 | Malware-aware processors: A framework for efficient online malware detection. | Meltem Ozsoy, Caleb Donovick, Iakov Gorelik, Nael B. Abu-Ghazaleh, Dmitry V. Ponomarev |
| 2015 | Scalable communication architecture for network-attached accelerators. | Sarah Neuwirth, Dirk Frey, Mondrian Nuessle, Ulrich Brning |