| 2010 | Modeling critical sections in Amdahl's law and its implications for multicore design. | Stijn Eyerman, Lieven Eeckhout |
| 2010 | Argia: exploiting packet latency slack in on-chip networks. | Reetuparna Das, Onur Mutlu, Thomas Moscibroda, Chita R. Das |
| 2010 | Moving the needle, computer architecture research in academe and industry. | William J. Dally |
| 2010 | LReplay: a pending period based deterministic replay scheme. | Yunji Chen, Weiwu Hu, Tianshi Chen, Ruiyang Wu |
| 2010 | A dynamically configurable coprocessor for convolutional neural networks. | Srimat T. Chakradhar, Murugan Sankaradass, Venkata Jakkula, Srihari Cadambi |
| 2010 | RETCON: transactional repair without replay. | Colin Blundell, Arun Raghavan, Milo M. K. Martin |
| 2010 | Evolution of thread-level parallelism in desktop applications. | Geoffrey Blake, Ronald G. Dreslinski, Trevor N. Mudge, Krisztin Flautner |
| 2010 | Predictive Power Management for Multi-core Processors. | William Lloyd Bircher, Lizy K. John |
| 2010 | Characteristics of Workloads Using the Pipeline Programming Model. | Christian Bienia, Kai Li |
| 2010 | Re-architecting DRAM memory systems with monolithically integrated silicon photonics. | Scott Beamer, Chen Sun, Yong-Jin Kwon, Ajay Joshi, Christopher Batten, Vladimir Stojanovic, Krste Asanovic |
| 2010 | Translation caching: skip, don't walk (the page table). | Thomas W. Barr, Alan L. Cox, Scott Rixner |
| 2010 | Energy-performance tradeoffs in processor architecture and circuit design: a marginal cost analysis. | Omid Azizi, Aqeel Mahesri, Benjamin C. Lee, Sanjay J. Patel, Mark Horowitz |
| 2010 | Necromancer: enhancing system throughput by animating dead cores. | Amin Ansari, Shuguang Feng, Shantanu Gupta, Scott A. Mahlke |
| 2010 | Achieving Power-Efficiency in Clusters without Distributed File System Complexity. | Hrishikesh Amur, Karsten Schwan |
| 2010 | IOMMU: Strategies for Mitigating the IOTLB Bottleneck. | Nadav Amit, Muli Ben-Yehuda, Ben-Ami Yassour |
| 2010 | Energy proportional datacenter networks. | Dennis Abts, Michael R. Marty, Philip M. Wells, Peter Klausler, Hong Liu |
| 2009 | A durable and energy efficient main memory using phase change memory technology. | Ping Zhou, Bo Zhao, Jun Yang, Youtao Zhang |
| 2009 | Decoupled DIMM: building high-bandwidth memory system using low-speed DRAM devices. | Hongzhong Zheng, Jiang Lin, Zhao Zhang, Zhichun Zhu |
| 2009 | A case for an interleaving constrained shared-memory multi-processor. | Jie Yu, Satish Narayanasamy |
| 2009 | Memory mapped ECC: low-cost error protection for last level caches. | Doe Hyun Yoon, Mattan Erez |
| 2009 | Ten ways to waste a parallel computer. | Katherine A. Yelick |
| 2009 | PIPP: promotion/insertion pseudo-partitioning of multi-core shared caches. | Yuejian Xie, Gabriel H. Loh |
| 2009 | Hybrid cache architecture with disparate memory technologies. | Xiaoxia Wu, Jian Li, Lixin Zhang, Evan Speight, Ramakrishnan Rajamony, Yuan Xie |
| 2009 | AnySP: anytime anywhere anyway signal processing. | Mark Woh, Sangwon Seo, Scott A. Mahlke, Trevor N. Mudge, Chaitali Chakrabarti, Krisztin Flautner |
| 2009 | A fault tolerant, area efficient architecture for Shor's factoring algorithm. | Mark Whitney, Nemanja Isailovic, Yatish Patel, John Kubiatowicz |