| 2011 | Universal rules guided design parameter selection for soft error resilient processors. | Lide Duan, Ying Zhang, Bin Li, Lu Peng |
| 2011 | Evaluation and optimization of multicore performance bottlenecks in supercomputing applications. | Jeffrey R. Diamond, Martin Burtscher, John D. McCalpin, Byoung-Do Kim, Stephen W. Keckler, James C. Browne |
| 2011 | Characterizing multi-threaded applications based on shared-resource contention. | Tanima Dey, Wei Wang, Jack W. Davidson, Mary Lou Soffa |
| 2011 | Storage I/O generation and replay for datacenter applications. | Christina Delimitrou, Sriram Sankar, Kushagra Vaid, Christos Kozyrakis |
| 2011 | Memory access pattern-aware DRAM performance model for multi-core systems. | Hyojin Choi, Jongbok Lee, Wonyong Sung |
| 2011 | Keynote II: Integrated modeling challenges in extreme-scale computing. | Pradip Bose |
| 2011 | Analyzing the impact of useless write-backs on the endurance and energy consumption of PCM main memory. | Santiago Bock, Bruce R. Childers, Rami G. Melhem, Daniel Moss, Youtao Zhang |
| 2011 | A comparative benchmarking of the FFT on Fermi and Evergreen GPUs. | Mohamed F. Ahmed, Omar Haridy |
| 2010 | Characterizing the design and performance of interactive java applications. | Dmitrijs Zaparanuks, Matthias Hauswirth |
| 2010 | Cache contention and application performance prediction for multi-core systems. | Chi Xu, Xi Chen, Robert P. Dick, Zhuoqing Morley Mao |
| 2010 | Demystifying GPU microarchitecture through microbenchmarking. | Henry Wong, Misel-Myrto Papadopoulou, Maryam Sadooghi-Alvandi, Andreas Moshovos |
| 2010 | Influences of SIMD architectures for scattered data interpolation algorithm. | Jean-Charles Tournier, Martin Naef |
| 2010 | Simulation environment for studying overlap of communication and computation. | Vladimir Subotic, Jess Labarta, Mateo Valero |
| 2010 | Dynamic program analysis of Microsoft Windows applications. | Alex Skaletsky, Tevi Devor, Nadav Chachmon, Robert S. Cohn, Kim M. Hazelwood, Vladimir Vladimirov, Moshe Bach |
| 2010 | Using special-purpose hardware to achieve a hundred-fold speedup in molecular dynamics simulations of proteins. | David Shaw |
| 2010 | The Hadoop distributed filesystem: Balancing portability and performance. | Jeffrey Shafer, Scott Rixner, Alan L. Cox |
| 2010 | Exploiting FPGAs for technology-aware system-level evaluation of multi-core architectures. | Simone Secchi, Paolo Meloni, Luigi Raffo |
| 2010 | Understanding transactional memory performance. | Donald E. Porter, Emmett Witchel |
| 2010 | An analysis of hard to predict branches. | Celal ztrk, Resit Sendag |
| 2010 | Hardware prediction of OS run-length for fine-grained resource customization. | David W. Nellans, Kshitij Sudan, Rajeev Balasubramonian, Erik Brunvand |
| 2010 | The big pileup. | Nick Mitchell |
| 2010 | Memphis: Finding and fixing NUMA-related performance problems on multi-core platforms. | Collin McCurdy, Jeffrey S. Vetter |
| 2010 | Modeling memory concurrency for multi-socket multi-core systems. | Anirban Mandal, Rob Fowler, Allan Porterfield |
| 2010 | PEBIL: Efficient static binary instrumentation for Linux. | Michael Laurenzano, Mustafa M. Tikir, Laura Carrington, Allan Snavely |
| 2010 | Performance-effective operation below Vcc-min. | Nikolas Ladas, Yiannakis Sazeides, Veerle Desmet |