| 2010 | A study of hardware assisted IP over InfiniBand and its impact on enterprise data center performance. | Ryan E. Grant, Pavan Balaji, Ahmad Afsahi |
| 2010 | Synthesizing memory-level parallelism aware miniature clones for SPEC CPU2006 and ImplantBench workloads. | Karthik Ganesan, Jungho Jo, Lizy K. John |
| 2010 | StatStack: Efficient modeling of LRU caches. | David Eklov, Erik Hagersten |
| 2010 | Program behavior characterization in large memory systems. | Parijat Dube, Michael Tsao, Dan E. Poff, Li Zhang, Alan Bivens |
| 2010 | Incorporating Instruction-Based Sampling into AMD CodeAnalyst. | Paul J. Drongowski, Lei Yu, Frank Swehosky, Suravee Suthikulpanit, Robert Richter |
| 2010 | ArchExplorer.org: A methodology for facilitating a fair Comparison of research ideas. | Veerle Desmet, Sylvain Girbal, Olivier Temam |
| 2010 | Scalability comparison of commodity operating systems on multi-cores. | Yan Cui, Yu Chen, Yuanchun Shi, Qingbo Wu |
| 2010 | Scaling OLTP applications on commodity multi-core platforms. | Yan Cui, Yu Chen, Yuanchun Shi |
| 2010 | Weak execution ordering - exploiting iterative methods on many-core GPUs. | Jianmin Chen, Zhuo Huang, Feiqi Su, Jih-Kwon Peir, Jeff Ho, Lu Peng |
| 2010 | Runahead execution vs. conventional data prefetching in the IBM POWER6 microprocessor. | Harold W. Cain, Priya Nagpurkar |
| 2010 | Visualizing complex dynamics in many-core accelerator architectures. | Aaron Ariel, Wilson W. L. Fung, Andrew E. Turner, Tor M. Aamodt |
| 2010 | High-level performance modeling of task-based algorithms. | Alexei Alexandrov, Douglas Armstrong, Hrabri Rajic, Michael Voss, Donald Hayes |
| 2010 | LagAlyzer: A latency profile analysis and visualization tool. | Andrea Adamoli, Milan Jovic, Matthias Hauswirth |
| 2009 | WARP: Enabling fast CPU scheduler development and evaluation. | Haoqiang Zheng, Jason Nieh |
| 2009 | Analyzing the impact of on-chip network traffic on program phases for CMPs. | Yu Zhang, Berkin zisikyilmaz, Gokhan Memik, John Kim, Alok N. Choudhary |
| 2009 | Accuracy of performance counter measurements. | Dmitrijs Zaparanuks, Milan Jovic, Matthias Hauswirth |
| 2009 | Experiment flows and microbenchmarks for reverse engineering of branch predictor structures. | Vladimir Uzelac, Aleksandar Milenkovic |
| 2009 | Understanding the cost of thread migration for multi-threaded Java applications running on a multicore platform. | Qiming Teng, Peter F. Sweeney, Evelyn Duesterwald |
| 2009 | QUICK: A flexible full-system functional model. | Dam Sunwoo, Joonsoo Kim, Derek Chiou |
| 2009 | Evaluating GPUs for network packet signature matching. | Randy Smith, Neelam Goyal, Justin Ormont, Karthikeyan Sankaralingam, Cristian Estan |
| 2009 | SuiteSpecks and SuiteSpots: A methodology for the automatic conversion of benchmarking programs into intrinsically checkpointed assembly code. | Jeff Ringenberg, Trevor N. Mudge |
| 2009 | Analysis of the TRIPS prototype block predictor. | Nitya Ranganathan, Doug Burger, Stephen W. Keckler |
| 2009 | Exploring speculative parallelism in SPEC2006. | Venkatesan Packirisamy, Antonia Zhai, Wei-Chung Hsu, Pen-Chung Yew, Tin-Fook Ngai |
| 2009 | The data-centricity of Web 2.0 workloads and its impact on server performance. | Moriyoshi Ohara, Priya Nagpurkar, Yohei Ueda, Kazuaki Ishizaki |
| 2009 | CMPSched$im: Evaluating OS/CMP interaction on shared cache management. | Jaideep Moses, Konstantinos Aisopos, Aamer Jaleel, Ravi R. Iyer, Ramesh Illikkal, Donald Newell, Srihari Makineni |