| 2011 | SecureME: a hardware-software approach to full system security. | Siddhartha Chhabra, Brian Rogers, Yan Solihin, Milos Prvulovic |
| 2011 | Coordinating processor and main memory for efficientserver power control. | Ming Chen, Xiaorui Wang, Xue Li |
| 2011 | Hystor: making the best use of solid state drives in high performance storage systems. | Feng Chen, David A. Koufaty, Xiaodong Zhang |
| 2011 | Predictive coordination of multiple on-chip resources for chip multiprocessors. | Jian Chen, Lizy Kurian John |
| 2011 | Poster: revisiting virtual channel memory for performance and fairness on multi-core architecture. | Licheng Chen, Yongbing Huang, Yungang Bao, Onur Mutlu, Guangming Tan, Mingyu Chen |
| 2011 | An idiom-finding tool for increasing productivity of accelerators. | Laura Carrington, Mustafa M. Tikir, Catherine Olschanowsky, Michael Laurenzano, Joshua Peraza, Allan Snavely, Stephen Poole |
| 2011 | Poster: programming clusters of GPUs with OMPSs. | Javier Bueno, Alejandro Duran, Xavier Martorell, Eduard Ayguad, Rosa M. Badia, Jess Labarta |
| 2011 | SRC: information retrieval as a persistent parallel service on supercomputer infrastructure. | Tobias Berka, Marin Vajtersic |
| 2011 | Karma: scalable deterministic record-replay. | Arkaprava Basu, Jayaram Bobba, Mark D. Hill |
| 2011 | Rethinking shared-memory languages and hardware. | Sarita V. Adve |
| 2010 | Timing local streams: improving timeliness in data prefetching. | Huaiyu Zhu, Yong Chen, Xian-He Sun |
| 2010 | Enigma: architectural and operating system support for reducing the impact of address translation. | Lixin Zhang, Evan Speight, Ramakrishnan Rajamony, Jiang Lin |
| 2010 | Streamlining GPU applications on the fly: thread divergence elimination through runtime thread-data remapping. | Eddy Z. Zhang, Yunlian Jiang, Ziyu Guo, Xipeng Shen |
| 2010 | Untitled record | Xuechen Zhang, Song Jiang |
| 2010 | Cache oblivious parallelograms in iterative stencil computations. | Robert Strzodka, Mohammed Shaheen, Dawid Pajak, Hans-Peter Seidel |
| 2010 | How to unleash array optimizations on code using recursive data structures. | Harmen L. A. van der Spek, C. W. Mattias Holm, Harry A. G. Wijshoff |
| 2010 | Speeding up Nek5000 with autotuning and specialization. | Jaewook Shin, Mary W. Hall, Jacqueline Chame, Chun Chen, Paul F. Fischer, Paul D. Hovland |
| 2010 | Compiler and runtime support for enabling generalized reduction computations on heterogeneous parallel configurations. | Vignesh T. Ravi, Wenjing Ma, David Chiu, Gagan Agrawal |
| 2010 | Adaptive multi-level cache allocation in distributed storage architectures. | Ramya Prabhakar, Shekhar Srikantaiah, Mahmut T. Kandemir, Christina M. Patrick |
| 2010 | Quantifying performance benefits of overlap using MPI-2 in a seismic modeling application. | Sreeram Potluri, Ping Lai, Karen A. Tomko, Sayantan Sur, Yifeng Cui, Mahidhar Tatineni, Karl W. Schulz, William L. Barth, Amitava Majumdar, Dhabaleswar K. Panda |
| 2010 | Handling task dependencies under strided and aliased references. | Josep M. Prez, Rosa M. Badia, Jess Labarta |
| 2010 | Exascale science: the next frontier in high performance computing. | Stephen S. Pawlowski |
| 2010 | A query language for understanding component interactions in production systems. | Adam J. Oliner, Alex Aiken |
| 2010 | Small-ruleset regular expression matching on GPGPUs: quantitative performance analysis and optimization. | Jamin Naghmouchi, Daniele Paolo Scarpazza, Mladen Berekovic |
| 2010 | Overlapping communication and computation by using a hybrid MPI/SMPSs approach. | Vladimir Marjanovic, Jess Labarta, Eduard Ayguad, Mateo Valero |