| 2012 | Quantifying the effectiveness of load balance algorithms. | Olga Pearce, Todd Gamblin, Bronis R. de Supinski, Martin Schulz, Nancy M. Amato |
| 2012 | High performance supercomputers: should the individual processor be more than a brick? | Yale N. Patt |
| 2012 | Overcoming single-thread performance hurdles in the core fusion reconfigurable multicore architecture. | Janani Mukundan, Saugata Ghose, Robert Karmazin, Engin Ipek, Jos F. Martnez |
| 2012 | Collective algorithms for sub-communicators. | Anshul Mittal, Nikhil Jain, Thomas George, Yogish Sabharwal, Sameer Kumar |
| 2012 | A file I/O system for many-core based clusters. | Yuki Matsuo, Taku Shimosawa, Yutaka Ishikawa |
| 2012 | Distributed replay protocol for distributed uniprocessors. | Mengjie Mao, Hong An, Bobin Deng, Tao Sun, Xuechao Wei, Wei Zhou, Wenting Han |
| 2012 | On the scalability of the clusters-booster concept: a critical assessment of the DEEP architecture. | Damin Alvarez Malln, Norbert Eicker, Maria Elena Innocenti, Giovanni Lapenta, Thomas Lippert, Estela Suarez |
| 2012 | Data-driven fault tolerance for work stealing computations. | Wenjing Ma, Sriram Krishnamoorthy |
| 2012 | Congestion avoidance on manycore high performance computing systems. | Miao Luo, Dhabaleswar K. Panda, Khaled Z. Ibrahim, Costin Iancu |
| 2012 | An optimized large-scale hybrid DGEMM design for CPUs and ATI GPUs. | Jiajia Li, Xingjian Li, Guangming Tan, Mingyu Chen, Ninghui Sun |
| 2012 | Node-based memory management for scalable NUMA architectures. | Stefan Lankes, Thomas Bemmerl, Thomas Roehl, Christian Terboven |
| 2012 | Optimizing latency and throughput for spawning processes on massively multicore processors. | Abhishek Kulkarni, Andrew Lumsdaine, Michael Lang, Latchesar Ionkov |
| 2012 | Better than native: using virtualization to improve compute node performance. | Brian Kocoloski, John R. Lange |
| 2012 | SnuCL: an OpenCL framework for heterogeneous CPU/GPU clusters. | Jungwon Kim, Sangmin Seo, Jun Lee, Jeongho Nah, Gangwon Jo, Jaejin Lee |
| 2012 | Integrated in-system storage architecture for high performance computing. | Dries Kimpe, Kathryn Mohror, Adam Moody, Brian Van Essen, Maya B. Gokhale, Robert B. Ross, Bronis R. de Supinski |
| 2012 | Enhancing the performance of assisted execution runtime systems through hardware/software techniques. | Gokcen Kestor, Roberto Gioiosa, Osman S. Unsal, Adrin Cristal, Mateo Valero |
| 2012 | An analysis of computational workloads for the ORNL Jaguar system. | Wayne Joubert, Shi-Quan Su |
| 2012 | Characterizing and improving the use of demand-fetched caches in GPUs. | Wenhao Jia, Kelly A. Shaw, Margaret Martonosi |
| 2012 | Unified memory optimizing architecture: memory subsystem control with a unified predictor. | Yasuo Ishii, Mary Inaba, Kei Hiraki |
| 2012 | High-performance code generation for stencil computations on GPU architectures. | Justin Holewinski, Louis-Nol Pouchet, P. Sadayappan |
| 2012 | HiRe: using hint & release to improve synchronization of speculative threads. | Liang Han, Xiaowei Jiang, Wei Liu, Youfeng Wu, James Tuck |
| 2012 | One stone two birds: synchronization relaxation and redundancy removal in GPU-CPU translation. | Ziyu Guo, Bo Wu, Xipeng Shen |
| 2012 | Multiple sub-row buffers in DRAM: unlocking performance and energy improvement opportunities. | Nagendra Dwarakanath Gulur, R. Manikantan, Mahesh Mehendale, R. Govindarajan |
| 2012 | Blue Gene/Q: design for sustained multi-petaflop computing. | Michael Gschwind |
| 2012 | GPU merge path: a GPU merging algorithm. | Oded Green, Robert McColl, David A. Bader |