| 2013 | CUPL: a compile-time uncoalesced memory access pattern locator for CUDA. | Madhur Amilkanthwar, Shankar Balachandran |
| 2013 | Improving performance of all-to-all communication through loop scheduling in PGAS environments. | Michail Alvanos, Gabriel Tanase, Montse Farreras, Ettore Tiotto, Jos Nelson Amaral, Xavier Martorell |
| 2013 | Improving communication in PGAS environments: static and dynamic coalescing in UPC. | Michail Alvanos, Montse Farreras, Ettore Tiotto, Jos Nelson Amaral, Xavier Martorell |
| 2013 | SemCache: semantics-aware caching for efficient GPU offloading. | Nabeel AlSaber, Milind Kulkarni |
| 2013 | Transparently consistent asynchronous shared memory. | Hakan Akkan, Latchesar Ionkov, Michael Lang |
| 2013 | Hybrid approach for data-flow analysis of MPI programs. | Sriram Aananthakrishnan, Greg Bronevetsky, Ganesh Gopalakrishnan |
| 2012 | Locality & utility co-optimization for practical capacity management of shared last level caches. | Dongyuan Zhan, Hong Jiang, Sharad C. Seth |
| 2012 | Fast loop-level data dependence profiling. | Hongtao Yu, Zhiyuan Li |
| 2012 | Channel borrowing: an energy-efficient nanophotonic crossbar architecture with light-weight arbitration. | Yi Xu, Jun Yang, Rami G. Melhem |
| 2012 | The RAMDISK storage accelerator: a method of accelerating I/O performance on HPC systems using RAMDISKs. | Tim Wickberg, Christopher D. Carothers |
| 2012 | Exploiting communication and packaging locality for cost-effective large scale networks. | Keith D. Underwood, Eric Borch |
| 2012 | CVP: an energy-efficient indirect branch prediction with compiler-guided value pattern. | Mingxing Tan, Xianhua Liu, Tong Tong, Xu Cheng |
| 2012 | Composable, non-blocking collective operations on power7 IH. | Gabriel Ilie Tanase, Gheorghe Almsi, Hanhong Xue, Charles Archer |
| 2012 | CRQ-based fair scheduling on composable multicore architectures. | Tao Sun, Hong An, Tao Wang, Haibo Zhang, Xiufeng Sui |
| 2012 | clSpMV: A Cross-Platform OpenCL SpMV Framework on GPUs. | Bor-Yiing Su, Kurt Keutzer |
| 2012 | Sparse matrix-vector multiply on the HICAMP architecture. | John P. Stevenson, Amin Firoozshahian, Alex Solomatnikov, Mark Horowitz, David R. Cheriton |
| 2012 | Enabling and scaling matrix computations on heterogeneous multi-core and multi-GPU systems. | Fengguang Song, Stanimire Tomov, Jack J. Dongarra |
| 2012 | Fault tolerant preconditioned conjugate gradient for sparse linear system solution. | Manu Shantharam, Sowmyalatha Srinivasmurthy, Padma Raghavan |
| 2012 | A design of hybrid operating system for a parallel computer with multi-core and many-core processors. | Mikiko Sato, Go Fukazawa, Kiyohiko Nagamine, Ryuichi Sakamoto, Mitaro Namiki, Kazumi Yoshinaga, Yuichi Tsujita, Atsushi Hori, Yutaka Ishikawa |
| 2012 | UniFI: leveraging non-volatile memories for a unified fault tolerance and idle power management technique. | Somayeh Sardashti, David A. Wood |
| 2012 | Supercomputing operating systems: a naive view from over the fence. | Timothy Roscoe |
| 2012 | An efficient work-distribution strategy for gridding radio-telescope data on GPUs. | John W. Romein |
| 2012 | Apricot: an optimizing compiler and productivity tool for x86-compatible many-core coprocessors. | Nishkam Ravi, Yi Yang, Tao Bao, Srimat T. Chakradhar |
| 2012 | Hardware support for enforcing isolation in lock-based parallel programs. | Paruj Ratanaworabhan, Martin Burtscher, Darko Kirovski, Benjamin G. Zorn |
| 2012 | Space-round tradeoffs for MapReduce computations. | Andrea Pietracaprina, Geppino Pucci, Matteo Riondato, Francesco Silvestri, Eli Upfal |