| 2014 | Thread-cooperative, bit-parallel computation of levenshtein distance on GPU. | Alejandro Chacn, Santiago Marco-Sola, Antonio Espinosa, Paolo Ribeca, Juan Carlos Moure |
| 2014 | Author retrospective for optimizing matrix multiply using PHiPAC: a portable high-performance ANSI C coding methodology. | Jeff A. Bilmes, Krste Asanovic, Chee-Whye Chin, Jim Demmel |
| 2014 | Parallelizing and optimizing sparse tensor computations. | Muthu Manikandan Baskaran, Benot Meister, Richard Lethin |
| 2014 | Scalable performance analysis of exascale MPI programs through signature-based clustering algorithms. | Amir Bahmani, Frank Mueller |
| 2014 | An efficient two-dimensional blocking strategy for sparse matrix-vector multiplication on GPUs. | Arash Ashari, Naser Sedaghati, John Eisenlohr, P. Sadayappan |
| 2013 | Exploiting uniform vector instructions for GPGPU performance, energy efficiency, and opportunistic reliability enhancement. | Ping Xiang, Yi Yang, Mike Mantor, Norm Rubin, Lisa R. Hsu, Huiyang Zhou |
| 2013 | Elastic and scalable tracing and accurate replay of non-deterministic events. | Xing Wu, Frank Mueller |
| 2013 | Power efficiency in a partially reconfigurable multiprocessor system. | Raymond J. Weber, Justin A. Hogan, Brock J. LaMeres, Todd Kaiser |
| 2013 | V-OpenCL: a method to use remote GPGPU. | Cong Wang, Tao Jiang, Rui Hou |
| 2013 | Bubble coloring: avoiding routing- and protocol-induced deadlocks with minimal virtual channel requirement. | Ruisheng Wang, Lizhong Chen, Timothy Mark Pinkston |
| 2013 | G-Charm: an adaptive runtime system for message-driven parallel applications on hybrid systems. | R. Vasudevan, Sathish S. Vadhiyar, Laxmikant V. Kal |
| 2013 | Exploiting reuse information to reduce refresh energy in on-chip eDRAM caches. | Alejandro Valero, Julio Sahuquillo, Salvador Petit, Jos Duato |
| 2013 | Evaluating on-die interconnects for a 4 TB/s router. | Keith D. Underwood, Eric Borch, John Sizer, Timothy Stremcha, Michael Strom |
| 2013 | Function, latency, bandwidth, power: towards a better computer. | Steven L. Teig |
| 2013 | HykSort: a new variant of hypercube quicksort on distributed memory architectures. | Hari Sundar, Dhairya Malhotra, George Biros |
| 2013 | Abstractions to separate concerns in semi-regular grids. | Andrew Stone, Michelle Mills Strout |
| 2013 | Holistic run-time parallelism management for time and energy efficiency. | Srinath Sridharan, Gagan Gupta, Gurindar S. Sohi |
| 2013 | Towards shared memory consistency models for GPUs. | Tyler Sorensen, Ganesh Gopalakrishnan, Vinod Grover |
| 2013 | A gossip-based approach to exascale system services. | Philip Soltero, Patrick G. Bridges, Dorian C. Arnold, Michael Lang |
| 2013 | The role of computer designers in reverse-engineering the brain. | James E. Smith |
| 2013 | Using platform-independent data locality analysis to predict cache performance on abstract hardware platforms. | Sonish Shrestha |
| 2013 | Scaling large-data computations on multi-GPU accelerators. | Amit Sabne, Putt Sakdhnagool, Rudolf Eigenmann |
| 2013 | The power 775 architecture at scale. | Ramakrishnan Rajamony, Mark W. Stephenson, William Evan Speight |
| 2013 | Bandwidth-optimal all-to-all exchanges in fat tree networks. | Bogdan Prisacari, Germn Rodrguez, Cyriel Minkenberg, Torsten Hoefler |
| 2013 | CMP off-chip bandwidth scheduling guided by instruction criticality. | Pablo Prieto, Valentin Puente, Jos-ngel Gregorio |