| 2013 | Improving performance of openSHMEM reference library by portable PE mapping technique. | Swaroop Pophale, Tony Curtis, Barbara M. Chapman |
| 2013 | Exploring hardware overprovisioning in power-constrained, high performance computing. | Tapasya Patki, David K. Lowenthal, Barry Rountree, Martin Schulz, Bronis R. de Supinski |
| 2013 | Scaling data race detection for partitioned global address space programs. | Chang-Seo Park, Koushik Sen, Costin Iancu |
| 2013 | Prefetching and cache management using task lifetimes. | Vassilis Papaefstathiou, Manolis Katevenis, Dimitrios S. Nikolopoulos, Dionisios N. Pnevmatikatos |
| 2013 | Inspector/executor load balancing algorithms for block-sparse tensor contractions. | David Ozog, Sameer Shende, Allen D. Malony, Jeff R. Hammond, James Dinan, Pavan Balaji |
| 2013 | Design and implementation of a customizable work stealing scheduler. | Jun Nakashima, Sho Nakatani, Kenjiro Taura |
| 2013 | Diagnosis and optimization of application prefetching performance. | Gabriel Marin, Collin McCurdy, Jeffrey S. Vetter |
| 2013 | Memory-conscious collective I/O for extreme scale HPC systems. | Yin Lu, Yong Chen, Yu Zhuang, Rajeev Thakur |
| 2013 | Efficient sparse matrix-vector multiplication on x86-based many-core processors. | Xing Liu, Mikhail Smelyanskiy, Edmond Chow, Pradeep Dubey |
| 2013 | A new approach for performance analysis of openMP programs. | Xu Liu, John M. Mellor-Crummey, Michael W. Fagan |
| 2013 | Exploiting domain knowledge to optimize parallel computational mechanics codes. | Chenyang Liu, Muhammad Hasan Jamal, Milind Kulkarni, Arun Prakash, Vijay S. Pai |
| 2013 | Address-aware fences. | Changhui Lin, Vijay Nagarajan, Rajiv Gupta |
| 2013 | SMIO: I/O similarity aware virtual machine management invirtual desktop environments. | Min Li, Sushil Mantri, Pin Zhou, Ali Raza Butt |
| 2013 | Evaluating the feasibility of using memory content similarity to improve system resilience. | Scott Levy, Patrick G. Bridges, Kurt B. Ferreira, Aidan P. Thompson, Christian R. Trott |
| 2013 | Automatically adapting programs for mixed-precision floating-point computation. | Michael O. Lam, Jeffrey K. Hollingsworth, Bronis R. de Supinski, Matthew P. LeGendre |
| 2013 | Towards more efficient execution: a decoupled access-execute approach. | Konstantinos Koukos, David Black-Schaffer, Vasileios Spiliopoulos, Stefanos Kaxiras |
| 2013 | Quantifying performance bottleneck cost through differential analysis. | Souad Koliai, Zakaria Bendifallah, Mathieu Tribalat, Cdric Valensi, Jean-Thomas Acquaviva, William Jalby |
| 2013 | An automatic input-sensitive approach for heterogeneous task partitioning. | Klaus Kofler, Ivan Grasso, Biagio Cosenza, Thomas Fahringer |
| 2013 | Enabling accurate power profiling of HPC applications on exascale systems. | Gokcen Kestor, Roberto Gioiosa, Darren J. Kerbyson, Adolfy Hoisie |
| 2013 | Imogen: a parallel 3D fluid and MHD code for GPUs. | Erik Keever, James N. Imamura |
| 2013 | Characteristics of | Laxmikant V. Kal |
| 2013 | Design of a large-scale storage-class RRAM system. | Myoungsoo Jung, John Shalf, Mahmut T. Kandemir |
| 2013 | Memorage: emerging persistent RAM based malleable main memory and storage architecture. | Ju-Young Jung, Sangyeun Cho |
| 2013 | Tuning the continual flow pipeline architecture. | Komal Jothi, Haitham Akkary |
| 2013 | The ARMv8 simulator. | Tao Jiang, Lele Zhang, Rui Hou, Yi Zhang, Qianlong Zhang, Lin Chai, Jing Han, Wuxiong Zhang, Cong Wang, Lixin Zhang |