| 2015 | CilkSpec: optimistic concurrency for Cilk. | Shaizeen Aga, Sriram Krishnamoorthy, Satish Narayanasamy |
| 2015 | A work-efficient algorithm for parallel unordered depth-first search. | Umut A. Acar, Arthur Charguraud, Mike Rainey |
| 2015 | Lost in heterogeneity: architectural selection based on code features. | Sameer AbuAsal, R. Tohid, J. Ramanujam |
| 2014 | Nonblocking Epochs in MPI One-Sided Communication. | Judicael A. Zounmevo, Xin Zhao, Pavan Balaji, William Gropp, Ahmad Afsahi |
| 2014 | BatchFS: scaling the file system control plane with client-funded metadata servers. | Qing Zheng, Kai Ren, Garth A. Gibson |
| 2014 | Trapped capacity: scheduling under a power cap to maximize machine-room throughput. | Ziming Zhang, Michael Lang, Scott Pakin, Song Fu |
| 2014 | CYPRESS: Combining Static and Dynamic Analysis for Top-Down Communication Trace Compression. | Jidong Zhai, Jianfei Hu, Xiongchao Tang, Xiaosong Ma, Wenguang Chen |
| 2014 | Teaching high performance computing: lessons from a flipped classroom, project-based course on finite element methods. | Jill Zarestky, Wolfgang Bangerth |
| 2014 | Distributed control: priority scheduling for single source shortest paths without synchronization. | Marcin Zalewski, Thejaka Amila Kanewala, Jesun Sahariar Firoz, Andrew Lumsdaine |
| 2014 | Quantitatively Modeling Application Resilience with the Data Vulnerability Factor. | Li Yu, Dong Li, Sparsh Mittal, Jeffrey S. Vetter |
| 2014 | Fast Iterative Graph Computation: A Path Centric Approach. | Pingpeng Yuan, Wenya Zhang, Changfeng Xie, Hai Jin, Ling Liu, Kisung Lee |
| 2014 | Rethinking key-value store for parallel I/O optimization. | Yanlong Yin, Antonios Kougkas, Kun Feng, Hassan Eslami, Yin Lu, Xian-He Sun, Rajeev Thakur, William Gropp |
| 2014 | Deflation strategies to improve the convergence of communication-avoiding GMRES. | Ichitaro Yamazaki, Stanimire Tomov, Jack J. Dongarra |
| 2014 | Domain Decomposition Preconditioners for Communication-Avoiding Krylov Methods on a Hybrid CPU/GPU Cluster. | Ichitaro Yamazaki, Sivasankaran Rajamanickam, Erik G. Boman, Mark Hoemmen, Michael A. Heroux, Stanimire Tomov |
| 2014 | MSL: A Synthesis Enabled Language for Distributed Implementations. | Zhilei Xu, Shoaib Kamil, Armando Solar-Lezama |
| 2014 | VSFS: a searchable distributed file system. | Lei Xu, Ziling Huang, Hong Jiang, Lei Tian, David R. Swanson |
| 2014 | Accelerating Kirchhoff migration on GPU using directives. | Rengan Xu, Maxime R. Hugues, Henri Calandra, Sunita Chandrasekaran, Barbara M. Chapman |
| 2014 | CommGram: a new visual analytics tool for large communication trace data. | Jieting Wu, Jianping Zeng, Hongfeng Yu, Joseph P. Kenny |
| 2014 | Integrating Pig with Harp to support iterative applications with fast cache and customized communication. | Tak-Lon Wu, Abhilash Koppula, Judy Qiu |
| 2014 | mPPM, viewed as a co-design effort. | Paul R. Woodward, Jagan Jayaraj, Richard Barrett |
| 2014 | Heterogeneous concurrent execution of Monte Carlo photon transport on CPU, GPU and MIC. | Noah Wolfe, Tianyu Liu, Christopher D. Carothers, Xie George Xu |
| 2014 | Visualization of memory access behavior on hierarchical NUMA architectures. | Benjamin Weyers, Christian Terboven, Dirk Schmidl, Joachim Herber, Torsten W. Kuhlen, Matthias S. Mller, Bernd Hentschel |
| 2014 | The anatomy of Mr. Scan: a dissection of performance of an extreme scale GPU-based clustering algorithm. | Benjamin Welton, Barton P. Miller |
| 2014 | A hierarchical tridiagonal system solver for heterogenous supercomputers. | Xinliang Wang, Yangtong Xu, Wei Xue |
| 2014 | BPAR: a bundle-based parallel aggregation framework for decoupled I/O execution. | Teng Wang, Kevin Vasko, Zhuo Liu, Hui Chen, Weikuan Yu |