| 2016 | GPU multisplit. | Saman Ashkiani, Andrew A. Davidson, Ulrich Meyer, John D. Owens |
| 2016 | Exploiting accelerators for efficient high dimensional similarity search. | Sandeep R. Agrawal, Christopher M. Dee, Alvin R. Lebeck |
| 2016 | Articulation points guided redundancy elimination for betweenness centrality. | Lei Wang, Fan Yang, Liangji Zhuang, Huimin Cui, Fang Lv, Xiaobing Feng |
| 2015 | Low-overhead software transactional memory with progress guarantees and strong semantics. | Minjia Zhang, Jipeng Huang, Man Cao, Michael D. Bond |
| 2015 | NUMA-aware graph-structured analytics. | Kaiyuan Zhang, Rong Chen, Haibo Chen |
| 2015 | Debugging parallel programs using fork handlers. | Javier Alczar Zapin |
| 2015 | High performance computing of fiber scattering simulation. | Leiming Yu, Yan Zhang, Xiang Gong, Nilay Roy, Lee Makowski, David R. Kaeli |
| 2015 | VirtCL: a framework for OpenCL device abstraction and management. | Yi-Ping You, Hen-Jung Wu, Yeh-Ning Tsai, Yen-Ting Chao |
| 2015 | SYNC or ASYNC: time to fuse for distributed graph-parallel computation. | Chenning Xie, Rong Chen, Haibing Guan, Binyu Zang, Haibo Chen |
| 2015 | Parallelizing a discrete event simulation application using the Habanero-Java multicore library. | Wei-Cheng Xiao, Jisheng Zhao, Vivek Sarkar |
| 2015 | Software partitioning of hardware transactions. | Lingxiang Xiang, Michael L. Scott |
| 2015 | Design and evaluation of a novel dataflow based bigdata solution. | Yao Wu, Long Zheng, Brian Heilig, Guang R. Gao |
| 2015 | Efficient and reasonable object-oriented concurrency. | Scott West, Sebastian Nanz, Bertrand Meyer |
| 2015 | CRA: a dynamic task allocation algorithm for many-core processor. | Chang Wang, Jiang Jiang, Yongxing Zhu, Xu Liu, Xing Han |
| 2015 | Gunrock: a high-performance graph processing library on the GPU. | Yangzihao Wang, Andrew A. Davidson, Yuechao Pan, Yuduo Wu, Andy Riffel, John D. Owens |
| 2015 | A programming model and runtime system for significance-aware energy-efficient computing. | Vassilis Vassiliadis, Konstantinos Parasyris, Charalambos Chalios, Christos D. Antonopoulos, Spyros Lalis, Nikolaos Bellas, Hans Vandierendonck, Dimitrios S. Nikolopoulos |
| 2015 | Adaptive GPU cache bypassing. | Yingying Tian, Sooraj Puthoor, Joseph L. Greathouse, Bradford M. Beckmann, Daniel A. Jimnez |
| 2015 | The lazy happens-before relation: better partial-order reduction for systematic concurrency testing. | Paul Thomson, Alastair F. Donaldson |
| 2015 | Scalable and efficient implementation of 3d unstructured meshes computation: a case study on matrix assembly. | Loc Thbault, Eric Petit, Quang Dinh |
| 2015 | A comparative investigation of device-specific mechanisms for exploiting HPC accelerators. | Ayman Tarakji, Lukas Brger, Rainer Leupers |
| 2015 | Cache-oblivious wavefront: improving parallelism of recursive dynamic programming algorithms without losing cache-efficiency. | Yuan Tang, Ronghui You, Haibin Kan, Jesmin Jahan Tithi, Pramod Ganapathi, Rezaul Alam Chowdhury |
| 2015 | Diagnosing the causes and severity of one-sided message contention. | Nathan R. Tallent, Abhinav Vishnu, Hubertus Van Dam, Jeff Daily, Darren J. Kerbyson, Adolfy Hoisie |
| 2015 | Optimization of asynchronous graph processing on GPU with hybrid coloring model. | Xuanhua Shi, Junling Liang, Sheng Di, Bingsheng He, Hai Jin, Lu Lu, Zhixiang Wang, Xuan Luo, Jianlong Zhong |
| 2015 | Thread-level parallelization and optimization of NWChem for the Intel MIC architecture. | Hongzhang Shan, Samuel Williams, Wibe de Jong, Leonid Oliker |
| 2015 | GStream: a graph streaming processing method for large-scale graphs on GPUs. | Hyunseok Seo, Jinwook Kim, Min-Soo Kim |