| 2014 | Writing scalable SIMD programs with ISPC. | James C. Brodman, Dmitry Babokin, Ilia Filippov, Peng Tu |
| 2014 | Graphs & networks: computing and analytics at lincoln laboratory. | Robert Bond |
| 2014 | Detecting silent data corruption through data dynamic monitoring for scientific applications. | Leonardo Arturo Bautista-Gomez, Franck Cappello |
| 2014 | Singe: leveraging warp specialization for high performance on GPUs. | Michael Bauer, Sean Treichler, Alex Aiken |
| 2014 | Future directions in analytic applications. | Edward J. Baranoski |
| 2014 | Rigorous specification and low-latency implementation of technical market indicators. | Konstantin Bakanov, Ivor T. A. Spence, Hans Vandierendonck, Charles J. Gillan |
| 2014 | Parallelization hints via code skeletonization. | Cfir Aguston, Yosi Ben-Asher, Gadi Haber |
| 2014 | Provably good scheduling for parallel programs that use data structures through implicit batching. | Kunal Agrawal, Jeremy T. Fineman, Brendan Sheridan, Jim Sukha, Robert Utterback |
| 2014 | Data structures for task-based priority scheduling. | Martin Wimmer, Francesco Versaci, Jesper Larsson Trff, Daniel Cederman, Philippas Tsigas |
| 2013 | WuKong: effective diagnosis of bugs at large system scales. | Bowen Zhou, Milind Kulkarni, Saurabh Bagchi |
| 2013 | Array dataflow analysis for polyhedral X10 programs. | Tomofumi Yuki, Paul Feautrier, Sanjay V. Rajopadhye, Vijay A. Saraswat |
| 2013 | Exploring different automata representations for efficient regular expression matching on GPUs. | Xiaodong Yu, Michela Becchi |
| 2013 | StreamScan: fast scan algorithms for GPUs without global barrier synchronization. | Shengen Yan, Guoping Long, Yunquan Zhang |
| 2013 | A peta-scalable CPU-GPU algorithm for global atmospheric simulations. | Chao Yang, Wei Xue, Haohuan Fu, Lin Gan, Linfeng Li, Yangtong Xu, Yutong Lu, Jiachang Sun, Guangwen Yang, Weimin Zheng |
| 2013 | X10-FT: transparent fault tolerance for APGAS language and runtime. | Chenning Xie, Zhijun Hao, Haibo Chen |
| 2013 | Compiler aided manual speculation for high performance concurrent data structures. | Lingxiang Xiang, Michael Lee Scott |
| 2013 | Complexity analysis and algorithm design for reorganizing data to minimize non-coalesced memory accesses on GPU. | Bo Wu, Zhijia Zhao, Eddy Zheng Zhang, Yunlian Jiang, Xipeng Shen |
| 2013 | Swift/T: scalable data flow programming for many-task applications. | Justin M. Wozniak, Timothy G. Armstrong, Michael Wilde, Daniel S. Katz, Ewing L. Lusk, Ian T. Foster |
| 2013 | CAP: co-scheduling based on asymptotic profiling in CPU+GPU hybrid systems. | Zhenning Wang, Long Zheng, Quan Chen, Minyi Guo |
| 2013 | Parallel time-space processing model based fast | Wei Wang, Hanli Wang, Dong Guo, Haoyang Wei, Guosun Zeng |
| 2013 | libEOMP: a portable OpenMP runtime library based on MCA APIs for embedded systems. | Cheng Wang, Sunita Chandrasekaran, Barbara M. Chapman, Jim Holt |
| 2013 | FastLane: improving performance of software transactional memory for low thread counts. | Jons-Tobias Wamhoff, Christof Fetzer, Pascal Felber, Etienne Rivire, Gilles Muller |
| 2013 | Pyjama: OpenMP-like implementation for Java, with GUI extensions. | Vikas, Nasser Giacaman, Oliver Sinnen |
| 2013 | The JStar language philosophy. | Mark Utting, Min-Hsien Weng, John G. Cleary |
| 2013 | Reducing contention through priority updates. | Julian Shun, Guy E. Blelloch, Jeremy T. Fineman, Phillip B. Gibbons |