| 2015 | Towards batched linear solvers on accelerated hardware platforms. | Azzam Haidar, Tingxing Dong, Piotr Luszczek, Stanimire Tomov, Jack J. Dongarra |
| 2015 | More than you ever wanted to know about synchronization: synchrobench, measuring the impact of the synchronization on concurrent algorithms. | Vincent Gramoli |
| 2015 | Automatic scalable atomicity via semantic locking. | Guy Golan-Gueta, G. Ramalingam, Mooly Sagiv, Eran Yahav |
| 2015 | Section based program analysis to reduce overhead of detecting unsynchronized thread communication. | Madan Mohan Das, Gabriel Southern, Jose Renau |
| 2015 | Effects of source-code optimizations on GPU performance and energy consumption. | Jared Coplin, Martin Burtscher |
| 2015 | Dynamic deadlock verification for general barrier synchronisation. | Tiago Cogumbreiro, Raymond Hu, Francisco Martins, Nobuko Yoshida |
| 2015 | Tiles: a new language mechanism for heterogeneous parallelism. | Yifeng Chen, Xiang Cui, Hong Mei |
| 2015 | A parallel algorithm for global states enumeration in concurrent systems. | Yen-Jung Chang, Vijay K. Garg |
| 2015 | Exploiting communication concurrency on high performance computing systems. | Nicholas Chaimov, Khaled Z. Ibrahim, Samuel Williams, Costin Iancu |
| 2015 | Barrier elision for production parallel programs. | Milind Chabbi, Wim Lavrijsen, Wibe de Jong, Koushik Sen, John M. Mellor-Crummey, Costin Iancu |
| 2015 | High performance locks for multi-level NUMA systems. | Milind Chabbi, Michael W. Fagan, John M. Mellor-Crummey |
| 2015 | GPU-SM: shared memory multi-GPU programming. | Javier Cabezas, Marc Jord, Isaac Gelado, Nacho Navarro, Wen-mei W. Hwu |
| 2015 | A framework for practical parallel fast matrix multiplication. | Austin R. Benson, Grey Ballard |
| 2015 | RaftLib: a C++ template library for high performance stream parallel processing. | Jonathan C. Beard, Peng Li, Roger D. Chamberlain |
| 2015 | Toward an evolutionary task parallel integrated MPI + X programming model. | Richard F. Barrett, Dylan T. Stark, Courtenay T. Vaughan, Ryan E. Grant, Stephen L. Olivier, Kevin T. Pedretti |
| 2015 | Performance implications of dynamic memory allocators on transactional memory systems. | Alexandro Baldassin, Edson Borin, Guido Araujo |
| 2015 | On optimizing machine learning workloads via kernel fusion. | Arash Ashari, Shirish Tatikonda, Matthias Boehm, Berthold Reinwald, Keith Campbell, John Keenleyside, P. Sadayappan |
| 2015 | Programming support for reconfigurable custom vector architectures. | Mehmet Ali Arslan, Krzysztof Kuchcinski, Flavius Gruian, Yangxurui Liu |
| 2015 | Predicate RCU: an RCU for scalable concurrent updates. | Maya Arbel, Adam Morrison |
| 2015 | Energy efficiency and performance frontiers for sparse computations on GPU supercomputers. | Hartwig Anzt, Stanimire Tomov, Jack J. Dongarra |
| 2015 | MPI+Threads: runtime contention and remedies. | Abdelhalim Amer, Huiwei Lu, Yanjie Wei, Pavan Balaji, Satoshi Matsuoka |
| 2015 | A performance study of Java garbage collectors on multicore architectures. | Maria Carpen-Amarie, Patrick Marlier, Pascal Felber, Gal Thomas |
| 2015 | SemCache++: semantics-aware caching for efficient multi-GPU offloading. | Nabeel AlSaber, Milind Kulkarni |
| 2015 | The SprayList: a scalable relaxed priority queue. | Dan Alistarh, Justin Kopinsky, Jerry Li, Nir Shavit |
| 2015 | PLUTO+: near-complete modeling of affine transformations for parallelism and locality. | Aravind Acharya, Uday Bondhugula |