| 2015 | Static/Dynamic validation of MPI collective communications in multi-threaded context. | Emmanuelle Saillard, Patrick Carribault, Denis Barthou |
| 2015 | Effective communication for a system of cluster-on-a-chip processors. | Pablo Reble, Stefan Lankes, Fabian Fischer, Matthias S. Mller |
| 2015 | Forma: a DSL for image processing applications to target GPUs and multi-core CPUs. | Mahesh Ravishankar, Justin Holewinski, Vinod Grover |
| 2015 | Distributed memory code generation for mixed Irregular/Regular computations. | Mahesh Ravishankar, Roshan Dathathri, Venmugil Elango, Louis-Nol Pouchet, J. Ramanujam, Atanas Rountev, P. Sadayappan |
| 2015 | CASTLE: fast concurrent internal binary search tree using edge-based locking. | Arunmoezhi Ramachandran, Neeraj Mittal |
| 2015 | Are web applications ready for parallelism? | Cosmin Radoi, Stephan Herhut, Jaswanth Sreeram, Danny Dig |
| 2015 | GPU technology applied to reverse time migration and seismic modeling via OpenACC. | Ahmad Qawasmeh, Barbara M. Chapman, Maxime R. Hugues, Henri Calandra |
| 2015 | JAWS: a JavaScript framework for adaptive CPU-GPU work sharing. | Xianglan Piao, Channoh Kim, Younghwan Oh, Huiying Li, Jincheon Kim, Hanjun Kim, Jae W. Lee |
| 2015 | Decoupled load balancing. | Olga Pearce, Todd Gamblin, Bronis R. de Supinski, Martin Schulz, Nancy M. Amato |
| 2015 | A Java util concurrent park contention tool. | Panagiotis Patros, Eric Aubanel, David Bremner, Michael Dawson |
| 2015 | Rethinking the parallelization of random-restart hill climbing: a case study in optimizing a 2-opt TSP solver for GPU execution. | Molly A. O'Neil, Martin Burtscher |
| 2015 | A collection-oriented programming model for performance portability. | Saurav Muralidharan, Michael Garland, Bryan Catanzaro, Albert Sidelnik, Mary W. Hall |
| 2015 | Patty: a pattern-based parallelization tool for the multicore age. | Korbinian Molitorisz, Tobias Mller, Walter F. Tichy |
| 2015 | Fence placement for legacy data-race-free programs via synchronization read detection. | Andrew J. McPherson, Vijay Nagarajan, Susmit Sarkar, Marcelo Cintra |
| 2015 | A library for portable and composable data locality optimizations for NUMA systems. | Zoltan Maj, Thomas R. Gross |
| 2015 | Parallelism vs. speculation: exploiting speculative genetic algorithm on GPU. | Yanchao Lu, Long Zheng, Li Li, Minyi Guo |
| 2015 | Helium: a transparent inter-kernel optimizer for OpenCL. | Thibaut Lutz, Christian Fensch, Murray Cole |
| 2015 | Deadlock-free buffer configuration for stream computing. | Peng Li, Jonathan C. Beard, Jeremy Buhler |
| 2015 | Energy-efficient computing for HPC workloads on heterogeneous manycore chips. | Akhil Langer, Ehsan Totoni, Udatta S. Palekar, Laxmikant V. Kal |
| 2015 | An OpenACC-based unified programming model for multi-accelerator systems. | Jungwon Kim, Seyong Lee, Jeffrey S. Vetter |
| 2015 | Efficient utilization of GPGPU cache hierarchy. | Mahmoud Khairy, Mohamed Zahran, Amr G. Wassal |
| 2015 | Stochastic gradient descent on GPUs. | Rashid Kaleem, Sreepathi Pai, Keshav Pingali |
| 2015 | Combining phase identification and statistic modeling for automated parallel benchmark generation. | Ye Jin, Mingliang Liu, Xiaosong Ma, Qing Liu, Jeremy Logan, Norbert Podhorszki, Jong Youl Choi, Scott Klasky |
| 2015 | A hierarchical approach to reducing communication in parallel graph algorithms. | Harshvardhan, Nancy M. Amato, Lawrence Rauchwerger |
| 2015 | Optimization for performance and energy for batched matrix computations on GPUs. | Azzam Haidar, Tingxing Dong, Piotr Luszczek, Stanimire Tomov, Jack J. Dongarra |