| 2016 | Exploiting Task-Parallelism in Message-Passing Sparse Linear System Solvers Using OmpSs. | Jos Ignacio Aliaga, Maria Barreda, Matthias Bollhfer, Enrique S. Quintana-Ort |
| 2016 | Parametric Multi-step Scheme for GPU-Accelerated Graph Decomposition into Strongly Connected Components. | Stefano Aldegheri, Jiri Barnat, Nicola Bombieri, Federico Busato, Milan Ceska |
| 2016 | Acceleration of Turbomachinery Steady Simulations on GPU. | Mohamed Hassanine Aissa, Lasse Mller, Tom Verstraete, Cornelis Vuik |
| 2016 | Task-Based Sparse Hybrid Linear Solver for Distributed Memory Heterogeneous Architectures. | Emmanuel Agullo, Luc Giraud, Stojce Nakov |
| 2016 | Task-Based Conjugate Gradient: From Multi-GPU Towards Heterogeneous Architectures. | Emmanuel Agullo, Luc Giraud, Abdou Guermouche, Stojce Nakov, Jean Roman |
| 2016 | Exploiting a Parametrized Task Graph Model for the Parallelization of a Sparse Direct Multifrontal Solver. | Emmanuel Agullo, George Bosilca, Alfredo Buttari, Abdou Guermouche, Florent Lopez |
| 2016 | Automatic Generation of OpenCL Code for ARM Architectures. | Sergio Afonso, Alejandro Acosta, Francisco Almeida |
| 2016 | Workflow Performance Profiles: Development and Analysis. | Dariusz Krl, Rafael Ferreira da Silva, Ewa Deelman, Vickie E. Lynch |
| 2016 | A Synchronization-Free Algorithm for Parallel Sparse Triangular Solves. | Weifeng Liu, Ang Li, Jonathan D. Hogg, Iain S. Duff, Brian Vinter |
| 2016 | Load-Sharing Policies in Parallel Simulation of Agent-Based Demographic Models. | Alessandro Pellegrini, Cristina Montaola-Sales, Francesco Quaglia, Josep Casanovas-Garca |
| 2015 | Leveraging MPI-3 Shared-Memory Extensions for Efficient PGAS Runtime Systems. | Huan Zhou, Kamran Idrees, Jos Gracia |
| 2015 | A Simplified TDP with Large Tables. | Yu Zhang |
| 2015 | Continuation Complexity: A Callback Hell for Distributed Systems. | Edgar Zamora-Gmez, Pedro Garca Lpez, Rubn Mondjar |
| 2015 | Superoptimizing Memory Subsystems for Multiple Objectives. | Joseph G. Wingbermuehle, Ron K. Cytron, Roger D. Chamberlain |
| 2015 | Canaries in a Coal Mine: Using Application-Level Checkpoints to Detect Memory Failures. | Patrick M. Widener, Kurt B. Ferreira, Scott Levy, Nathan Fabian |
| 2015 | Scalable Data-Driven PageRank: Algorithms, System Issues, and Lessons Learned. | Joyce Jiyoung Whang, Andrew Lenharth, Inderjit S. Dhillon, Keshav Pingali |
| 2015 | Fast Parallel Suffix Array on the GPU. | Leyuan Wang, Sean Baxter, John D. Owens |
| 2015 | GPGPU Virtualisation with Multi-API Support Using Containers. | John Walsh, Jonathan Dukes |
| 2015 | 10, 000 Performance Models per Minute - Scalability of the UG4 Simulation Framework. | Andreas Vogel, Alexandru Calotoiu, Alexandre Strube, Sebastian Reiter, Arne Ngel, Felix Wolf, Gabriel Wittum |
| 2015 | Challenges of a Systematic Approach to Parallel Computing and Supercomputing Education. | Vladimir V. Voevodin, Victor Gergel, Nina Popova |
| 2015 | Quantifying the Performance Impact of Graph Structure on Neighbour Iteration Strategies for PageRank. | Merijn Verstraaten, Ana Lucia Varbanescu, Cees de Laat |
| 2015 | Integration of ICT in Concurrent and Parallel Programming Lectures. | Antonio J. Tomeu-Hardasmal, Alberto G. Salguero, Manuel I. Capel |
| 2015 | A Practical Transactional Memory Interface. | Shahar Timnat, Maurice Herlihy, Erez Petrank |
| 2015 | Optimizing Task Parallelism with Library-Semantics-Aware Compilation. | Peter Thoman, Stefan Moosbrugger, Thomas Fahringer |
| 2015 | Optimized Force Calculation in Molecular Dynamics Simulations for the Intel Xeon Phi. | Nikola Tchipev, Amer Wafai, Colin W. Glass, Wolfgang Eckhardt, Alexander Heinecke, Hans-Joachim Bungartz, Philipp Neumann |