| 2016 | A Benchmark on Multi Improvement Neighborhood Search Strategies in CPU/GPU Systems. | Eyder Rios, Igor Machado Coelho, Luiz Satoru Ochi, Cristina Boeres, Ricardo C. Farias |
| 2016 | A Scalable Algorithm for Simulating the Structural Plasticity of the Brain. | Sebastian Rinke, Markus Butz-Ostendorf, Marc-Andr Hermanns, Mikal Naveau, Felix Wolf |
| 2016 | Towards a GPU Abstraction for Lua. | Raphael Ribeiro, Paulo Motta |
| 2016 | Breadth-First Search on Heterogeneous Platforms: A Case of Study on Social Networks. | Luis Remis, Mara Jess Garzarn, Rafael Asenjo, Angeles G. Navarro |
| 2016 | REPP-H: Runtime Estimation of Power and Performance on Heterogeneous Data Centers. | Rajiv Nishtala, Xavier Martorell, Vinicius Petrucci, Daniel Moss |
| 2016 | Using Balanced Data Placement to Address I/O Contention in Production Environments. | Sarah Neuwirth, Feiyi Wang, Sarp Oral, Sudharshan Vazhkudai, James H. Rogers, Ulrich Brning |
| 2016 | A Hybrid Parallel Algorithm for the Auction Algorithm in Multicore Systems. | Aline de Paula Nascimento, Cristina Nader Vasconcelos, F. S. Jamel, Alexandre da Costa Sena |
| 2016 | Value Reuse Potential in ARM Architectures. | Rodrigo Costa de Moura, Giovane O. Torres, Maurcio L. Pilla, Larcio Lima Pilla, Amarildo T. da Costa, Felipe Maia Galvo Frana |
| 2016 | STOMP: Statistical Techniques for Optimizing and Modeling Performance of Blocked Sparse Matrix Vector Multiplication. | Steena Monteiro, Forrest N. Iandola, Daniel Wong |
| 2016 | MAGC: A Mapping Approach for GPU Clusters. | Seyed Hessam Mirsadeghi, Iman Faraji, Ahmad Afsahi |
| 2016 | A Processor Workload Distribution Algorithm for Massively Parallel Applications. | Serge Midonnet, Achille Wattelar |
| 2016 | Automatic Insertion of Copy Annotation in Data-Parallel Programs. | Gleison Souza Diniz Mendonca, Breno Campos Ferreira Guimares, Pricles Rafael Oliveira Alves, Fernando Magno Quinto Pereira, Mrcio Machado Pereira, Guido Araujo |
| 2016 | Parallel Pairwise Correlation Computation on Intel Xeon Phi Clusters. | Yongchao Liu, Tony Pan, Srinivas Aluru |
| 2016 | Planning Your SQL-on-Hadoop Deployment Using a Low-Cost Simulation-Based Approach. | Jun Liu, Bianny Bian, Samantika Subramaniam Sury |
| 2016 | Optimisation of a Molecular Dynamics Simulation of Chromosome Condensation. | Timothy R. Law, Jonny Hancox, Tammy M. K. Cheng, Raphael A. G. Chaleil, Steven A. Wright, Paul A. Bates, Stephen A. Jarvis |
| 2016 | Synchronization-Free Automatic Parallelization for Arbitrarily Nested Affine Loops. | Tomasz Klimek, Marek Palkowski, Wlodzimierz Bielecki |
| 2016 | Dataflow to Hardware Synthesis Framework on FPGAs. | Youngsoo Kim, Shrikant Jadhav, Clay S. Gloster Jr. |
| 2016 | Dynamic Inter-Thread Vectorization Architecture: Extracting DLP from TLP. | Sajith Kalathingal, Caroline Collange, Bharath Narasimha Swamy, Andr Seznec |
| 2016 | Parallelism and Scalability: A Solution Focused on the Cloud Computing Processing Service Billing. | Emmanoel M. De Sousa Junior, Idalmis Milin Sardia, Frederico Lopes |
| 2016 | Empirical, Analytical Study of Hardware-Based Page Swap in Hybrid Main Memory System. | Ju-Young Jung, Rami G. Melhem |
| 2016 | HYPPO: A Hybrid, Piecewise Polynomial Modeling Technique for Non-Smooth Surfaces. | Travis Johnston, Connor Zanin, Michela Taufer |
| 2016 | Partitioning GPUs for Improved Scalability. | Johan Janzen, David Black-Schaffer, Andra Hugo |
| 2016 | Speeding Up Stencil Computations with Kernel Convolution. | Guilherme C. Januario, Bryan S. Rosenburg, Yoonho Park, Michael Perrone, Jos E. Moreira, Tereza Cristina M. B. Carvalho |
| 2016 | An Image Processing VLIW Architecture for Real-Time Depth Detection. | Dan Iorga, Razvan Nane, Yi Lu, Edwin van Dalen, Koen Bertels |
| 2016 | Performance Optimization for SpMV on Multi-GPU Systems Using Threads and Multiple Streams. | Ping Guo, Changjiang Zhang |