| 2017 | Exploiting Vector and Multicore Parallelism for Recursive, Data- and Task-Parallel Programs. | Bin Ren, Sriram Krishnamoorthy, Kunal Agrawal, Milind Kulkarni |
| 2017 | POSTER: A Wait-Free Queue with Wait-Free Memory Reclamation. | Pedro Ramalhete, Andreia Correia |
| 2017 | POSTER: Poor Man's URCU. | Pedro Ramalhete, Andreia Correia |
| 2017 | Optimizing the Four-Index Integral Transform Using Data Movement Lower Bounds Analysis. | Samyam Rajbhandari, Fabrice Rastello, Karol Kowalski, Sriram Krishnamoorthy, P. Sadayappan |
| 2017 | Simple, Accurate, Analytical Time Modeling and Optimal Tile Size Selection for GPGPU Stencils. | Nirmal Prajapati, Waruna Ranasinghe, Sanjay V. Rajopadhye, Rumen Andonov, Hristo N. Djidjev, Tobias Grosser |
| 2017 | Checking Concurrent Data Structures Under the C/C++11 Memory Model. | Peizhao Ou, Brian Demsky |
| 2017 | Parallel CCD++ on GPU for Matrix Factorization. | Israt Nisa, Aravind Sukumaran-Rajam, Rakshith Kunchum, P. Sadayappan |
| 2017 | POSTER: A GPU-Friendly Skiplist Algorithm. | Nurit Moscovici, Nachshon Cohen, Erez Petrank |
| 2017 | Function Call Re-Vectorization. | Rubens E. A. Moreira, Caroline Collange, Fernando Magno Quinto Pereira |
| 2017 | POSTER: Automated Load Balancer Selection Based on Application Characteristics. | Harshitha Menon, Kavitha Chandrasekar, Laxmikant V. Kal |
| 2017 | A Multicore Path to Connectomics-on-Demand. | Alexander Matveev, Yaron Meirovitch, Hayk Saribekyan, Wiktor Jakubiuk, Tim Kaler, Gergely dor, David M. Budden, Aleksandar Zlateski, Nir Shavit |
| 2017 | Thread Data Sharing in Cache: Theory and Measurement. | Hao Luo, Pengcheng Li, Chen Ding |
| 2017 | A Framework for Developing Parallel Applications with high level Tasks on Heterogeneous Platforms. | Chao Liu, Miriam Leeser |
| 2017 | High Performance Detection of Strongly Connected Components in Sparse Graphs on GPUs. | Pingfan Li, Xuhao Chen, Jie Shen, Jianbin Fang, Tao Tang, Canqun Yang |
| 2017 | Launch-Time Optimization of OpenCL GPU Kernels. | Andrew S. D. Lee, Tarek S. Abdelrahman |
| 2017 | A high-performance portable abstract interface for explicit SIMD vectorization. | Przemyslaw Karpinski, J. McDonald |
| 2017 | POSTER: MAPA: An Automatic Memory Access Pattern Analyzer for GPU Applications. | Gangwon Jo, Jaehoon Jung, Jiyoung Park, Jaejin Lee |
| 2017 | Grammar-aware Parallelization for Scalable XPath Querying. | Lin Jiang, Zhijia Zhao |
| 2017 | Combining SIMD and Many/Multi-core Parallelism for Finite State Machines with Enumerative Speculation. | Peng Jiang, Gagan Agrawal |
| 2017 | Towards Composable GPU Programming: Programming GPUs with Eager Actions and Lazy Views. | Michael Haidl, Michel Steuwer, Hendrik Dirks, Tim Humernbrum, Sergei Gorlatch |
| 2017 | High-performance Cholesky factorization for GPU-only execution. | Azzam Haidar, Ahmad Abdelfattah, Stanimire Tomov, Jack J. Dongarra |
| 2017 | POSTER: Distributed Control: The Benefits of Eliminating Global Synchronization via Effective Scheduling. | Jesun Sahariar Firoz, Thejaka Amila Kanewala, Marcin Zalewski, Martina Barnas, Andrew Lumsdaine |
| 2017 | DNNMark: A Deep Neural Network Benchmark Suite for GPUs. | Shi Dong, David R. Kaeli |
| 2017 | POSTER: IOGP: An Incremental Online Graph Partitioning for Large-Scale Distributed Graph Databases. | Dong Dai, Wei Zhang, Yong Chen |
| 2017 | Layout Lock: A Scalable Locking Paradigm for Concurrent Data Layout Modifications. | Nachshon Cohen, Arie Tal, Erez Petrank |