| 2017 | POSTER: State Teleportation via Hardware Transactional Memory. | Nachshon Cohen, Maurice Herlihy, Erez Petrank, Elias Wald |
| 2017 | POSTER: Provably Efficient Scheduling of Cache-Oblivious Wavefront Algorithms. | Rezaul Chowdhury, Pramod Ganapathi, Yuan Tang, Jesmin Jahan Tithi |
| 2017 | EffiSha: A Software Framework for Enabling Effficient Preemptive Scheduling of GPU. | Guoyang Chen, Yue Zhao, Xipeng Shen, Huiyang Zhou |
| 2017 | An Efficient Abortable-locking Protocol for Multi-level NUMA Systems. | Milind Chabbi, Abdelhalim Amer, Shasha Wen, Xu Liu |
| 2017 | TaskInsight: Understanding Task Schedules Effects on Memory and Performance. | Germn Ceballos, Thomas Grass, Andra Hugo, David Black-Schaffer |
| 2017 | POSTER: HythTM: Extending the Applicability of Intel TSX Hardware Transactional Support. | Arnamoy Bhattacharyya, Mike Dai Wang, Mihai Burcea, Yi Ding, Allen Deng, Sai Varikooty, Shafaaf Hossain, Cristiana Amza |
| 2017 | Groute: An Asynchronous Multi-GPU Programming Model for Irregular Computations. | Tal Ben-Nun, Michael Sutton, Sreepathi Pai, Keshav Pingali |
| 2017 | Synchronized-by-Default Concurrency for Shared-Memory Systems. | Martin Bttig, Thomas R. Gross |
| 2017 | KiWi: A Key-Value Map for Scalable Real-Time Analytics. | Dmitry Basin, Edward Bortnikov, Anastasia Braginsky, Guy Golan-Gueta, Eshcar Hillel, Idit Keidar, Moshe Sulamy |
| 2017 | POSTER: An Architecture and Programming Model for Accelerating Parallel Commutative Computations via Privatization. | Vignesh Balaji, Dhruva Tirumala, Brandon Lucia |
| 2017 | S-Caffe: Co-designing MPI Runtimes and Caffe for Scalable Deep Learning on Modern GPU Clusters. | Ammar Ahmad Awan, Khaled Hamidouche, Jahanzeb Maqbool Hashmi, Dhabaleswar K. Panda |
| 2017 | POSTER: Reuse, don't Recycle: Transforming Algorithms that Throw Away Descriptors. | Maya Arbel-Raviv, Trevor Brown |
| 2017 | Batched Gauss-Jordan Elimination for Block-Jacobi Preconditioner Generation on GPUs. | Hartwig Anzt, Jack J. Dongarra, Goran Flegar, Enrique S. Quintana-Ort |
| 2017 | PETRAS: Performance, Energy and Thermal Aware Resource Allocation and Scheduling for Heterogeneous Systems. | Shouq Alsubaihi, Jean-Luc Gaudiot |
| 2017 | Reduction to Tridiagonal Form for Symmetric Eigenproblems on Asymmetric Multicore Processors. | Pedro Alonso, Sandra Cataln, Jos R. Herrero, Enrique S. Quintana-Ort, Rafael Rodrguez-Snchez |
| 2017 | Contention in Structured Concurrency: Provably Efficient Dynamic Non-Zero Indicators for Nested Parallelism. | Umut A. Acar, Naama Ben-David, Mike Rainey |
| 2017 | POSTER: Cache-Oblivious MPI All-to-All Communications on Many-Core Architectures. | Shigang Li, Yunquan Zhang, Torsten Hoefler |
| 2016 | General-purpose join algorithms for large graph triangle listing on heterogeneous systems. | Daniel Zinn, Haicheng Wu, Jin Wang, Molham Aref, Sudhakar Yalamanchili |
| 2016 | Scalable adaptive NUMA-aware lock: combining local locking and remote locking for efficient concurrency. | Mingzhe Zhang, Francis C. M. Lau, Cho-Li Wang, Luwei Cheng, Haibo Chen |
| 2016 | GPUpIO: the case for I/O-driven preemption on GPUs. | Lior Zeno, Avi Mendelson, Mark Silberstein |
| 2016 | A wait-free queue as fast as fetch-and-add. | Chaoran Yang, John M. Mellor-Crummey |
| 2016 | High performance model based image reconstruction. | Xiao Wang, Amit Sabne, Sherman J. Kisner, Anand Raghunathan, Charles A. Bouman, Samuel P. Midkiff |
| 2016 | Gunrock: a high-performance graph processing library on the GPU. | Yangzihao Wang, Andrew A. Davidson, Yuechao Pan, Yuduo Wu, Andy Riffel, John D. Owens |
| 2016 | Be my guest: MCS lock now welcomes guests. | Tianzheng Wang, Milind Chabbi, Hideaki Kimura |
| 2016 | Compilers, hands-off my hands-on optimizations. | Richard Veras, Doru-Thom Popovici, Tze Meng Low, Franz Franchetti |