| 2018 | Extending ILUPACK with a Task-Parallel Version of BiCG for Dual-GPU Servers. | Jos Ignacio Aliaga, Matthias Bollhfer, Ernesto Dufrechou, Pablo Ezzatti, Enrique S. Quintana-Ort |
| 2018 | Performance challenges in modular parallel programs. | Umut A. Acar, Vitaly Aksenov, Arthur Charguraud, Mike Rainey |
| 2018 | Griffin: uniting CPU and GPU in information retrieval systems for intra-query parallelism. | Yang Liu, Jianguo Wang, Steven Swanson |
| 2017 | Understanding The Security of Discrete GPUs. | Zhiting Zhu, Sangman Kim, Yuri Rozhanski, Yige Hu, Emmett Witchel, Mark Silberstein |
| 2017 | POSTER: An Infrastructure for HPC Knowledge Sharing and Reuse. | Yue Zhao, Chunhua Liao, Xipeng Shen |
| 2017 | Understanding the GPU Microarchitecture to Achieve Bare-Metal Performance Tuning. | Xiuxia Zhang, Guangming Tan, Shuangbai Xue, Jiajia Li, Keren Zhou, Mingyu Chen |
| 2017 | POSTER: On the Problem of Consistency Exceptions in the Context of Strong Memory Models. | Minjia Zhang, Swarnendu Biswas, Michael D. Bond |
| 2017 | Efficient Convex Optimization on GPUs for Embedded Model Predictive Control. | Leiming Yu, Abraham Goldsmith, Stefano Di Cairano |
| 2017 | Pagoda: Fine-Grained GPU Resource Virtualization for Narrow Tasks. | Tsung Tai Yeh, Amit Sabne, Putt Sakdhnagool, Rudolf Eigenmann, Timothy G. Rogers |
| 2017 | POSTER: Recovering Performance for Vector-based Machine Learning on Managed Runtime. | Mingyu Wu, Haibing Guan, Binyu Zang, Haibo Chen |
| 2017 | Silent Data Corruption Resilient Two-sided Matrix Factorizations. | Panruo Wu, Nathan DeBardeleben, Qiang Guan, Sean Blanchard, Jieyang Chen, Dingwen Tao, Xin Liang, Kaiming Ouyang, Zizhong Chen |
| 2017 | Merge or Separate?: Multi-job Scheduling for OpenCL Kernels on CPU/GPU Platforms. | Yuan Wen, Michael F. P. O'Boyle |
| 2017 | Eunomia: Scaling Concurrent Search Trees under Contention Using HTM. | Xin Wang, Weihua Zhang, Zhaoguo Wang, Ziyun Wei, Haibo Chen, Wenyun Zhao |
| 2017 | SC-Haskell: Sequential Consistency in Languages That Minimize Mutable Shared Heap. | Michael Vollmer, Ryan G. Scott, Madanlal Musuvathi, Ryan R. Newton |
| 2017 | Directive-based tile abstraction to distribute loops on accelerators. | Tristan Vanderbruggen, John Cavazos, Chunhua Liao, Daniel J. Quinlan |
| 2017 | Processor-Oblivious Record and Replay. | Robert Utterback, Kunal Agrawal, I-Ting Angelina Lee, Milind Kulkarni |
| 2017 | Self-Checkpoint: An In-Memory Checkpoint Method Using Less Space and Its Practice on Fault-Tolerant HPL. | Xiongchao Tang, Jidong Zhai, Bowen Yu, Wenguang Chen, Weimin Zheng |
| 2017 | POSTER: STAR (Space-Time Adaptive and Reductive) Algorithms for Real-World Space-Time Optimality. | Yuan Tang, Ronghui You |
| 2017 | Using Butterfly-Patterned Partial Sums to Draw from Discrete Distributions. | Guy L. Steele Jr., Jean-Baptiste Tristan |
| 2017 | It's Time for a New Old Language. | Guy L. Steele Jr. |
| 2017 | Isoefficiency in Practice: Configuring and Understanding the Performance of Task-based Applications. | Sergei Shudler, Alexandru Calotoiu, Torsten Hoefler, Felix Wolf |
| 2017 | Assessing One-to-One Parallelism Levels Mapping for OpenMP Offloading to GPUs. | Chen Shen, Xiaonan Tian, Dounia Khaldi, Barbara M. Chapman |
| 2017 | Tapir: Embedding Fork-Join Parallelism into LLVM's Intermediate Representation. | Tao B. Schardl, William S. Moses, Charles E. Leiserson |
| 2017 | Noise Injection Techniques to Expose Subtle and Unintended Message Races. | Kento Sato, Dong H. Ahn, Ignacio Laguna, Gregory L. Lee, Martin Schulz, Christopher M. Chambreau |
| 2017 | Model-based Iterative CT Image Reconstruction on GPUs. | Amit Sabne, Xiao Wang, Sherman J. Kisner, Charles A. Bouman, Anand Raghunathan, Samuel P. Midkiff |