| 2017 | Optimizing recursive task parallel programs. | Suyash Gupta, Rahul Shrivastava, V. Krishna Nandivada |
| 2017 | Way-combining directory: an adaptive and scalable low-cost coherence directory. | J. Rubn Titos Gil, Antonio Flores, Ricardo Fernndez-Pascual, Alberto Ros, Manuel E. Acacio |
| 2017 | Automatic topology mapping of diverse large-scale parallel applications. | Juan J. Galvez, Nikhil Jain, Laxmikant V. Kal |
| 2017 | Dynamic scheduling for efficient hierarchical sparse matrix operations on the GPU. | Andreas Derler, Rhaleb Zayer, Hans-Peter Seidel, Markus Steinberger |
| 2017 | Revisiting phased transactional memory. | Joao P. L. de Carvalho, Guido Araujo, Alexandro Baldassin |
| 2017 | HiPA: history-based piecewise approximation for functions. | Aurangzeb, Rudolf Eigenmann |
| 2017 | Novel HPC techniques to batch execution of many variable size BLAS computations on GPUs. | Ahmad Abdelfattah, Azzam Haidar, Stanimire Tomov, Jack J. Dongarra |
| 2016 | GreenGear: Leveraging and Managing Server Heterogeneity for Improving Energy Efficiency in Green Data Centers. | Xu Zhou, Haoran Cai, Qiang Cao, Hong Jiang, Lei Tian, Changsheng Xie |
| 2016 | Efficient Timestamp-Based Cache Coherence Protocol for Many-Core Architectures. | Yuan Yao, Guanhua Wang, Zhiguo Ge, Tulika Mitra, Wenzhi Chen, Naxin Zhang |
| 2016 | Scheduling Tasks with Mixed Timing Constraints in GPU-Powered Real-Time Systems. | Yunlong Xu, Rui Wang, Tao Li, Mingcong Song, Lan Gao, Zhongzhi Luan, Depei Qian |
| 2016 | GCaR: Garbage Collection aware Cache Management with Improved Performance for Flash-based SSDs. | Suzhen Wu, Yanping Lin, Bo Mao, Hong Jiang |
| 2016 | BLASX: A High Performance Level-3 BLAS Library for Heterogeneous Multi-GPU Computing. | Linnan Wang, Wei Wu, Zenglin Xu, Jianxiong Xiao, Yi Yang |
| 2016 | Parallel Transposition of Sparse Data Structures. | Hao Wang, Weifeng Liu, Kaixi Hou, Wu-chun Feng |
| 2016 | Noise Aware Scheduling in Data Centers. | Hameedah Sultan, Arpit Katiyar, Smruti R. Sarangi |
| 2016 | AEQUITAS: Coordinated Energy Management Across Parallel Applications. | Haris Ribic, Yu David Liu |
| 2016 | Exploiting Dynamic Reuse Probability to Manage Shared Last-level Caches in CPU-GPU Heterogeneous Processors. | Siddharth Rai, Mainak Chaudhuri |
| 2016 | SReplay: Deterministic Sub-Group Replay for One-Sided Communication. | Xuehai Qian, Koushik Sen, Paul Hargrove, Costin Iancu |
| 2016 | Prefetching Techniques for Near-memory Throughput Processors. | Reena Panda, Yasuko Eckert, Nuwan Jayasena, Onur Kayiran, Michael Boyer, Lizy Kurian John |
| 2016 | Lynx: Using OS and Hardware Support for Fast Fine-Grained Inter-Core Communication. | Konstantina Mitropoulou, Vasileios Porpodas, Xiaochun Zhang, Timothy M. Jones |
| 2016 | TurboTiling: Leveraging Prefetching to Boost Performance of Tiled Codes. | Sanyam Mehta, Rajat Garg, Nishad Trivedi, Pen-Chung Yew |
| 2016 | DSMR: A Parallel Algorithm for Single-Source Shortest Path Problem. | Saeed Maleki, Donald Nguyen, Andrew Lenharth, Mara Jess Garzarn, David A. Padua, Keshav Pingali |
| 2016 | SARVAVID: A Domain Specific Language for Developing Scalable Computational Genomics Applications. | Kanak Mahadik, Christopher Wright, Jinyi Zhang, Milind Kulkarni, Saurabh Bagchi, Somali Chaterji |
| 2016 | Replichard: Towards Tradeoff between Consistency and Performance for Metadata. | Zhiying Li, Ruini Xue, Lixiang Ao |
| 2016 | Barrier-Aware Warp Scheduling for Throughput Processors. | Yuxi Liu, Zhibin Yu, Lieven Eeckhout, Vijay Janapa Reddi, Yingwei Luo, Xiaolin Wang, Zhenlin Wang, Cheng-Zhong Xu |
| 2016 | Towards an Adaptive Multi-Power-Source Datacenter. | Longjun Liu, Hongbin Sun, Chao Li, Yang Hu, Nanning Zheng, Tao Li |