| 2016 | BlackBox: lightweight security monitoring for COTS binaries. | Byron Hawkins, Brian Demsky, Michael B. Taylor |
| 2016 | Visual Instance Inlining and Specialization: Building Domain-Specific Diagrams from Reusable Types. | Niklas Fors, Grel Hedin |
| 2016 | Flexible on-stack replacement in LLVM. | Daniele Cono D'Elia, Camil Demetrescu |
| 2016 | AutoFDO: automatic feedback-directed optimization for warehouse-scale applications. | Dehao Chen, Xinliang David Li, Tipp Moseley |
| 2016 | Exploiting recent SIMD architectural advances for irregular applications. | Linchuan Chen, Peng Jiang, Gagan Agrawal |
| 2016 | Validating optimizations of concurrent C/C++ programs. | Soham Chakraborty, Viktor Vafeiadis |
| 2016 | Have abstraction and eat performance, too: optimized heterogeneous computing with parallel patterns. | Kevin J. Brown, HyoukJoong Lee, Tiark Rompf, Arvind K. Sujeeth, Christopher De Sa, Christopher R. Aberger, Kunle Olukotun |
| 2016 | A black-box approach to energy-aware scheduling on integrated CPU-GPU systems. | Rajkishore Barik, Naila Farooqui, Brian T. Lewis, Chunling Hu, Tatiana Shpeisman |
| 2016 | Opening polyhedral compiler's black box. | Lnac Bagnres, Oleksandr Zinenko, Stphane Huot, Cdric Bastoul |
| 2016 | Trace-based affine reconstruction of codes. | Gabriel Rodrguez, Jos M. Andin, Mahmut T. Kandemir, Juan Tourio |
| 2015 | On performance debugging of unnecessary lock contentions on multicore processors: a replay-based approach. | Long Zheng, Xiaofei Liao, Bingsheng He, Song Wu, Hai Jin |
| 2015 | Dependence-Based Code Transformation for Coarse-Grained Parallelism. | Bo Zhao, Zhen Li, Ali Jannesari, Felix Wolf, Weiguo Wu |
| 2015 | HERMES: a fast cross-ISA binary translator with post-optimization. | Xiaochun Zhang, Qi Guo, Yunji Chen, Tianshi Chen, Weiwu Hu |
| 2015 | A Roadmap for a Type Architecture Based Parallel Programming Language. | Muhammad Nur Yanhaona, Andrew S. Grimshaw |
| 2015 | An Evaluation of Memory Sharing Performance for Heterogeneous Embedded SoCs with Many-Core Accelerators. | Pirmin Vogel, Andrea Marongiu, Luca Benini |
| 2015 | Optimizing and auto-tuning scale-free sparse matrix-vector multiplication on Intel Xeon Phi. | Wai Teng Tang, Ruizhe Zhao, Mian Lu, Yun Liang, Huynh Phung Huyng, Xibai Li, Rick Siow Mong Goh |
| 2015 | MemorySanitizer: fast detector of uninitialized memory use in C++. | Evgeniy Stepanov, Konstantin Serebryany |
| 2015 | Reactive tiling. | Jithendra Srinivas, Wei Ding, Mahmut T. Kandemir |
| 2015 | Locality aware concurrent start for stencil applications. | Sunil Shrestha, Guang R. Gao, Joseph B. Manzano, Andrs Mrquez, John Feo |
| 2015 | Branch prediction and the performance of interpreters: don't trust folklore. | Erven Rohou, Bharath Narasimha Swamy, Andr Seznec |
| 2015 | PSLP: padded SLP automatic vectorization. | Vasileios Porpodas, Alberto Magni, Timothy M. Jones |
| 2015 | Optimizing the flash-RAM energy trade-off in deeply embedded systems. | James Pallister, Kerstin Eder, Simon J. Hollis |
| 2015 | Snapshot-based loading-time acceleration for web applications. | JinSeok Oh, Soo-Mook Moon |
| 2015 | Scalable conditional induction variables (CIV) analysis. | Cosmin E. Oancea, Lawrence Rauchwerger |
| 2015 | Approximating flow-sensitive pointer analysis using frequent itemset mining. | Vaivaswatha Nagaraj, R. Govindarajan |