| 2016 | Neural network transformation under hardware constraints. | Youhui Zhang, Yu Ji, Wenguang Chen, Yuan Xie |
| 2016 | Towards the design of fault-tolerant mixed-criticality systems on multicores. | Luyuan Zeng, Pengcheng Huang, Lothar Thiele |
| 2016 | RRAM based learning acceleration. | Yu Wang, Lixue Xia, Ming Cheng, Tianqi Tang, Boxun Li, Huazhong Yang |
| 2016 | Runtime management of adaptive MPSoCs for graceful degradation. | Stavros Tzilis, Ioannis Sourdis, Vasileios Vasilikos, Dimitrios Rodopoulos, Dimitrios Soudris |
| 2016 | LOCUS: low-power customizable many-core architecture for wearables. | Cheng Tan, Aditi Kulkarni Mohite, Vanchinathan Venkataramani, Manupa Karunaratne, Tulika Mitra, Li-Shiuan Peh |
| 2016 | Enabling OpenVX support in mW-scale parallel accelerators. | Giuseppe Tagliavini, Germain Haugou, Andrea Marongiu, Luca Benini |
| 2016 | D-PUF: an intrinsically reconfigurable DRAM PUF for device authentication in embedded systems. | Soubhagya Sutar, Arnab Raha, Vijay Raghunathan |
| 2016 | Matrix multiplication beyond auto-tuning: rewrite-based GPU code generation. | Michel Steuwer, Toomas Remmelg, Christophe Dubach |
| 2016 | Redesigning a tagless access buffer to require minimal ISA changes. | Carlos Sanchez, Peter Gavin, Daniel Moreau, Magnus Sjlander, David B. Whalley, Per Larsson-Edefors, Sally A. McKee |
| 2016 | On-the-fly load data value tracing in multicores. | Mounika Ponugoti, Amrish K. Tewar, Aleksandar Milenkovic |
| 2016 | ILP-based modulo scheduling for high-level synthesis. | Julian Oppermann, Andreas Koch, Melanie Reuter-Oppermann, Oliver Sinnen |
| 2016 | COMET: communication-optimised multi-threaded error-detection technique. | Konstantina Mitropoulou, Vasileios Porpodas, Timothy M. Jones |
| 2016 | Handling large data sets for high-performance embedded applications in heterogeneous systems-on-chip. | Paolo Mantovani, Emilio G. Cota, Christian Pilato, Giuseppe Di Guglielmo, Luca P. Carloni |
| 2016 | Speculative disassembly of binary code. | M. Ammar Ben Khadra, Dominik Stoffel, Wolfgang Kunz |
| 2016 | A real-time digital-microfluidic platform for epigenetics. | Mohamed Ibrahim, Craig Boswell, Krishnendu Chakrabarty, Kristin Scott, Miroslav Pajic |
| 2016 | CaffePresso: an optimized library for deep learning on embedded accelerator-based platforms. | Gopalakrishna Hegde, Siddhartha, Nachiappan Ramasamy, Nachiket Kapre |
| 2016 | A jump-target identification method for multi-architecture static binary translation. | Alessandro Di Federico, Giovanni Agosta |
| 2016 | Hybrid network-on-chip architectures for accelerating deep learning kernels on heterogeneous manycore platforms. | Wonje Choi, Karthi Duraisamy, Ryan Gary Kim, Janardhan Rao Doppa, Partha Pratim Pande, Radu Marculescu, Diana Marculescu |
| 2016 | Thrifty-malloc: A HW/SW codesign for the dynamic management of hardware transactional memory in embedded multicore systems. | Thomas Carle, Dimitra Papagiannopoulou, Tali Moreshet, Andrea Marongiu, Maurice Herlihy, R. Iris Bahar |
| 2016 | Power and thermal management in massive multicore chips: theoretical foundation meets architectural innovation and resource allocation. | Paul Bogdan, Partha Pratim Pande, Hussam Amrouch, Muhammad Shafique, Jrg Henkel |
| 2016 | FastCollect: offloading generational garbage collection to integrated GPUs. | Abhinav, Rupesh Nasre |
| 2015 | Vector-aware register allocation for GPU shader processors. | Yi-Ping You, Szu-Chieh Chen |
| 2015 | Timing characterization of OpenMP4 tasking model. | Maria A. Serrano, Alessandra Melani, Roberto Vargas, Andrea Marongiu, Marko Bertogna, Eduardo Quiones |
| 2015 | Optimizing mobile display brightness by leveraging human visual perception. | Matthew Schuchhardt, Susmit Jha, Raid Ayoub, Michael Kishinevsky, Gokhan Memik |
| 2015 | Exploiting cache conflicts to reduce radiation sensitivity of operating systems on embedded systems. | Thiago Santini, Paolo Rech, Luigi Carro, Flvio Rech Wagner |