| 2016 | Throughput-Optimized OpenCL-based FPGA Accelerator for Large-Scale Convolutional Neural Networks. | Naveen Suda, Vikas Chandra, Ganesh Dasika, Abinash Mohanty, Yufei Ma, Sarma B. K. Vrudhula, Jae-sun Seo, Yu Cao |
| 2016 | OLAF'16: Second International Workshop on Overlay Architectures for FPGAs. | Hayden Kwok-Hay So, John Wawrzynek |
| 2016 | Low-Swing Signaling for FPGA Power Reduction (Abstract Only). | Sayeh Sharifymoghaddam, Ali Sheikholeslami |
| 2016 | Spatial Debug & Debug Without Re-programming in FPGAs: On-Chip debugging in FPGAs. | Pankaj Shanker |
| 2016 | Optimal Circuits for Streamed Linear Permutations Using RAM. | Franois Serre, Thomas Holenstein, Markus Pschel |
| 2016 | A Case for Work-stealing on FPGAs with OpenCL Atomics. | Nadesh Ramanathan, John Wickerson, Felix Winterstein, George A. Constantinides |
| 2016 | Going Deeper with Embedded FPGA Platform for Convolutional Neural Network. | Jiantao Qiu, Jie Wang, Song Yao, Kaiyuan Guo, Boxun Li, Erjin Zhou, Jincheng Yu, Tianqi Tang, Ningyi Xu, Sen Song, Yu Wang, Huazhong Yang |
| 2016 | ENFIRE: An Energy-efficient Fine-grained Spatio-temporal Reconfigurable Computing Fabric (Abstact Only). | Wenchao Qian, Christopher Babecki, Robert Karam, Swarup Bhunia |
| 2016 | An FPGA-SOC Based Accelerating Solution for N-body Simulations in MOND (Abstract Only). | Bo Peng, Tianqi Wang, Xi Jin, Chuanjun Wang |
| 2016 | GraphOps: A Dataflow Library for Graph Analytics Acceleration. | Tayo Oguntebi, Kunle Olukotun |
| 2016 | PRFloor: An Automatic Floorplanner for Partially Reconfigurable FPGA Systems. | Tuan D. A. Nguyen, Akash Kumar |
| 2016 | FCUDA-SoC: Platform Integration for Field-Programmable SoC with the CUDA-to-FPGA Compiler. | Tan Nguyen, Swathi T. Gurumani, Kyle Rupnow, Deming Chen |
| 2016 | Resolve: Generation of High-Performance Sorting Architectures from High-Level Synthesis. | Janarbek Matai, Dustin Richmond, Dajung Lee, Zac Blair, Qiongzhi Wu, Amin Abazari, Ryan Kastner |
| 2016 | Just In Time Assembly of Accelerators. | Sen Ma, Zeyad Aklah, David Andrews |
| 2016 | High Level Synthesis of Complex Applications: An H.264 Video Decoder. | Xinheng Liu, Yao Chen, Tan Nguyen, Swathi T. Gurumani, Kyle Rupnow, Deming Chen |
| 2016 | Pitfalls and Tradeoffs in Simultaneous, On-Chip FPGA Delay Measurement. | Timothy A. Linscott, Benjamin Gojman, Raphael Rubin, Andr DeHon |
| 2016 | Using Stochastic Computing to Reduce the Hardware Requirements for a Restricted Boltzmann Machine Classifier. | Bingzhe Li, M. Hassan Najafi, David J. Lilja |
| 2016 | The Stratix™ 10 Highly Pipelined FPGA Architecture. | David M. Lewis, Gordon R. Chiu, Jeffrey Chromczak, David R. Galloway, Ben Gamsa, Valavan Manohararajah, Ian Milton, Tim Vanderhoek, John Van Dyken |
| 2016 | A High-throughput Architecture for Lossless Decompression on FPGA Designed Using HLS (Abstract Only). | Jie Lei, Yu-Ting Chen, Yunsong Li, Jason Cong |
| 2016 | Doubling FPGA Throughput via a Soft SerDes Architecture for Full-Bandwidth Serial Pipelining (Abstract Only). | Aaron Landy, Greg Stitt |
| 2016 | Re-targeting Optimization Sequences from Scalar Processors to FPGAs in HLS compilers (Abstract Only). | Ronak Kogta, Suresh Purini, Ajit Mathew |
| 2016 | SDSoC: A Higher-level Programming Environment for Zynq SoC and Ultrascale+ MPSoC. | Vinod Kathail, James Hwang, Welson Sun, Yogesh Chobe, Tom Shui, Jorge Carrillo |
| 2016 | GPU-Accelerated High-Level Synthesis for Bitwidth Optimization of FPGA Datapaths. | Nachiket Kapre, Deheng Ye |
| 2016 | FGPU: An SIMT-Architecture for FPGAs. | Muhammed Al Kadi, Benedikt Janen, Michael Hbner |
| 2016 | A Low DDR Bandwidth 100FPS 1080p Video 2D Discrete Wavelet Transform Implementation on FPGA (Abstract Only). | Mohammed Shaaban Ibraheem, Syed Zahid Ahmed, Khalil Hachicha, Sylvain Hochberg, Patrick Garda |