| 2017 | Performance Analysis and Optimization of Sparse Matrix-Vector Multiplication on Modern Multi- and Many-Core Processors. | Athena Elafrou, Georgios I. Goumas, Nectarios Koziris |
| 2017 | GCN: GPU-Based Cube CNN Framework for Hyperspectral Image Classification. | Han Dong, Tao Li, Jiabing Leng, Lingyan Kong, Gang Bai |
| 2017 | A Parallel TSP-Based Algorithm for Balanced Graph Partitioning. | Harshvardhan Das, Subodh Kumar |
| 2017 | Scalable Write Allocation in the WAFL File System. | Matthew Curtis-Maury, Ram Kesavan, Mrinal K. Bhattacharjee |
| 2017 | An Efficient, Distributed Stochastic Gradient Descent Algorithm for Deep-Learning Applications. | Guojing Cong, Onkar Bhardwaj, Minwei Feng |
| 2017 | Efficient and Scalable Multi-Source Streaming Broadcast on GPU Clusters for Deep Learning. | Ching-Hsiang Chu, Xiaoyi Lu, Ammar Ahmad Awan, Hari Subramoni, Jahanzeb Maqbool Hashmi, Bracy Elton, Dhabaleswar K. Panda |
| 2017 | A Coflow-Based Co-Optimization Framework for High-Performance Data Analytics. | Long Cheng, Ying Wang, Yulong Pei, Dick H. J. Epema |
| 2017 | A Pareto Framework for Data Analytics on Heterogeneous Systems: Implications for Green Energy Usage and Performance. | Aniket Chakrabarti, Srinivasan Parthasarathy, Christopher Stewart |
| 2017 | GLTO: On the Adequacy of Lightweight Thread Approaches for OpenMP Implementations. | Adrin Castell, Sangmin Seo, Rafael Mayo, Pavan Balaji, Enrique S. Quintana-Ort, Antonio J. Pea |
| 2017 | Exploiting GPUs for Fast Force-Directed Visualization of Large-Scale Networks. | Govert G. Brinkmann, Kristian F. D. Rietveld, Frank W. Takes |
| 2017 | Overlapping Data Transfers with Computation on GPU with Tiles. | Burak Bastem, Didem Unat, Weiqun Zhang, Ann S. Almgren, John Shalf |
| 2017 | High Performance Query Processing for Web Scale RDF Data using BSP Style Communication and Balanced Distribution. | Minho Bae, Junho Eum, Donghoon Kim, Sangyoon Oh |
| 2017 | High-Performance Recommender System Training Using Co-Clustering on CPU/GPU Clusters. | Kubilay Atasu, Thomas P. Parnell, Celestine Dnner, Michail Vlachos, Haralampos Pozidis |
| 2017 | A Machine Learning Approach for Efficient Parallel Simulation of Beam Dynamics on GPUs. | Kamesh Arumugam, Desh Ranjan, Mohammad Zubair, Balsa Terzic, Alexander N. Godunov, Tunazzina Islam |
| 2017 | Variable-Size Batched LU for Small Matrices and Its Integration into Block-Jacobi Preconditioning. | Hartwig Anzt, Jack J. Dongarra, Goran Flegar, Enrique S. Quintana-Ort |
| 2017 | WA-Dataspaces: Exploring the Data Staging Abstractions for Wide-Area Distributed Scientific Workflows. | Mehmet Fatih Aktas, Javier Diaz Montes, Ivan Rodero, Manish Parashar |
| 2017 | Application-Aware Power Coordination on Power Bounded NUMA Multicore Systems. | Rong Ge, Pengfei Zou, Xizhou Feng |
| 2016 | MIC: An Efficient Anonymous Communication System in Data Center Networks. | Tingwei Zhu, Dan Feng, Yu Hua, Fang Wang, Qingyu Shi, Jiahao Liu |
| 2016 | TECH: A Thermal-Aware and Cost Efficient Mechanism for Colocation Demand Response. | Ziqi Zhao, Fan Wu, Shaolei Ren, Xiaofeng Gao, Guihai Chen, Yong Cui |
| 2016 | High Performance MPI Library for Container-Based HPC Cloud on InfiniBand Clusters. | Jie Zhang, Xiaoyi Lu, Dhabaleswar K. Panda |
| 2016 | RMD: A Resemblance and Mergence Based Approach for High Performance Deduplication. | Panfeng Zhang, Ping Huang, Xubin He, Hua Wang, Lingyu Yan, Ke Zhou |
| 2016 | RegTT: Accelerating Tree Traversals on GPUs by Exploiting Regularities. | Feng Zhang, Peng Di, Hao Zhou, Xiangke Liao, Jingling Xue |
| 2016 | The Future(s) of Transactional Memory. | Jingna Zeng, Joo Pedro Barreto, Seif Haridi, Lus E. T. Rodrigues, Paolo Romano |
| 2016 | Thread Similarity Matrix: Visualizing Branch Divergence in GPGPU Programs. | Zhibin Yu, Lieven Eeckhout, Cheng-Zhong Xu |
| 2016 | HppCnn: A High-Performance, Portable Deep-Learning Library for GPGPUs. | Yi Yang, Min Feng, Srimat T. Chakradhar |