| 2018 | Interference between I/O and MPI Traffic on Fat-tree Networks. | Kevin A. Brown, Nikhil Jain, Satoshi Matsuoka, Martin Schulz, Abhinav Bhatele |
| 2018 | CSTF: Large-Scale Sparse Tensor Factorizations on Distributed Platforms. | Zachary Blanco, Bangtian Liu, Maryam Mehri Dehnavi |
| 2018 | A Performance Model to Execute Workflows on High-Bandwidth-Memory Architectures. | Anne Benoit, Swann Perarnau, Loc Pottier, Yves Robert |
| 2018 | A Communication-Efficient Causal Broadcast Protocol. | Joo Paulo de Araujo, Luciana Arantes, Elias P. Duarte Jr., Luiz A. Rodrigues, Pierre Sens |
| 2018 | An Empirical Comparison of k-Shortest Simple Path Algorithms on Multicores. | Deepak Ajwani, Erika Duriakova, Neil Hurley, Ulrich Meyer, Alexander Schickedanz |
| 2018 | Massively Scaling the Metal Microscopic Damage Simulation on Sunway TaihuLight Supercomputer. | Shigang Li, Baodong Wu, Yunquan Zhang, Xianmeng Wang, Jianjiang Li, Changjun Hu, Jue Wang, Yangde Feng, Ningming Nie |
| 2018 | Dual-Paradigm Stream Processing. | Song Wu, Zhiyi Liu, Shadi Ibrahim, Lin Gu, Hai Jin, Fei Chen |
| 2017 | Practical Experience with Transactional Lock Elision. | Tingzhe Zhou, Pantea Zardoshti, Michael F. Spear |
| 2017 | The Cloud as an OpenMP Offloading Device. | Herv Yviquel, Guido Araujo |
| 2017 | Runtime Data Layout Scheduling for Machine Learning Dataset. | Yang You, James Demmel |
| 2017 | Order/Radix Problem: Towards Low End-to-End Latency Interconnection Networks. | Ryota Yasudo, Michihiro Koibuchi, Koji Nakano, Hiroki Matsutani, Hideharu Amano |
| 2017 | Network Aware Multi-User Computation Partitioning in Mobile Edge Clouds. | Lei Yang, Jiannong Cao, Zhenyu Wang, Weigang Wu |
| 2017 | Bitslice Vectors: A Software Approach to Customizable Data Precision on Processors with SIMD Extensions. | Shixiong Xu, David Gregg |
| 2017 | Non-Sequential Striping for Distributed Storage Systems with Different Redundancy Schemes. | Yanwen Xie, Dan Feng, Fang Wang |
| 2017 | Large-Scale Parallelization of Smoothed Particle Hydrodynamics Method on Heterogeneous Cluster. | Yingrui Wang, Leisheng Li, Rong Tian |
| 2017 | Data Caching in Next Generation Mobile Cloud Services, Online vs. Off-Line. | Yang Wang, Shuibing He, Xiaopeng Fan, Chengzhong Xu, Joseph C. Culberson, Joseph Horton |
| 2017 | Boosting the Efficiency of HPCG and Graph500 with Near-Data Processing. | Erik Vermij, Leandro Fiorin, Christoph Hagleitner, Koen Bertels |
| 2017 | MPI-GDS: High Performance MPI Designs with GPUDirect-aSync for CPU-GPU Control Flow Decoupling. | Akshay Venkatesh, Khaled Hamidouche, Sreeram Potluri, Davide Rossetti, Ching-Hsiang Chu, Dhabaleswar K. Panda |
| 2017 | Greed Is Good: Parallel Algorithms for Bipartite-Graph Partial Coloring on Multicore Architectures. | Mustafa Kemal Tas, Kamer Kaya, Erik Saule |
| 2017 | Accelerating Graph Analytics by Utilising the Memory Locality of Graph Partitioning. | Jiawen Sun, Hans Vandierendonck, Dimitrios S. Nikolopoulos |
| 2017 | Parallel Algorithms for the Computation of Cycles in Relative Neighborhood Graphs. | Hari Sundar, Parmeshwar Khurd |
| 2017 | Predicting Response Latency Percentiles for Cloud Object Storage Systems. | Yi Su, Dan Feng, Yu Hua, Zhan Shi |
| 2017 | Multiple Pattern Matching for Network Security Applications: Acceleration through Vectorization. | Charalampos Stylianopoulos, Magnus Almgren, Olaf Landsiedel, Marina Papatriantafilou |
| 2017 | Constrained Tensor Factorization with Accelerated AO-ADMM. | Shaden Smith, Alec Beri, George Karypis |
| 2017 | Parallel Reconstruction of Three Dimensional Magnetohydrodynamic Equilibria in Plasma Confinement Devices. | Sudip K. Seal, Mark R. Cianciosa, Steven P. Hirshman, Andreas Wingen, Robert S. Wilcox, Ezekial A. Unterberg |