| 2009 | Barcelona OpenMP Tasks Suite: A Set of Benchmarks Targeting the Exploitation of Task Parallelism in OpenMP. | Alejandro Duran, Xavier Teruel, Roger Ferrer, Xavier Martorell, Eduard Ayguad |
| 2009 | Improving Resource Availability by Relaxing Network Allocation Constraints on Blue Gene/P. | Narayan Desai, Darius Buntinas, Daniel Buettner, Pavan Balaji, Anthony Chan |
| 2009 | Exploiting Simulation Slack to Improve Parallel Simulation Speed. | Jianwei Chen, Murali Annavaram, Michel Dubois |
| 2009 | End-User Diagnosis of Communication Paths in Sensor Network Systems. | Qing Cao, Dong Wang, Tarek F. Abdelzaher |
| 2009 | A Heuristic for Mapping Virtual Machines and Links in Emulation Testbeds. | Rodrigo N. Calheiros, Rajkumar Buyya, Csar A. F. De Rose |
| 2009 | Cache-Efficient, Intranode, Large-Message MPI Communication with MPICH2-Nemesis. | Darius Buntinas, Brice Goglin, David Goodell, Guillaume Mercier, Stphanie Moreaud |
| 2009 | Parallel Phase Model: A Programming Model for High-end Parallel Machines with Manycores. | Ron Brightwell, Mike Heroux, Zhaofang Wen, Junfeng Wu |
| 2009 | Integrated Performance Views in Charm++: Projections Meets TAU. | Scott Biersdorff, Chee Wai Lee, Allen D. Malony, Laxmikant V. Kal |
| 2009 | Optimizing the Latency of Streaming Applications under Throughput and Reliability Constraints. | Anne Benoit, Mourad Hakem, Yves Robert |
| 2009 | Computing the Throughput of Replicated Workflows on Heterogeneous Platforms. | Anne Benoit, Matthieu Gallet, Bruno Gaujal, Yves Robert |
| 2009 | Speeding Up Distributed MapReduce Applications Using Hardware Accelerators. | Yolanda Becerra, Vicen Beltran, David Carrera, Marc Gonzlez, Jordi Torres, Eduard Ayguad |
| 2009 | Accelerating Lattice Boltzmann Fluid Flow Simulations Using Graphics Processors. | Peter Bailey, Joe Myre, Stuart D. C. Walsh, David J. Lilja, Martin O. Saar |
| 2009 | Direct N-body Kernels for Multicore Platforms. | Nitin Arora, Aashay Shringarpure, Richard W. Vuduc |
| 2009 | Performance Characterization of a Hierarchical MPI Implementation on Large-scale Distributed-memory Platforms. | Sadaf R. Alam, Richard F. Barrett, Jeffery A. Kuehn, Steve Poole |
| 2008 | Resource Allocation for Distributed Streaming Applications. | Qian Zhu, Gagan Agrawal |
| 2008 | Bandwidth-Efficient Continuous Query Processing over DHTs. | Yingwu Zhu |
| 2008 | Thermal Management for 3D Processors via Task Scheduling. | Xiuyi Zhou, Yi Xu, Yu Du, Youtao Zhang, Jun Yang |
| 2008 | Memory Access Scheduling Schemes for Systems with Multi-Core Processors. | Hongzhong Zheng, Jiang Lin, Zhao Zhang, Zhichun Zhu |
| 2008 | Maotai: View-Oriented Parallel Programming on CMT Processors. | Jiaqi Zhang, Zhiyi Huang, Wenguang Chen, Qihang Huang, Weimin Zheng |
| 2008 | ParColl: Partitioned Collective I/O on the Cray XT. | Weikuan Yu, Jeffrey S. Vetter |
| 2008 | Impacts of Indirect Blocks on Buffer Cache Energy Efficiency. | Jianhui Yue, Yifeng Zhu, Zhao Cai |
| 2008 | A Replication Overlay Assisted Resource Discovery Service for Federated Systems. | Hao Yang, Fan Ye, Zhen Liu |
| 2008 | The Content Pollution in Peer-to-Peer Live Streaming Systems: Analysis and Implications. | Sirui Yang, Hai Jin, Bo Li, Xiaofei Liao, Hong Yao, Xuping Tu |
| 2008 | Deadlock-Free Fully Adaptive Routing in Tori Based on a New Virtual Network Partitioning Scheme. | Dong Xiang, Qi Wang, Yi Pan |
| 2008 | A Distributed Context-Free Language Constrained Shortest Path Algorithm. | Charles B. Ward, Nathan M. Wiegand, Phillip G. Bradford |