| 2009 | Auto-vectorization through code generation for stream processing applications. | Huayong Wang, Henrique Andrade, Bugra Gedik, Kun-Lung Wu |
| 2009 | Tuned and wildly asynchronous stencil kernels for hybrid CPU/GPU systems. | Sundaresan Venkatasubramanian, Richard W. Vuduc |
| 2009 | A european perspective on supercomputing. | Mateo Valero |
| 2009 | Single-particle 3d reconstruction from cryo-electron microscopy images on GPU. | Guangming Tan, Ziyu Guo, Mingyu Chen, Dan Meng |
| 2009 | Maximizing MPI point-to-point communication performance on RDMA-enabled clusters with customized protocols. | Matthew Small, Xin Yuan |
| 2009 | Prediction-based power estimation and scheduling for CMPs. | Karan Singh, Major Bhadauria, Sally A. McKee |
| 2009 | Refereeing conflicts in hardware transactional memory. | Arrvindh Shriraman, Sandhya Dwarkadas |
| 2009 | Chunking parallel loops in the presence of synchronization. | Jun Shirako, Jisheng M. Zhao, V. Krishna Nandivada, Vivek Sarkar |
| 2009 | FTL design exploration in reconfigurable high-performance SSD for server applications. | Ji-Yong Shin, Zenglin Xia, Ning-Yi Xu, Rui Gao, Xiongfei Cai, Seungryoul Maeng, Feng-Hsiung Hsu |
| 2009 | High-performance regular expression scanning on the Cell/B.E. processor. | Daniele Paolo Scarpazza, Gregory F. Russell |
| 2009 | Adagio: making DVS practical for complex HPC applications. | Barry Rountree, David K. Lowenthal, Bronis R. de Supinski, Martin Schulz, Vincent W. Freeh, Tyler K. Bletsch |
| 2009 | Exploring pattern-aware routing in generalized fat tree networks. | Germn Rodrguez, Ramn Beivide, Cyriel Minkenberg, Jess Labarta, Mateo Valero |
| 2009 | Fast and scalable list ranking on the GPU. | M. Suhail Rehman, Kishore Kothapalli, P. J. Narayanan |
| 2009 | Creating artificial global history to improve branch prediction accuracy. | Leo Porter, Dean M. Tullsen |
| 2009 | TransMetric: architecture independent workload characterization for transactional memory benchmarks. | James Poe, Clay Hughes, Tao Li |
| 2009 | Towards 100 gbit/s ethernet: multicore-based parallel communication protocol design. | Stavros Passas, Kostas Magoutis, Angelos Bilas |
| 2009 | High-performance CUDA kernel execution on FPGAs. | Alexandros Papakonstantinou, Karthik Gururaj, John A. Stratton, Deming Chen, Jason Cong, Wen-mei W. Hwu |
| 2009 | Limited early value communication to improve performance of transactional memory. | Salil Mohan Pant, Gregory T. Byrd |
| 2009 | Subdomain communication to increase scalability in large-scale scientific applications. | Aleksandr Ovcharenko, Onkar Sahni, Christopher D. Carothers, Kenneth E. Jansen, Mark S. Shephard |
| 2009 | Using many-core hardware to correlate radio astronomy signals. | Rob van Nieuwpoort, John W. Romein |
| 2009 | Synchronization optimizations for efficient execution on multi-cores. | Alexandru Nicolau, Guangqiang Li, Alexander V. Veidenbaum, Arun Kejariwal |
| 2009 | Load balancing using work-stealing for pipeline parallelism in emerging applications. | Angeles G. Navarro, Rafael Asenjo, Siham Tabik, Calin Cascaval |
| 2009 | Understanding the interconnection network of SpiNNaker. | Javier Navaridas, Mikel Lujn, Jos Miguel-Alonso, Luis A. Plana, Steve B. Furber |
| 2009 | OhHelp: a scalable domain-decomposing dynamic load balancing for particle-in-cell simulations. | Hiroshi Nakashima, Yohei Miyake, Hideyuki Usui, Yoshiharu Omura |
| 2009 | /scratch as a cache: rethinking HPC center scratch storage. | Henry M. Monti, Ali Raza Butt, Sudharshan S. Vazhkudai |