| 2010 | An experimental approach to performance measurement of heterogeneous parallel applications using CUDA. | Allen D. Malony, Scott Biersdorff, Wyatt Spear, Shangkar Mayanglambam |
| 2010 | Fast and accurate NCBI BLASTP: acceleration with multiphase FPGA-based prefiltering. | Atabak Mahram, Martin C. Herbordt |
| 2010 | A compiler-automated array compression scheme for optimizing memory intensive programs. | Lixia Liu, Zhiyuan Li |
| 2010 | The auction: optimizing banks usage in Non-Uniform Cache Architectures. | Javier Lira, Carlos Molina, Antonio Gonzlez |
| 2010 | High-throughput Bayesian network learning using heterogeneous multicore computers. | Michael D. Linderman, Robert V. Bruggner, Vivek Athalye, Teresa H. Meng, Narges Bani Asadi, Garry P. Nolan |
| 2010 | Optimal bucket algorithms for large MPI collectives on torus interconnects. | Nikhil Jain, Yogish Sabharwal |
| 2010 | The next-generation supercomputer project and a plan for the advanced institute for computational science. | Kimihiko Hirao |
| 2010 | An empirically tuned 2D and 3D FFT library on CUDA GPU. | Liang Gu, Xiaoming Li, Jakob Siegel |
| 2010 | SAMS multi-layout memory: providing multiple views of data to boost SIMD performance. | Chunyang Gou, Georgi Kuzmanov, Georgi Gaydadjiev |
| 2010 | Clustering performance data efficiently at massive scales. | Todd Gamblin, Bronis R. de Supinski, Martin Schulz, Robert J. Fowler, Daniel A. Reed |
| 2010 | FPGA accelerating double/quad-double high precision floating-point applications for ExaScale computing. | Yong Dou, Yuanwu Lei, Guiming Wu, Song Guo, Jie Zhou, Li Shen |
| 2010 | Throughput computing. | William J. Dally |
| 2010 | Evaluation of parallel H.264 decoding strategies for the Cell Broadband Engine. | Chi Ching Chi, Ben H. H. Juurlink, Cor Meenderinck |
| 2010 | Large-scale FFT on GPU clusters. | Yifeng Chen, Xiang Cui, Hong Mei |
| 2010 | Static reuse distances for locality-based optimizations in MATLAB. | Arun Chauhan, Chun-Yu Shei |
| 2010 | Indemics: an interactive data intensive framework for high performance epidemic simulation. | Keith R. Bisset, Jiangzhuo Chen, Xizhou Feng, Yifei Ma, Madhav V. Marathe |
| 2010 | An approach to resource-aware co-scheduling for CMPs. | Major Bhadauria, Sally A. McKee |
| 2010 | Decomposable and responsive power models for multicore processors using performance counters. | Ramon Bertran, Marc Gonzlez, Xavier Martorell, Nacho Navarro, Eduard Ayguad |
| 2010 | Making nested parallel transactions practical using lightweight hardware support. | Woongki Baek, Nathan Grasso Bronson, Christos Kozyrakis, Kunle Olukotun |
| 2010 | Untitled record | Narges Bani Asadi, Christopher W. Fletcher, Greg Gibeling, John Wawrzynek, Wing H. Wong, Garry P. Nolan |
| 2009 | Divide-and-conquer: a bubble replacement for low level caches. | Chuanjun Zhang, Bing Xue |
| 2009 | Less reused filter: improving l2 cache performance via filtering less reused lines. | Lingxiang Xiang, Tianzhou Chen, Qingsong Shi, Wei Hu |
| 2009 | Combining thread level speculation helper threads and runahead execution. | Polychronis Xekalakis, Nikolas Ioannou, Marcelo Cintra |
| 2009 | Dynamic parallelization of single-threaded binary programs using speculative slicing. | Cheng Wang, Youfeng Wu, Edson Borin, Shiliang Hu, Wei Liu, Dave Sager, Tin-Fook Ngai, Jesse Fang |
| 2009 | Practice of parallelizing network applications on multi-core architectures. | Junchang Wang, Haipeng Cheng, Bei Hua, Xinan Tang |