| 2008 | Stencil computation optimization and auto-tuning on state-of-the-art multicore architectures. | Kaushik Datta, Mark Murphy, Vasily Volkov, Samuel Williams, Jonathan Carter, Leonid Oliker, David A. Patterson, John Shalf, Katherine A. Yelick |
| 2008 | Performance potential of molecular dynamics simulations on high performance reconfigurable computing systems. | Matt Chiu, Martin C. Herbordt, Martin Langhammer |
| 2008 | Extending CC-NUMA systems to support write update optimizations. | Liqun Cheng, John B. Carter |
| 2008 | Hiding I/O latency with pre-execution prefetching for parallel applications. | Yong Chen, Surendra Byna, Xian-He Sun, Rajeev Thakur, William Gropp |
| 2008 | High-frequency simulations of global seismic wave propagation using SPECFEM3D_GLOBE on 62K processors. | Laura Carrington, Dimitri Komatitsch, Michael Laurenzano, Mustafa M. Tikir, David Micha, Nicolas Le Goff, Allan Snavely, Jeroen Tromp |
| 2008 | Using server-to-server communication in parallel file systems to simplify consistency and improve performance. | Philip H. Carns, Bradley W. Settlemyer, Walter B. Ligon III |
| 2008 | Parallel I/O prefetching using MPI file caching and I/O signatures. | Surendra Byna, Yong Chen, Xian-He Sun, Rajeev Thakur, William Gropp |
| 2008 | Scalable adaptive mantle convection simulation on petascale supercomputers. | Carsten Burstedde, Omar Ghattas, Michael Gurnis, Georg Stadler, Eh Tan, Tiankai Tu, Lucas C. Wilcox, Shijie Zhong |
| 2008 | Analysis of application heartbeats: learning structural and temporal features in time series data for identification of performance problems. | Emma S. Buneci, Daniel A. Reed |
| 2008 | SMARTMAP: operating system support for efficient data sharing among processes on a multi-core processor. | Ron Brightwell, Kevin T. Pedretti, Trammell Hudson |
| 2008 | Implementing phase unwrapping using Field Programmable Gate Arrays or Graphics Processing Units: A comparison. | Sherman Braganza, Miriam Leeser |
| 2008 | 0.374 Pflop/s trillion-particle kinetic modeling of laser plasma interaction on Roadrunner. | Kevin J. Bowers, Brian J. Albright, Ben Bergen, Lin Yin, Kevin J. Barker, Darren J. Kerbyson |
| 2008 | A dynamic scheduler for balancing HPC applications. | Carlos Boneti, Roberto Gioiosa, Francisco J. Cazorla, Mateo Valero |
| 2008 | EpiSimdemics: an efficient algorithm for simulating the spread of infectious disease over large realistic social networks. | Christopher L. Barrett, Keith R. Bisset, Stephen G. Eubank, Xizhou Feng, Madhav V. Marathe |
| 2008 | Entering the petaflop era: the architecture and performance of Roadrunner. | Kevin J. Barker, Kei Davis, Adolfy Hoisie, Darren J. Kerbyson, Michael Lang, Scott Pakin, Jos Carlos Sancho |
| 2008 | PAM: a novel performance/power aware meta-scheduler for multi-core systems. | Mohammad Banikazemi, Dan E. Poff, Blent Abali |
| 2008 | New algorithm to enable 400+ TFlop/s sustained performance in simulations of disorder effects in high- | Gonzalo Alvarez, Michael S. Summers, Don E. Maxwell, Markus Eisenbach, Jeremy S. Meredith, Jeffrey M. Larkin, John M. Levesque, Thomas A. Maier, Paul R. C. Kent, Eduardo F. D'Azevedo, Thomas C. Schulthess |
| 2008 | Early evaluation of IBM BlueGene/P. | Sadaf R. Alam, Richard F. Barrett, M. Bast, Mark R. Fahey, Jeffery A. Kuehn, Collin McCurdy, James H. Rogers, Philip C. Roth, Ramanan Sankaran, Jeffrey S. Vetter, Patrick H. Worley, Weikuan Yu |
| 2008 | Embarrassingly parallel jobs are not embarrassingly easy to schedule on the grid. | Enis Afgan, Purushotham V. Bangalore |
| 2008 | Nimrod/K: towards massively parallel dynamic grid workflows. | David Abramson, Colin Enticott, Ilkay Altintas |
| 2008 | Massively Parallelized Quasi-Monte Carlo financial Simulation on a FPGA Supercomputer. | Xiang Tian, Khaled Benkrid |
| 2007 | A user-level secure grid file system. | Ming Zhao, Renato J. O. Figueiredo |
| 2007 | The design methodology of Phoenix cluster system software stack. | Jianfeng Zhan, Lei Wang, Bibo Tu, Hui Wang, Zhihong Zhang, Yi Jin, Yu Wen, Yuansheng Chen, Peng Wang, Bizhu Qiu, Dan Meng, Ninghui Sun |
| 2007 | Optimizing center performance through coordinated data staging, scheduling and recovery. | Zhe Zhang, Chao Wang, Sudharshan S. Vazhkudai, Xiaosong Ma, Gregory G. Pike, John W. Cobb, Frank Mueller |
| 2007 | Implementation of the Smith-Waterman algorithm on a reconfigurable supercomputing platform. | Peiheng Zhang, Guangming Tan, Guang R. Gao |