| 2006 | A distributed system based on web services for computational science simulations. | Keshav Pingali, Paul Stodghill |
| 2006 | Cooperative checkpointing: a robust approach to large-scale systems reliability. | Adam J. Oliner, Larry Rudolph, Ramendra K. Sahoo |
| 2006 | Wide and efficient trace prediction using the local trace predictor. | Juan C. Moure, Domingo Benitez, Dolores Rexachs, Emilio Luque |
| 2006 | Design space exploration for multicore architectures: a power/performance/thermal view. | Matteo Monchiero, Ramon Canal, Antonio Gonzlez |
| 2006 | Efficient remote block-level I/O over an RDMA-capable NIC. | Manolis Marazakis, Konstantinos Xinidis, Vassilis Papaefstathiou, Angelos Bilas |
| 2006 | Coupling prefix caching and collective downloads for remote dataset access. | Xiaosong Ma, Vincent W. Freeh, Tao Yang, Sudharshan Vazhkudai, Tyler A. Simon, Stephen L. Scott |
| 2006 | Accelerator design for protein sequence HMM search. | Rahul P. Maddimsetty, Jeremy Buhler, Roger D. Chamberlain, Mark A. Franklin, Brandon Harris |
| 2006 | On the performance potential of different types of speculative thread-level parallelism: The DL version of this paper includes corrections that were not made available in the printed proceedings. | Arun Kejariwal, Xinmin Tian, Wei Li, Milind Girkar, Sergey Kozhukhov, Hideki Saito, Utpal Banerjee, Alexandru Nicolau, Alexander V. Veidenbaum, Constantine D. Polychronopoulos |
| 2006 | Lightweight lock-free synchronization methods for multithreading. | Arun Kejariwal, Hideki Saito, Xinmin Tian, Milind Girkar, Wei Li, Utpal Banerjee, Alexandru Nicolau, Constantine D. Polychronopoulos |
| 2006 | A case for high performance computing with virtual machines. | Wei Huang, Jiuxing Liu, Blent Abali, Dhabaleswar K. Panda |
| 2006 | Large files, small writes, and pNFS. | Dean Hildebrand, Lee Ward, Peter Honeyman |
| 2006 | Implementing virtual memory in a vector processor with software restart markers. | Mark Hampton, Krste Asanovic |
| 2006 | Accurate memory data flow modeling in statistical simulation. | Davy Genbrugge, Lieven Eeckhout, Koen De Bosschere |
| 2006 | Scalable algorithms for global snapshots in distributed systems. | Rahul Garg, Vijay K. Garg, Yogish Sabharwal |
| 2006 | Scaling MPI to short-memory MPPs such as BG/L. | Montse Farreras, Toni Cortes, Jess Labarta, George Almsi |
| 2006 | STAR-MPI: self tuned adaptive routines for MPI collective operations. | Ahmad Faraj, Xin Yuan, David K. Lowenthal |
| 2006 | Feedback-directed memory disambiguation through store distance analysis. | Changpeng Fang, Steve Carr, Soner nder, Zhenlin Wang |
| 2006 | Online power-performance adaptation of multithreaded programs using hardware event-based prediction. | Matthew Curtis-Maury, James Dzierwa, Christos D. Antonopoulos, Dimitrios S. Nikolopoulos |
| 2006 | MPIPP: an automatic profile-guided parallel process placement toolset for SMP clusters and multiclusters. | Hu Chen, Wenguang Chen, Jian Huang, Bob Robert H. Kuhn |
| 2006 | Experimental evaluation of application-level checkpointing for OpenMP programs. | Greg Bronevetsky, Keshav Pingali, Paul Stodghill |
| 2006 | Design tradeoffs for tiled CMP on-chip networks. | James D. Balfour, William J. Dally |
| 2006 | BranchTap: improving performance with very few checkpoints through adaptive speculation control. | Patrick Akl, Andreas Moshovos |
| 2006 | Heterogeneous way-size cache. | Jaume Abella, Antonio Gonzlez |
| 2005 | What is worth learning from parallel workloads?: a user and session based analysis. | Julia Zilber, Ofer Amit, David Talby |
| 2005 | Fast branch misprediction recovery in out-of-order superscalar processors. | Peng Zhou, Soner nder, Steve Carr |