| 2000 | Optimized unrolling of nested loops. | Vivek Sarkar |
| 2000 | Improving parallel system performance by changing the arrangement of the network links. | Valentin Puente, Cruz Izu, Jos A. Gregorio, Ramn Beivide, J. M. Prellezo, Fernando Vallejo |
| 2000 | Hardware spatial forwarding for widely shared data. | Marius Pirvu, Laxmi N. Bhuyan |
| 2000 | Boosting superpage utilization with the shadow memory and the partial-subblock TLB. | Cheol Ho Park, JaeWoong Chung, Byeong Hag Seong, Yangwoo Roh, Daeyeon Park |
| 2000 | Comparative study of page-based and segment-based software DSM through compiler optimization. | Junpei Niwa, Takashi Matsumoto, Kei Hiraki |
| 2000 | A case for use-level dynamic page migration. | Dimitrios S. Nikolopoulos, Theodore S. Papatheodorou, Constantine D. Polychronopoulos, Jess Labarta, Eduard Ayguad |
| 2000 | An adaptive software library for fast Fourier transforms. | Dragan Mirkovic, Rishad Mahasoom, S. Lennart Johnsson |
| 2000 | A general performance model for parallel sweeps on orthogonal grids for particle transport calculations. | Mark M. Mathis, Nancy M. Amato, Marvin L. Adams |
| 2000 | Next-generation generic programming and its application to sparse matrix computations. | Nikolay Mateev, Keshav Pingali, Paul Stodghill, Vladimir Kotlyar |
| 2000 | Using profiling to reduce branch misprediction costs on a dynamically scheduled processor. | Srinivas Mantripragada, Alexandru Nicolau |
| 2000 | Using complete system simulation to characterize SPECjvm98 benchmarks. | Tao Li, Lizy Kurian John, Narayanan Vijaykrishnan, Anand Sivasubramaniam, Jyotsna Sabarinathan, Anupama Murthy |
| 2000 | Unroll-based register coalescing. | Suhyun Kim, Soo-Mook Moon, Jinpyo Park, Kemal Ebcioglu |
| 2000 | Fast greedy weighted fusion. | Ken Kennedy |
| 2000 | Using accurate arithmetics to improve numerical reproducibility and stability in parallel applications. | Yun He, Chris H. Q. Ding |
| 2000 | A compiler method for the parallel execution of irregular reductions in scalable shared memory multiprocessors. | Eladio Gutirrez, Oscar G. Plata, Emilio L. Zapata |
| 2000 | Binary translation and architecture convergence issues for IBM system/390. | Michael Gschwind, Kemal Ebcioglu, Erik R. Altman, Sumedh W. Sathaye |
| 2000 | Automated cache optimizations using CME driven diagnosis. | Somnath Ghosh, Margaret Martonosi, Sharad Malik |
| 2000 | Performance evaluation of a new routing strategy for irregular networks with source routing. | Jos Flich, Manuel P. Malumbres, Pedro Lpez, Jos Duato |
| 2000 | Compiling object-oriented data intensive applications. | Renato Ferreira, Gagan Agrawal, Joel H. Saltz |
| 2000 | Design of dynamic load-balancing tools for parallel applications. | Karen D. Devine, Bruce Hendrickson, Erik G. Boman, Matthew St. John, Courtenay T. Vaughan |
| 2000 | Characterizing processor architectures for programmable network interfaces. | Patrick Crowley, Marc E. Fiuczynski, Jean-Loup Baer, Brian N. Bershad |
| 2000 | A low-complexity issue logic. | Ramon Canal, Antonio Gonzlez |
| 2000 | Automatic loop transformations and parallelization for Java. | Pedro V. Artigas, Manish Gupta, Samuel P. Midkiff, Jos E. Moreira |
| 2000 | Synthesizing transformations for locality enhancement of imperfectly-nested loop nests. | Nawaaz Ahmed, Nikolay Mateev, Keshav Pingali |
| 1999 | Fast cluster failover using virtual memory-mapped communication. | Yuanyuan Zhou, Peter M. Chen, Kai Li |