| 1999 | The scalability of multigrain systems. | Donald Yeung |
| 1999 | Adapting cache line size to application behavior. | Alexander V. Veidenbaum, Weiyu Tang, Rajesh K. Gupta, Alexandru Nicolau, Xiaomei Ji |
| 1999 | SMARTS: exploiting temporal locality and parallelism through vertical execution. | Suvas Vajracharya, Steve Karmesin, Peter H. Beckman, James Crotinger, Allen D. Malony, Sameer Shende, R. R. Oldehoeft, Stephen Smith |
| 1999 | The design and evaluation of high performance communication using a Gigabit Ethernet. | Shinji Sumimoto, Hiroshi Tezuka, Atsushi Hori, Hiroshi Harada, Toshiyuki Takahashi, Yutaka Ishikawa |
| 1999 | A design analysis of a hybrid technology multithreaded architecture for petaflops scale computation3. | Thomas L. Sterling, Larry A. Bergman |
| 1999 | A new "quad-tree-based" sub-system allocation technique for mesh-connected parallel machines. | Jeeraporn Srisawat, Nikitas A. Alexandridis |
| 1999 | Realizing the performance potential of the virtual interface architecture. | Evan Speight, Hazim Abdel-Shafi, John K. Bennett |
| 1999 | Reducing cache misses using hardware and software page placement. | Timothy Sherwood, Brad Calder, Joel S. Emer |
| 1999 | CACHET: an adaptive cache coherence protocol for distributed shared-memory systems. | Xiaowei Shen, Arvind, Larry Rudolph |
| 1999 | A comparison of MPI, SHMEM and cache-coherent shared address space programming models on the SGI Origin2000. | Hongzhang Shan, Jaswinder Pal Singh |
| 1999 | Efficient management of memory hierarchies in embedded DRAM systems. | Ashley Saulsbury, Su-Jaen Huang, Fredrik Dahlgren |
| 1999 | A locality sensitive multi-module cache with explicit management. | F. Jess Snchez, Antonio Gonzlez |
| 1999 | Mechanisms and policies for supporting fine-grained cycle stealing. | Kyung Dong Ryu, Jeffrey K. Hollingsworth, Peter J. Keleher |
| 1999 | Improving virtual function call target prediction via dependence-based pre-computation. | Amir Roth, Andreas Moshovos, Gurindar S. Sohi |
| 1999 | The pool of subsectors cache design. | Jeffrey B. Rothman, Alan Jay Smith |
| 1999 | Eliminating synchronization bottlenecks in object-based programs using adaptive replication. | Martin C. Rinard, Pedro C. Diniz |
| 1999 | Classifying load and store instructions for memory renaming. | Glenn Reinman, Brad Calder, Dean M. Tullsen, Gary S. Tyson, Todd M. Austin |
| 1999 | Software trace cache. | Alex Ramrez, Josep Llus Larriba-Pey, Carlos Navarro, Josep Torrellas, Mateo Valero |
| 1999 | Resource usage models for instruction scheduling: two new models and a classification. | V. Janaki Ramanan, Ramaswamy Govindarajan |
| 1999 | On the complexity of list scheduling algorithms for distributed-memory systems. | Andrei Radulescu, Arjan J. C. van Gemund |
| 1999 | Adding a vector unit to a superscalar processor. | Francisca Quintana, Jess Corbal, Roger Espasa, Mateo Valero |
| 1999 | Low-level router design and its impact on supercomputer system performance. | Valentin Puente, Jos A. Gregorio, Cruz Izu, Ramn Beivide, Fernando Vallejo |
| 1999 | Responsiveness without interrupts. | Dejan Perkovic, Peter J. Keleher |
| 1999 | Improving the performance of speculatively parallel applications on the Hydra CMP. | Kunle Olukotun, Lance Hammond, Mark Willey |
| 1999 | Dynamic remote memory acquisition for parallel data mining on ATM-connected PC cluster. | Masato Oguchi, Masaru Kitsuregawa |