| 1997 | Improving Data Cache Performance by Pre-Executing Instructions Under a Cache Miss. | James Dundas, Trevor N. Mudge |
| 1997 | Performance considerations in software multicasts. | Jrg Cordsen, Hans Werner Pohl, Wolfgang Schrder-Preikschat |
| 1997 | Compiler and Run-Time Support for Semi-Structured Applications. | Nikos Chrisochoides, Induprakas Kodukula, Keshav Pingali |
| 1997 | Optimizing Collective I/O Performance on Parallel Computers: A Multisystem Study. | Ying Chen, Jarek Nieplocha, Ian T. Foster, Marianne Winslett |
| 1997 | Implementation of Collective I/O in the Intel Paragon Parallel File System: Initial Experiences. | Rajesh Bordawekar |
| 1997 | CP-PACS: A Massively Parallel Processor for Large Scale Scientific Calculations. | Taisuke Boku, Ken'ichi Itakura, Hiroshi Nakamura, Kisaburo Nakazawa |
| 1997 | Optimizing Matrix Multiply Using PHiPAC: A Portable, High-Performance, ANSI C Coding Methodology. | Jeff A. Bilmes, Krste Asanovic, Chee-Whye Chin, James Demmel |
| 1997 | Multiprocessor Scheduling with Client Resources to Improve the Response Time of WWW Applications. | Daniel Andresen, Tao Yang |
| 1996 | Data-Localization for Fortran Macro-Dataflow Computation Using Partial Static Task Assignment. | Akimasa Yoshida, Kenichi Koshizuka, Hironori Kasahara |
| 1996 | Optimizing Primary Data Caches for Parallel Scientific Applications: The Pool Buffer Approach. | Liuxi Yang, Josep Torrellas |
| 1996 | Memory Organization in Multi-Channel Optical Networks: NUMA and COMA Revisited. | Yan Yang Xiao, John K. Bennett |
| 1996 | The Effect of Interrupts on Software Pipeline Execution on Message-Passing Architectures. | Rob F. Van der Wijngaart, Sekhar R. Sarukkai, Pankaj Mehra |
| 1996 | Benchmark Tests on the Digital Equipment Corporation Alpha AXP 21164-based AlphaServer 8400, Including a Comparison of Optimized Vector and Superscalar Processing. | Harvey J. Wasserman |
| 1996 | Experimental Evaluation of Efficient Sparse Matrix Distributions. | Manuel Ujaldon, Shamik D. Sharma, Emilio L. Zapata, Joel H. Saltz |
| 1996 | Profile Driven Weighted Decomposition. | Karen A. Tomko, Edward S. Davidson |
| 1996 | An Efficient Steepest-Edge Simplex Algorithm for SIMD Computers. | Michael E. Thomadakis, Jyh-Charn Liu |
| 1996 | Hybrid Algorithms for Complete Exchange in 2D Meshes. | N. S. Sundar, Doddaballapur Narasimha-Murthy Jayasimha, Dhabaleswar K. Panda, P. Sadayappan |
| 1996 | Detection and Global Optimization of Reduction Operations for Distributed Parallel Machines. | Toshio Suganuma, Hideaki Komatsu, Toshio Nakatani |
| 1996 | Satisfiability Test with Synchronous Simulated Annealing on the Fujitsu AP1000 Massively-Parallel Multiprocessor. | Andrew Sohn, Rupak Biswas |
| 1996 | Analysis of Local Enumeration and Storage Schemes in HPF. | Henk J. Sips, Kees van Reeuwijk, Will Denissen |
| 1996 | A MATLAB to Fortran 90 Translator and Its Effectiveness. | Luiz De Rose, David A. Padua |
| 1996 | Runtime Coupling of Data-Parallel Programs. | M. Ranganathan, Anurag Acharya, Guy Edjlali, Alan Sussman, Joel H. Saltz |
| 1996 | A Template for Non-Uniform Parallel Loops Based on Dynamic Scheduling and Prefetching Techniques. | Salvatore Orlando, Raffaele Perego |
| 1996 | The Galley Parallel File System. | Nils Nieuwejaar, David Kotz |
| 1996 | Block Algorithms for Sparse Matrix Computations on High Performance Workstations. | Juan J. Navarro, Elena Garca-Diego, Josep Llus Larriba-Pey, Toni Juan |