| 1996 | Data Prefetching and Multilevel Blocking for Linear Algebra Operations. | Juan J. Navarro, Elena Garca-Diego, Jos R. Herrero |
| 1996 | Fine Grain Parallel Communication on General Purpose LANs. | Todd W. Mummert, Corey Kosak, Peter Steenkiste, Allan Fisher |
| 1996 | Design and Evaluation of Dynamic Access Ordering Hardware. | Sally A. McKee, Assaji Aluwihare, Benjamin H. Clark, Robert H. Klenke, Trevor C. Landon, Christopher W. Oliver, Maximo H. Salinas, Adam E. Szymkowiak, Kenneth L. Wright, William A. Wulf, James H. Aylor |
| 1996 | Compiler Support for Hybrid Irregular Accesses on Multicomputers. | Antonio Lain, Prithviraj Banerjee |
| 1996 | Evaluating Virtual Channels for Cache-Coherent Shared-Memory Multiprocessors. | Akhilesh Kumar, Laxmi N. Bhuyan |
| 1996 | Automating Parallel Runtime Optimizations Using Post-Mortem Analysis. | Sanjeev Krishnan, Laxmikant V. Kal |
| 1996 | A Register Allocation Technique Using Guarded PDG. | Akira Koseki, Hideaki Komatsu, Yoshiaki Fukazawa |
| 1996 | A Cost-Comparison Approach for Adaptive Distributed Shared Memory. | Jai-Hoon Kim, Nitin H. Vaidya |
| 1996 | Amon2: A Parallel Wire Routing Algorithm on a Torus Network Parallel Computer. | Hesham Keshk, Shin-ichiro Mori, Hiroshi Nakashima, Shinji Tomita |
| 1996 | Minimizing Communication While Preserving Parallelism. | Wayne Kelly, William W. Pugh |
| 1996 | The GLOW Cache Coherence Protocol Extensions for Widely Shared Data. | Stefanos Kaxiras, James R. Goodman |
| 1996 | Performance of the Vectorial Processor VEC-SM2 Using Serial Multiport Memory. | Jacques Jorda, Abdelaziz Mzoughi, O. Lafontaine, Daniel Litaize |
| 1996 | Mapping Performance Data for High-Level and Data Views of Parallel Program Performance. | R. Bruce Irvin, Barton P. Miller |
| 1996 | Synchronization Hardware for Networks of Workstations: Performance vs. Cost. | Rahmat S. Hyder, David A. Wood |
| 1996 | Integrating Task and Data Parallelism Using Shared Objects. | Saniya Ben Hassen, Henri E. Bal |
| 1996 | Examination of a Memory Access Classification Scheme for Pointer-Intensive and Numeric Programs. | Luddy Harrison |
| 1996 | Evaluating the Limits of Message Passing via the Shared Attraction Memory on CC-COMA Machines: Experiences with TCGMSG and PVM. | Kaushik Ghosh, Stephen R. Breit |
| 1996 | Parallel Implementation of the Lanczos Method for Sparse Matrices: Analysis of Data Distributions. | Ester M. Garzn, Inmaculada Garca |
| 1996 | Automatic Partitioning Techniques for Solving Partial Differential Equations on Irregular Adaptive Meshes. | Jrme Galtier |
| 1996 | Run-Time Compilation for Parallel Sparse Matrix Computations. | Cong Fu, Tao Yang |
| 1996 | Improving Single-Process Performance with Multithreaded Processors. | Alexandre Farcy, Olivier Temam |
| 1996 | CTADEL: A Generator of Multi-Platform High Performance Codes for PDE-Based Scientific Applications. | Robert van Engelen, Lex Wolters, Gerard Cats |
| 1996 | Are There Advantages to High-Dimension Architectures? Analysis of | Shantanu Dutt, Nam Trinh |
| 1996 | ParInt: A Software Package for Parallel Integration. | Elise de Doncker, Ajay Gupta, Jay Ball, Patricia Ealy, Alan Genz |
| 1996 | A Performance Study of Cosmological Simulations on Message-Passing and Shared-Memory Multiprocessors. | Marios D. Dikaiakos, Joachim Stadel |