| 1996 | Don't Use the Page Number, But a Pointer To It. | Andr Seznec |
| 1996 | Correlation and Aliasing in Dynamic Branch Predictors. | Stuart Sechrest, Chih-Chieh Lee, Trevor N. Mudge |
| 1996 | Missing the Memory Wall: The Case for Processor/Memory Integration. | Ashley Saulsbury, Fong Pong, Andreas Nowatzyk |
| 1996 | A Router Architecture for Real-Time Point-to-Point Networks. | Jennifer Rexford, John Hall, Kang G. Shin |
| 1996 | Decoupled Hardware Support for Distributed Shared Memory. | Steven K. Reinhardt, Robert W. Pfile, David A. Wood |
| 1996 | Evaluation of Design Alternatives for a Multiprocessor Microprocessor. | Basem A. Nayfeh, Lance Hammond, Kunle Olukotun |
| 1996 | Coherent Network Interfaces for Fine-Grain Communication. | Shubhendu S. Mukherjee, Babak Falsafi, Mark D. Hill, David A. Wood |
| 1996 | COMA: An Opportunity for Building Fault-Tolerant Scalable Shared Memory Multiprocessors. | Christine Morin, Alain Gefflaut, Michel Bantre, Anne-Marie Kermarrec |
| 1996 | Polling Watchdog: Combining Polling and Interrupts for Efficient Message Handling. | Olivier Maquelin, Guang R. Gao, Herbert H. J. Hum, Kevin B. Theobald, Xinmin Tian |
| 1996 | STiNG: A CC-NUMA Computer System for the Commercial Marketplace. | Tom Lovett, Russell M. Clapp |
| 1996 | Rotating Combined Queueing (RCQ): Bandwidth and Latency Guarantees in Low-Cost, High-Performance Networks. | Jae H. Kim, Andrew A. Chien |
| 1996 | The Difference-bit Cache. | Toni Juan, Toms Lang, Juan J. Navarro |
| 1996 | Understanding Application Performance on Shared Virtual Memory Systems. | Liviu Iftode, Jaswinder Pal Singh, Kai Li |
| 1996 | DCD - Disk Caching Disk: A New Approach for Boosting I/O Performance. | Yiming Hu, Qing Yang |
| 1996 | Informing Memory Operations: Providing Memory Performance Feedback in Modern Processors. | Mark Horowitz, Margaret Martonosi, Todd C. Mowry, Michael D. Smith |
| 1996 | Application and Architectural Bottlenecks in Large Scale Distributed Shared Memory Machines. | Chris Holt, Jaswinder Pal Singh, John L. Hennessy |
| 1996 | Performance Comparison of ILP Machines with Cycle Time Evaluation. | Tetsuya Hara, Hideki Ando, Chikako Nakanishi, Masao Nakaya |
| 1996 | An Analysis of Dynamic Branch Prediction Schemes on System Workloads. | Nicholas C. Gloy, Cliff Young, J. Bradley Chen, Michael D. Smith |
| 1996 | Early Experience with Message-Passing on the SHRIMP Multicomputer. | Edward W. Felten, Richard Alpert, Angelos Bilas, Matthias A. Blumrich, Douglas W. Clark, Stefanos N. Damianakis, Cezary Dubnicki, Liviu Iftode, Kai Li |
| 1996 | Using Hybrid Branch Predictors to Improve Branch Prediction Accuracy in the Presence of Context Switches. | Marius Evers, Po-Yung Chang, Yale N. Patt |
| 1996 | Evaluation of Multithreaded Uniprocessors for Commercial Application Environments. | Richard J. Eickemeyer, Ross E. Johnson, Steven R. Kunkel, Mark S. Squillante, Shiafun Liu |
| 1996 | Compiler and Hardware Support for Cache Coherence in Large-Scale Multiprocessors: Design Considerations and Performance Study. | Lynn Choi, Pen-Chung Yew |
| 1996 | Memory Bandwidth Limitations of Future Microprocessors. | Doug Burger, James R. Goodman, Alain Kgi |
| 1996 | High-Bandwidth Address Translation for Multiple-Issue Processors. | Todd M. Austin, Gurindar S. Sohi |
| 1995 | Speeding Up Irregular Applications in Shared-Memory Multiprocessors: Memory Binding and Group Prefetching. | Zheng Zhang, Josep Torrellas |