| 2025 | Multi-Core Aware Evaluation of Prefetchers. | Mart Torrents, Paul Caheny, Stijn Eyerman, Wim Heirman |
| 2025 | FinGraV: Methodology for Fine-Grain GPU Power Visibility and Insights. | Varsha Singhania, Shaizeen Aga, Mohamed Assem Ibrahim |
| 2025 | RayFlex: An Open-Source RTL Implementation of the Hardware Ray Tracer Datapath. | Fangjia Shen, Aaron Barnes, Anusuya Nallathambi, Timothy G. Rogers |
| 2025 | Evaluating Compute in Memory Architectures for Matrix Multiplication: A Dataflow-Centric Perspective. | Tanvi Sharma, Indranil Chakraborty, Mustafa Fayez Ali, Kaushik Roy |
| 2025 | Exploring Constrained Dataflow Accelerators for Real-Time Multi-Task Multi-Model Ml Workloads. | Jamin Seo, Jianming Tong, Tushar Krishna, Hyoukjun Kwon |
| 2025 | Benchmarking 3D Gaussian Splatting Rendering. | Saichand Samudrala, Sushant Kondguli, Paul Gratz |
| 2025 | Evaluation and Comparison of the Energy Efficiency of Several Intel Multicore Processors. | Thomas Rauber, Gudula Rnger |
| 2025 | SCALE-Sim V3: a Modular Cycle-Accurate Systolic Accelerator Simulator for End-To-End System Analysis. | Ritik Raj, Sarbartha Banerjee, Nikhil Chandra, Zishen Wan, Jianming Tong, Ananda Samajdar, Tushar Krishna |
| 2025 | An Analytical Cost Model for Fast Evaluation of Multiple Compute-Engine CNN Accelerators. | Fareed Qararyah, Mohammad Ali Maleki, Pedro Trancoso |
| 2025 | La Superba: Leveraging a Self-Comparison Method to Understand the Performance Benefits of Sparse Acceleration Optimizations. | Nebil Ozer, Gregory Kollmer, Ramyad Hadidi, Bahar Asgari |
| 2025 | Performance Analysis of GEMM Workloads on the AMD Versal Platform. | Kaustubh Manohar Mhatre, Venkata Guru Prashanth Mulleti, Curt John Bansil, Endri Taka, Aman Arora |
| 2025 | Beyond the Numbers: Measuring Android Performance Through User Perception. | Jaeheon Lee, Juhyung Park, Seonggyun Oh, Jinhyung Koo, Sungjin Lee |
| 2025 | Characterizing Compute-Communication Overlap in GPU-Accelerated Distributed Deep Learning: Performance and Power Implications. | Seonho Lee, Jihwan Oh, Seokjin Go, Divya Mahajan |
| 2025 | COSMOS: An LLC Contention Slowdown Model for Heterogeneous Multi-Core Systems. | Yongju Lee, Jaewon Kwon, Cheolhwan Kim, Enhyeok Jang, Jiwon Lee, Hyunwuk Lee, Won Woo Ro |
| 2025 | Beethoven: A Heterogeneous Multi-Core Accelerator System Composer. | Chris Kjellqvist, Brendan Peercy, Alvin R. Lebeck, Lisa Wu Wills |
| 2025 | ADOR: A Design Exploration Framework for LLM Serving with Enhanced Latency and Throughput. | Junsoo Kim, Hunjong Lee, Geonwoo Ko, Gyubin Choi, Seri Ham, Seongmin Hong, Joo-Young Kim |
| 2025 | Understanding the Performance Horizon of the Latest ML Workloads with NonGEMM Workloads. | Rachid Karami, Sheng-Chun Kao, Hyoukjun Kwon |
| 2025 | Intel in-Memory Analytics Accelerator: Performance Characterization and Guidelines. | Jaeyoung Kang, Qirong Xia, Ipoom Jeong, Yongjoo Park, Nam Sung Kim |
| 2025 | Hierarchical Traversal Stack Design Using Shared Memory for GPU Ray Tracing. | Eunsoo Jung, Eunbi Jeong, Gunjae Koo, Yunho Oh, Myung Kuk Yoon |
| 2025 | PIM-BEACON: A Benchmarking and Emulation Framework Supporting Adaptive CONfigurations in DRAM-Based Processing-in-Memory Systems. | Inseong Hwang, Jihoon Jang, Chaewon Park, Hyun Kim |
| 2025 | TPNM: A CXL Based General Purpose Tiered Process Near Memory Framework. | Pingyi Huo, Anusha Devulapally, Hasan Al Maruf, Meena Arunachalam, Mahmut Taylan Kandemir, Vijaykrishnan Narayanan |
| 2025 | GPU Simulation Acceleration via Parallelization. | Rodrigo Huerta, Antonio Gonzlez |
| 2025 | MeMo: Enhancing Representative Sampling via Mechanistic Micro-Model Signatures. | Chenji Han, Huai Xu, Guangyao Guo, Yuxuan Wu, Fuxin Zhang |
| 2025 | Concurrent PIM and Load/Store Servicing in PIM-Enabled Memory. | Sudhanshu Gupta, Niti Madan, Sooraj Puthoor, Nuwan Jayasena, Sandhya Dwarkadas |
| 2025 | Energon: A Sustainability-Driven Modeling Framework for AI Data Centers. | Wenzhe Guo, Joyjit Kundu, Uras Tos, Giuliano Sisto, Cedric Rolin, Lars-ke Ragnarsson, Timon Evenblij |