Skip to content

Per Stenstrm

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

72

Venues

22

Active years

1989–2025

Best venue rank

A*

Where they publish

Papers

72 indexed papers, newest first.

YearVenueTitleAuthors
2025EuroParBATCH-DNN: Adaptive and Dynamic Batching for Multi-DNN Accelerators.Piyumal Ranawaka, Per Stenstrm
2025SCASaP: Automatic Software Prefetching for Sparse Tensor Computations in MLIR.Konstantinos Sotiropoulos, Jonas Skeppstedt, Per Stenstrm
2024ICSHMComp: Extending Near-Memory Capacity using Compression in Hybrid Memory.Qi Shao, Angelos Arelakis, Per Stenstrm
2022HPCAGBDI: Going Beyond Base-Delta-Immediate Compression with Global Bases.Alexandra Angerd, Angelos Arelakis, Vasilis Spiliopoulos, Erik Sintorn, Per Stenstrm
2020ICPPA GPU Register File using Static Data Compression.Alexandra Angerd, Erik Sintorn, Per Stenstrm
2019ICPPSaC: Exploiting Execution-Time Slack to Save Energy in Heterogeneous Multicore Systems.Muhammad Waqar Azhar, Miquel Perics, Per Stenstrm
2018HPCAProFess: A Probabilistic Hybrid Main Memory Management Framework for High Performance and Fairness.Dmitry Knyaginin, Vassilis Papaefstathiou, Per Stenstrm
2017RTASTiming-Anomaly Free Dynamic Scheduling of Task-Based Parallel Applications.Petros Voudouris, Per Stenstrm, Risat Pathan
2016DATEEUROSERVER: Share-anything scale-out micro-server design.Manolis Marazakis, John Goodacre, Didier Fuin, Paul M. Carpenter, John Thomson, Emil Mats, Antimo Bruno, Per Stenstrm, Jrme Martin, Yves Durand, Isabelle Dor
2016HPCARADAR: Runtime-assisted dead region management for last-level caches.Madhavan Manivannan, Vassilis Papaefstathiou, Miquel Perics, Per Stenstrm
2016RTSSTiming-anomaly free dynamic scheduling of task-based parallel applications.Petros Voudouris, Per Stenstrm, Risat Pathan
2015ICPPEnhancing Garbage Collection Synchronization Using Explicit Bit Barriers.Jochen Hollmann, J. Rubn Titos Gil, Per Stenstrm
2015MICROHyComp: a hybrid cache compression method for selection of data-type-specific compression methods.Angelos Arelakis, Fredrik Dahlgren, Per Stenstrm
2014DATEEffective resource management towards efficient computing.Per Stenstrm
2014ICPPCrystal: A Design-Time Resource Partitioning Method for Hybrid Main Memory.Dmitry Knyaginin, Georgi Gaydadjiev, Per Stenstrm
2014ISCASCAngelos Arelakis, Per Stenstrm
2014RTASOverhead-aware temporal partitioning on multicore processors.Risat Mahmud Pathan, Per Stenstrm, Lars-Goran Green, Torbjorn Hult, Patrik Sandin
2013CGOImproving data access efficiency by using a tagless access buffer (TAB).Alen Bardizbanyan, Peter Gavin, David B. Whalley, Magnus Sjlander, Per Larsson-Edefors, Sally A. McKee, Per Stenstrm
2013HiPCHARP: Adaptive abort recurrence prediction for Hardware Transactional Memory.Adri Armejach, Anurag Negi, Adrin Cristal, Osman S. Unsal, Per Stenstrm, Tim Harris
2013ICPPEfficient Forwarding of Producer-Consumer Data in Task-Based Programs.Madhavan Manivannan, Anurag Negi, Per Stenstrm
2012HPCAπ-TM: Pessimistic invalidation for scalable lazy hardware transactional memory.Anurag Negi, J. Rubn Titos Gil, Manuel E. Acacio, Jos M. Garca, Per Stenstrm
2012SCCritical lock analysis: diagnosing critical section bottlenecks in multithreaded applications.Guancheng Chen, Per Stenstrm
2011CASESA unified approach to eliminate memory accesses early.Mafijul Md. Islam, Per Stenstrm
2011ICPPImplications of Merging Phases on Scalability of Multi-core Architectures.Madhavan Manivannan, Ben H. H. Juurlink, Per Stenstrm
2011ICPPEager Meets Lazy: The Impact of Write-Buffering on Hardware Transactional Memory.Anurag Negi, J. Rubn Titos Gil, Manuel E. Acacio, Jos M. Garca, Per Stenstrm
2011ICSZEBRA: a data-centric, hybrid-policy hardware transactional memory design.J. Rubn Titos Gil, Anurag Negi, Manuel E. Acacio, Jos M. Garca, Per Stenstrm
2011ICSPoster: implications of merging phases on scalability of multi-core architectures.Madhavan Manivannan, Ben H. H. Juurlink, Per Stenstrm
2011SBAC-PADClassification and Elimination of Conflicts in Hardware Transactional Memory Systems.Mridha-Mohammad Waliullah, Per Stenstrm
2010CASESCharacterization and exploitation of narrow-width loads: the narrow-width cache approach.Mafijul Md. Islam, Per Stenstrm
2009ICSCancellation of loads that return zero using zero-value caches.Md. Mafijul Islam, Sally A. McKee, Per Stenstrm
2008DSDLeveraging Data Promotion for Low Power D-NUCA Caches.Alessandro Bardine, Manuel Comparetti, Pierfrancesco Foglia, Giacomo Gabrielli, Cosimo Antonio Prete, Per Stenstrm
2008ICPPAccommodation of the Bandwidth of Large Cache Blocks Using Cache/Memory Link Compression.Martin Thuresson, Per Stenstrm
2007DATEMicroprocessors in the era of terascale integration.Shekhar Borkar, Norman P. Jouppi, Per Stenstrm
2007EuroParStarvation-Free Transactional Memory-System Protocols.M. M. Waliullah, Per Stenstrm
2007HPCAAn Adaptive Shared/Private NUCA Cache Partitioning Scheme for Chip Multiprocessors.Haakon Dybdahl, Per Stenstrm
2007ICPPLoop-level Speculative Parallelism in Embedded Applications.Md. Mafijul Islam, Alexander Busck, Mikael Engbom, Simji Lee, Michel Dubois, Per Stenstrm
2006HiPCA Cache-Partitioning Aware Replacement Policy for Chip Multiprocessors.Haakon Dybdahl, Per Stenstrm, Lasse Natvig
2006HPCAChip-multiprocessing and beyond.Per Stenstrm
2006SBAC-PADScalable Value-Cache Based Compression Schemes for Multiprocessors.Martin Thuresson, Per Stenstrm
2006SBAC-PADDual-Thread Speculation: Two Threads in the Machine are Worth Eight in the Bush.Fredrik Warg, Per Stenstrm
2005ISCAA Robust Main-Memory Compression Scheme.Magnus Ekman, Per Stenstrm
2005ISPASSEnhancing Multiprocessor Architecture Simulation Speed Using Matched-Pair Comparison.Magnus Ekman, Per Stenstrm
2003HiPCOne Chip, One Server: How Do We Exploit Its Power?Per Stenstrm
2003ICPPPerformance and Power Impact of Issue-width in Chip-Multiprocessor Cores.Magnus Ekman, Per Stenstrm
2003ICPPA Novel Approach to Cache Block Reuse Predictions.Jonas Jalminger, Per Stenstrm
2003ISPASSIntegrating complete-system and user-level performance/power simulators: the SimWattch approach.Jianwei Chen, Michel Dubois, Per Stenstrm
2002HPCAThe FAB Predictor: Using Fourier Analysis to Predict the Outcome of Conditional Branches.Martin Kmpe, Per Stenstrm, Michel Dubois
2002ISLPEDTLB and snoop energy-reduction using virtual caches in low-power chip-multiprocessors.Magnus Ekman, Per Stenstrm, Fredrik Dahlgren
2001EuroParA Case Study of Load Distribution in Parallel View Frustum Culling and Collision Detection.Ulf Assarsson, Per Stenstrm
2000EuroParParallel Computer Architecture.Silvia M. Mller, Per Stenstrm, Mateo Valero, Stamatis Vassiliadis
2000HPCAA Prefetching Technique for Irregular Accesses to Linked Data Structures.Magnus Karlsson, Fredrik Dahlgren, Per Stenstrm
2000ISCARecency-based TLB preloading.Ashley Saulsbury, Fredrik Dahlgren, Per Stenstrm
2000SIGMETRICSAn analytical model of the working-set sizes in decision-support systems.Magnus Karlsson, Per Stenstrm
1999RTSSTiming Anomalies in Dynamically Scheduled Microprocessors.Thomas Lundqvist, Per Stenstrm
1999RTCSAA Method to Improve the Estimated Worst-Case Performance of Data Caching.Thomas Lundqvist, Per Stenstrm
1998USENIXSimICS/Sun4m: A Virtual Workstation.Peter S. Magnusson, Fredrik Larsson, Andreas Moestedt, Bengt Werner, Jim Nilsson, Per Stenstrm, Fredrik Lundholm, Magnus Karlsson, Fredrik Dahlgren, Hkan Grahn
1998WCAEA holistic approach to computer system design education based on system simulation techniques.Per Stenstrm, Fredrik Dahlgren
1997EuroParA Performance Tuning Approach for Shared-Memory Multiprocessors.Per Stenstrm, Jonas Skeppstedt
1996HPCAPerformance Evaluation of a Cluster-Based Multiprocessor Built from ATM Switches and Bus-Based Multiprocessor Servers.Magnus Karlsson, Per Stenstrm
1995HPCAEffectiveness of Hardware-Based Stride and Sequential Prefetching in Shared-Memory Multiprocessors.Fredrik Dahlgren, Per Stenstrm
1995ISCAEfficient Strategies for Software-Only Protocols in Shared-Memory Multiprocessors.Hkan Grahn, Per Stenstrm
1994ASPLOSSimple Compiler Algorithms to Reduce Ownership Operhead in Cache Coherence Protocols.Jonas Skeppstedt, Per Stenstrm
1994ICPPReducing the Write Traffic for a Hybrid Cache Protocol.Fredrik Dahlgren, Per Stenstrm
1994ICPPAn Integrated Methodology for the Verification of Directory-Based Cache Protocols.Fong Pong, Per Stenstrm, Michel Dubois
1994ISCACombined Performance Gains of Simple Cache Protocol Extensions.Fredrik Dahlgren, Michel Dubois, Per Stenstrm
1993ICPPFixed and Adaptive Sequential Prefetching in Shared Memory Multiprocessors.Fredrik Dahlgren, Michel Dubois, Per Stenstrm
1993ISCAThe Detection and Elimination of Useless Misses in Multiprocessors.Michel Dubois, Jonas Skeppstedt, Livio Ricciulli, Krishnan Ramamurthy, Per Stenstrm
1993ISCAAn Adaptive Cache Coherence Protocol Optimized for Migratory Sharing.Per Stenstrm, Mats Brorsson, Lars Sandberg
1992ISCAComparative Performance Evaluation of Cache-Coherent NUMA and COMA Architectures.Per Stenstrm, Truman Joe, Anoop Gupta
1991ICPPA Lockup-Free Multiprocessor Cache Design.Per Stenstrm, Fredrik Dahlgren, Lars Lundberg
1991MICROOn Reconfigurable On-Chip Data Caches.Fredrik Dahlgren, Per Stenstrm
1989ISCAA Cache Consistency Protocol for Multiprocessors with Multistage Networks.Per Stenstrm