P. Sadayappan
Publication record assembled from the DBLP archive of ranked conferences.
Papers indexed
179
Venues
30
Active years
1985–2025
Best venue rank
A*
Where they publish
- ASC34 papers
- AICS18 papers
- BICPP18 papers
- BPPoPP14 papers
- NationalHiPC14 papers
- A*PLDI11 papers
- CCLUSTER11 papers
- BCC7 papers
- ACGO6 papers
- AHPDC6 papers
- CJSSPP6 papers
- BSPAA4 papers
- MulticonferenceICCS4 papers
- A*ASPLOS3 papers
- A*POPL3 papers
- BCCGRID3 papers
- A*DAC3 papers
- BEuroPar2 papers
- A*KDD1 paper
- A*AAAI1 paper
- AOOPSLA1 paper
- AFPGA1 paper
- A*ICDE1 paper
- A*HPCA1 paper
- NationalFCCM1 paper
- AICDCS1 paper
- CHCW1 paper
- BICPADS1 paper
- CICCD1 paper
- A*ICRA1 paper
Papers
179 indexed papers, newest first.
| Year | Venue | Title | Authors |
|---|---|---|---|
| 2025 | SC | FaSTCC: Fast Sparse Tensor Contractions on CPUs. | Saurabh Raje, Hunter McCoy, Atanas Rountev, Prashant Pandey, P. Sadayappan |
| 2025 | SC | Guiding Application Users via Estimation of Computational Resources for Massively Parallel Chemistry Computations. | Tanzila Tabassum, Omer Subasi, Ajay Panyala, Epiya Ebiapia, Gerald Baumgartner, Erdal Mutlu, P. Sadayappan, Karol Kowalski |
| 2025 | SC | Distributed Sparse Tensor Computations in MLIR. | Miheer Vaidya, Shreya Singh, Devanshu Mantri, Michael Shannon Eydenberg, Brian Michael Kelley, Sivasankaran Rajamanickam, Atanas Rountev, P. Sadayappan |
| 2023 | ICS | Scalable parallelization for the solution of phonon Boltzmann Transport Equation. | Han D. Tran, Siddharth Saurav, P. Sadayappan, Sandip Mazumder, Hari Sundar |
| 2023 | PPoPP | TDC: Towards Extremely Efficient CNNs on GPUs via Hardware-Aware Tucker Decomposition. | Lizhi Xiang, Miao Yin, Chengming Zhang, Aravind Sukumaran-Rajam, P. Sadayappan, Bo Yuan, Dingwen Tao |
| 2023 | SC | Automatic Generation of Distributed-Memory Mappings for Tensor Computations. | Martin Kong, Raneem Abu Yosef, Atanas Rountev, P. Sadayappan |
| 2022 | CC | Training of deep learning pipelines on memory-constrained GPUs via segmented fused-tiled execution. | Yufan Xu, Saurabh Raje, Atanas Rountev, Gerald Sabin, Aravind Sukumaran-Rajam, P. Sadayappan |
| 2022 | CGO | Comprehensive Accelerator-Dataflow Co-design Optimization for Convolutional Neural Networks. | Miheer Vaidya, Aravind Sukumaran-Rajam, Atanas Rountev, P. Sadayappan |
| 2021 | ASPLOS | Analytical characterization and design space exploration for optimization of CNNs. | Rui Li, Yufan Xu, Aravind Sukumaran-Rajam, Atanas Rountev, P. Sadayappan |
| 2021 | PLDI | IOOpt: automatic derivation of I/O complexity bounds for affine programs. | Auguste Olivry, Guillaume Iooss, Nicolas Tollenaere, Atanas Rountev, P. Sadayappan, Fabrice Rastello |
| 2021 | SPAA | Efficient Distributed Algorithms for Convolutional Neural Networks. | Rui Li, Yufan Xu, Aravind Sukumaran-Rajam, Atanas Rountev, P. Sadayappan |
| 2020 | KDD | ALO-NMF: Accelerated Locality-Optimized Non-negative Matrix Factorization. | Gordon Euhyun Moon, J. Austin Ellis, Aravind Sukumaran-Rajam, Srinivasan Parthasarathy, P. Sadayappan |
| 2020 | PLDI | Automated derivation of parametric data movement lower bounds for affine programs. | Auguste Olivry, Julien Langou, Louis-Nol Pouchet, P. Sadayappan, Fabrice Rastello |
| 2020 | SC | Scalable heterogeneous execution of a coupled-cluster model with perturbative triples. | Jinsung Kim, Ajay Panyala, Bo Peng, Karol Kowalski, P. Sadayappan, Sriram Krishnamoorthy |
| 2020 | SC | Efficient tiled sparse matrix multiplication through matrix signatures. | Sreyya Emre Kurt, Aravind Sukumaran-Rajam, Fabrice Rastello, P. Sadayappan |
| 2019 | AAAI | ATP: Directed Graph Embedding with Asymmetric Transitivity Preservation. | Jiankai Sun, Bortik Bandyopadhyay, Armin Bashizade, Jiongqian Liang, P. Sadayappan, Srinivasan Parthasarathy |
| 2019 | CGO | A Code Generator for High-Performance Tensor Contractions on GPUs. | Jinsung Kim, Aravind Sukumaran-Rajam, Vineeth Thumma, Sriram Krishnamoorthy, Ajay Panyala, Louis-Nol Pouchet, Atanas Rountev, P. Sadayappan |
| 2019 | PPoPP | Adaptive sparse tiling for sparse matrix multiplication. | Changwan Hong, Aravind Sukumaran-Rajam, Israt Nisa, Kunal Singh, P. Sadayappan |
| 2019 | SC | Analytical cache modeling and tilesize optimization for tensor contractions. | Rui Li, Aravind Sukumaran-Rajam, Richard Veras, Tze Meng Low, Fabrice Rastello, Atanas Rountev, P. Sadayappan |
| 2019 | SC | Parallel Data-Local Training for Optimizing Word2Vec Embeddings for Word and Graph Embeddings. | Gordon Euhyun Moon, Denis Newman-Griffis, Jinsung Kim, Aravind Sukumaran-Rajam, Eric Fosler-Lussier, P. Sadayappan |
| 2019 | SC | An efficient mixed-mode representation of sparse tensors. | Israt Nisa, Jiajia Li, Aravind Sukumaran-Rajam, Prashant Singh Rawat, Sriram Krishnamoorthy, P. Sadayappan |
| 2018 | HiPC | Sampled Dense Matrix Multiplication for High-Performance Machine Learning. | Israt Nisa, Aravind Sukumaran-Rajam, Sreyya Emre Kurt, Changwan Hong, P. Sadayappan |
| 2018 | HPDC | Efficient sparse-matrix multi-vector product on GPUs. | Changwan Hong, Aravind Sukumaran-Rajam, Bortik Bandyopadhyay, Jinsung Kim, Sreyya Emre Kurt, Israt Nisa, Shivani Sabhlok, mit V. atalyrek, Srinivasan Parthasarathy, P. Sadayappan |
| 2018 | ICCS | Parallel Latent Dirichlet Allocation on GPUs. | Gordon Euhyun Moon, Israt Nisa, Aravind Sukumaran-Rajam, Bortik Bandyopadhyay, Srinivasan Parthasarathy, P. Sadayappan |
| 2018 | ICS | Optimizing Tensor Contractions in CCSD(T) for Efficient Execution on GPUs. | Jinsung Kim, Aravind Sukumaran-Rajam, Changwan Hong, Ajay Panyala, Rohit Kumar Srivastava, Sriram Krishnamoorthy, P. Sadayappan |
| 2018 | PLDI | GPU code optimization using abstract kernel emulation and sensitivity analysis. | Changwan Hong, Aravind Sukumaran-Rajam, Jinsung Kim, Prashant Singh Rawat, Sriram Krishnamoorthy, Louis-Nol Pouchet, Fabrice Rastello, P. Sadayappan |
| 2018 | PPoPP | Performance modeling for GPUs using abstract kernel emulation. | Changwan Hong, Aravind Sukumaran-Rajam, Jinsung Kim, Prashant Singh Rawat, Sriram Krishnamoorthy, Louis-Nol Pouchet, Fabrice Rastello, P. Sadayappan |
| 2018 | PPoPP | Register optimizations for stencils on GPUs. | Prashant Singh Rawat, Fabrice Rastello, Aravind Sukumaran-Rajam, Louis-Nol Pouchet, Atanas Rountev, P. Sadayappan |
| 2018 | SC | Associative instruction reordering to alleviate register pressure. | Prashant Singh Rawat, Aravind Sukumaran-Rajam, Atanas Rountev, Fabrice Rastello, Louis-Nol Pouchet, P. Sadayappan |
| 2017 | HiPC | Characterization of Data Movement Requirements for Sparse Matrix Computations on GPUs. | Sreyya Emre Kurt, Vineeth Thumma, Changwan Hong, Aravind Sukumaran-Rajam, P. Sadayappan |
| 2017 | HiPC | Parallel LDA with Over-Decomposition. | Gordon Euhyun Moon, Aravind Sukumaran-Rajam, P. Sadayappan |
| 2017 | ICS | On improving performance of sparse matrix-matrix multiplication on GPUs. | Rakshith Kunchum, Ankur Chaudhry, Aravind Sukumaran-Rajam, Qingpeng Niu, Israt Nisa, P. Sadayappan |
| 2017 | PPoPP | Parallel CCD++ on GPU for Matrix Factorization. | Israt Nisa, Aravind Sukumaran-Rajam, Rakshith Kunchum, P. Sadayappan |
| 2017 | PPoPP | Optimizing the Four-Index Integral Transform Using Data Movement Lower Bounds Analysis. | Samyam Rajbhandari, Fabrice Rastello, Karol Kowalski, Sriram Krishnamoorthy, P. Sadayappan |
| 2016 | CC | Register allocation and promotion through combined instruction scheduling and loop unrolling. | Lukasz Domagala, Duco van Amstel, Fabrice Rastello, P. Sadayappan |
| 2016 | CC | On fusing recursive traversals of K-d trees. | Samyam Rajbhandari, Jinsung Kim, Sriram Krishnamoorthy, Louis-Nol Pouchet, Fabrice Rastello, Robert J. Harrison, P. Sadayappan |
| 2016 | HiPC | Compiler Support for Software Cache Coherence. | Sanket Tavarageri, Wooil Kim, Josep Torrellas, P. Sadayappan |
| 2016 | PLDI | Effective padding of multidimensional arrays to avoid cache conflict misses. | Changwan Hong, Wenlei Bao, Albert Cohen, Sriram Krishnamoorthy, Louis-Nol Pouchet, Fabrice Rastello, J. Ramanujam, P. Sadayappan |
| 2016 | POPL | PolyCheck: dynamic verification of iteration space transformations on affine programs. | Wenlei Bao, Sriram Krishnamoorthy, Louis-Nol Pouchet, Fabrice Rastello, P. Sadayappan |
| 2016 | PPoPP | Effective resource management for enhancing performance of 2D and 3D stencils on GPUs. | Prashant Singh Rawat, Changwan Hong, Mahesh Ravishankar, Vinod Grover, Louis-Nol Pouchet, P. Sadayappan |
| 2016 | SC | PIPES: a language and compiler for task-based programming on distributed-memory clusters. | Martin Kong, Louis-Nol Pouchet, P. Sadayappan, Vivek Sarkar |
| 2016 | SC | A domain-specific compiler for a parallel multiresolution adaptive numerical simulation environment. | Samyam Rajbhandari, Jinsung Kim, Sriram Krishnamoorthy, Louis-Nol Pouchet, Fabrice Rastello, Robert J. Harrison, P. Sadayappan |
| 2016 | SPAA | Brief Announcement: Approximating the I/O Complexity of One-Shot Red-Blue Pebbling. | Timothy Carpenter, Fabrice Rastello, P. Sadayappan, Anastasios Sidiropoulos |
| 2015 | CGO | Characterizing and enhancing global memory data coalescing on GPUs. | Naznin Fauzia, Louis-Nol Pouchet, P. Sadayappan |
| 2015 | ICS | Optimistic Delinearization of Parametrically Sized Arrays. | Tobias Grosser, Jagannathan Ramanujam, Louis-Nol Pouchet, P. Sadayappan, Sebastian Pop |
| 2015 | ICS | Automatic Selection of Sparse Matrix Representation on GPUs. | Naser Sedaghati, Te Mu, Louis-Nol Pouchet, Srinivasan Parthasarathy, P. Sadayappan |
| 2015 | POPL | On Characterizing the Data Access Complexity of Programs. | Venmugil Elango, Fabrice Rastello, Louis-Nol Pouchet, J. Ramanujam, P. Sadayappan |
| 2015 | PPoPP | On optimizing machine learning workloads via kernel fusion. | Arash Ashari, Shirish Tatikonda, Matthias Boehm, Berthold Reinwald, Keith Campbell, John Keenleyside, P. Sadayappan |
| 2015 | PPoPP | Distributed memory code generation for mixed Irregular/Regular computations. | Mahesh Ravishankar, Roshan Dathathri, Venmugil Elango, Louis-Nol Pouchet, J. Ramanujam, Atanas Rountev, P. Sadayappan |
| 2015 | SC | An elegant sufficiency: load-aware differentiated scheduling of data transfers. | Rajkumar Kettimuthu, Gayane Vardoyan, Gagan Agrawal, P. Sadayappan, Ian T. Foster |
| 2015 | SC | SDSLc: a multi-target domain-specific compiler for stencil computations. | Prashant Singh Rawat, Martin Kong, Thomas Henretty, Justin Holewinski, Kevin Stock, Louis-Nol Pouchet, J. Ramanujam, Atanas Rountev, P. Sadayappan |
| 2014 | CCGRID | Modeling and Optimizing Large-Scale Wide-Area Data Transfers. | Rajkumar Kettimuthu, Gayane Vardoyan, Gagan Agrawal, P. Sadayappan |
| 2014 | CGO | Hybrid Hexagonal/Classical Tiling for GPUs. | Tobias Grosser, Albert Cohen, Justin Holewinski, P. Sadayappan, Sven Verdoolaege |
| 2014 | HiPC | A fast implementation of MLR-MCL algorithm on multi-core processors. | Qingpeng Niu, Pai-Wei Lai, S. M. Faisal, Srinivasan Parthasarathy, P. Sadayappan |
| 2014 | ICPP | CAST: Contraction Algorithm for Symmetric Tensors. | Samyam Rajbhandari, Akshay Nikam, Pai-Wei Lai, Kevin Stock, Sriram Krishnamoorthy, P. Sadayappan |
| 2014 | ICS | An efficient two-dimensional blocking strategy for sparse matrix-vector multiplication on GPUs. | Arash Ashari, Naser Sedaghati, John Eisenlohr, P. Sadayappan |
| 2014 | OOPSLA | WOSC 2014: second workshop on optimizing stencil computations. | Shoaib Kamil, Saman P. Amarasinghe, P. Sadayappan |
| 2014 | PLDI | A framework for enhancing data reuse via associative reordering. | Kevin Stock, Martin Kong, Tobias Grosser, Louis-Nol Pouchet, Fabrice Rastello, J. Ramanujam, P. Sadayappan |
| 2014 | PLDI | Compiler-assisted detection of transient memory errors. | Sanket Tavarageri, Sriram Krishnamoorthy, P. Sadayappan |
| 2014 | SC | Fast Sparse Matrix-Vector Multiplication on GPUs for Graph Applications. | Arash Ashari, Naser Sedaghati, John Eisenlohr, Srinivasan Parthasarathy, P. Sadayappan |
| 2014 | SC | A Communication-Optimal Framework for Contracting Distributed Tensors. | Samyam Rajbhandari, Akshay Nikam, Pai-Wei Lai, Kevin Stock, Sriram Krishnamoorthy, P. Sadayappan |
| 2014 | SPAA | On characterizing the data movement complexity of computational DAGs for parallel execution. | Venmugil Elango, Fabrice Rastello, Louis-Nol Pouchet, J. Ramanujam, P. Sadayappan |
| 2013 | ASPLOS | Split tiling for GPUs: automatic parallelization using trapezoidal tiles. | Tobias Grosser, Albert Cohen, Paul H. J. Kelly, J. Ramanujam, P. Sadayappan, Sven Verdoolaege |
| 2013 | FPGA | Polyhedral-based data reuse optimization for configurable computing. | Louis-Nol Pouchet, Peng Zhang, P. Sadayappan, Jason Cong |
| 2013 | ICDE | Stratification driven placement of complex data: A framework for distributed data analytics. | Ye Wang, Srinivasan Parthasarathy, P. Sadayappan |
| 2013 | ICS | A stencil compiler for short-vector SIMD architectures. | Thomas Henretty, Richard Veras, Franz Franchetti, Louis-Nol Pouchet, J. Ramanujam, P. Sadayappan |
| 2013 | PLDI | When polyhedral transformations meet SIMD code generation. | Martin Kong, Richard Veras, Kevin Stock, Franz Franchetti, Louis-Nol Pouchet, P. Sadayappan |
| 2013 | SC | A framework for load balancing of tensor contraction expressions via dynamic task partitioning. | Pai-Wei Lai, Kevin Stock, Samyam Rajbhandari, Sriram Krishnamoorthy, P. Sadayappan |
| 2012 | ASPLOS | High-performance sparse matrix-vector multiplication on GPUs for structured grid computations. | Jeswin Godwin, Justin Holewinski, P. Sadayappan |
| 2012 | CC | Analytical Bounds for Optimal Tile Size Selection. | Jun Shirako, Kamal Sharma, Naznin Fauzia, Louis-Nol Pouchet, J. Ramanujam, P. Sadayappan, Vivek Sarkar |
| 2012 | HiPC | A global address space approach to automated data management for parallel Quantum Monte Carlo applications. | Qingpeng Niu, James Dinan, Sravya Tirukkovalur, Lubos Mitas, Lucas K. Wagner, P. Sadayappan |
| 2012 | ICS | High-performance code generation for stencil computations on GPU architectures. | Justin Holewinski, Louis-Nol Pouchet, P. Sadayappan |
| 2012 | PLDI | Dynamic trace-based analysis of vectorization potential of applications. | Justin Holewinski, Ragavendar Ramamurthi, Mahesh Ravishankar, Naznin Fauzia, Louis-Nol Pouchet, Atanas Rountev, P. Sadayappan |
| 2012 | SC | GADBMS: A Framework for Scalable Array Analytics. | Tyler Clemons, Srinivasan Parthasarathy, P. Sadayappan |
| 2012 | SC | Code generation for parallel execution of a class of irregular loops on distributed memory systems. | Mahesh Ravishankar, John Eisenlohr, Louis-Nol Pouchet, J. Ramanujam, Atanas Rountev, P. Sadayappan |
| 2011 | CC | Data Layout Transformation for Stencil Computations on Short-Vector SIMD Architectures. | Thomas Henretty, Kevin Stock, Louis-Nol Pouchet, Franz Franchetti, J. Ramanujam, P. Sadayappan |
| 2011 | CGO | Predictive modeling in a polyhedral optimization space. | Eunjung Park, Louis-Nol Pouchet, John Cavazos, Albert Cohen, P. Sadayappan |
| 2011 | HiPC | Dynamic selection of tile sizes. | Sanket Tavarageri, Louis-Nol Pouchet, J. Ramanujam, Atanas Rountev, P. Sadayappan |
| 2011 | POPL | Loop transformations: convexity, pruning and optimization. | Louis-Nol Pouchet, Uday Bondhugula, Cdric Bastoul, Albert Cohen, J. Ramanujam, P. Sadayappan, Nicolas Vasilache |
| 2010 | CC | Automatic C-to-CUDA Code Generation for Affine Programs. | Muthu Manikandan Baskaran, J. Ramanujam, P. Sadayappan |
| 2010 | CCGRID | Selective Recovery from Failures in a Task Parallel Programming Model. | James Dinan, Arjun Singri, P. Sadayappan, Sriram Krishnamoorthy |
| 2010 | CGO | Parameterized tiling revisited. | Muthu Manikandan Baskaran, Albert Hartono, Sanket Tavarageri, Thomas Henretty, J. Ramanujam, P. Sadayappan |
| 2010 | SC | Combined Iterative and Model-driven Optimization in an Automatic Parallelization Framework. | Louis-Nol Pouchet, Uday Bondhugula, Cdric Bastoul, Albert Cohen, J. Ramanujam, P. Sadayappan |
| 2009 | CLUSTER | Scalable I/O forwarding framework for high-performance computing systems. | Nawab Ali, Philip H. Carns, Kamil Iskra, Dries Kimpe, Samuel Lang, Robert Latham, Robert B. Ross, Lee Ward, P. Sadayappan |
| 2009 | HPDC | An integrated framework for performance-based optimization of scientific workflows. | Vijay S. Kumar, P. Sadayappan, Gaurang Mehta, Karan Vahi, Ewa Deelman, Varun Ratnakar, Jihie Kim, Yolanda Gil, Mary W. Hall, Tahsin M. Kur, Joel H. Saltz |
| 2009 | ICS | Parametric multi-level tiling of imperfectly nested loops. | Albert Hartono, Muthu Manikandan Baskaran, Cdric Bastoul, Albert Cohen, Sriram Krishnamoorthy, Boyana Norris, J. Ramanujam, P. Sadayappan |
| 2009 | PPoPP | Compiler-assisted dynamic scheduling for effective parallelization of loop nests on multicore processors. | Muthu Manikandan Baskaran, Nagavijayalakshmi Vydyanathan, Uday Bondhugula, J. Ramanujam, Atanas Rountev, P. Sadayappan |
| 2009 | SC | Scalable work stealing. | James Dinan, D. Brian Larkins, P. Sadayappan, Sriram Krishnamoorthy, Jarek Nieplocha |
| 2009 | SC | Enabling software management for multicore caches with a lightweight hardware support. | Jiang Lin, Qingda Lu, Xiaoning Ding, Zhao Zhang, Xiaodong Zhang, P. Sadayappan |
| 2008 | CC | Automatic Transformations for Communication-Minimized Parallelization and Locality Optimization in the Polyhedral Model. | Uday Bondhugula, Muthu Manikandan Baskaran, Sriram Krishnamoorthy, J. Ramanujam, Atanas Rountev, P. Sadayappan |
| 2008 | CLUSTER | An OSD-based approach to managing directory operations in parallel file systems. | Nawab Ali, Ananth Devulapalli, Dennis Dalessandro, Pete Wyckoff, P. Sadayappan |
| 2008 | CLUSTER | Are nonblocking networks really needed for high-end-computing workloads? | Narayan Desai, Pavan Balaji, P. Sadayappan, Mohammad Islam |
| 2008 | HPCA | Gaining insights into multicore cache partitioning: Bridging the gap between simulation and real systems. | Jiang Lin, Qingda Lu, Xiaoning Ding, Zhao Zhang, Xiaodong Zhang, P. Sadayappan |
| 2008 | HPDC | Multi-hop path splitting and multi-pathing optimizations for data transfers over shared wide-area networks using gridFTP. | Gaurav Khanna, mit V. atalyrek, Tahsin M. Kur, P. Sadayappan, Joel H. Saltz, Rajkumar Kettimuthu, Ian T. Foster |
| 2008 | ICCS | Integrated Data and Task Management for Scientific Applications. | Jarek Nieplocha, Sriram Krishnamoorthy, Marat Valiev, Manojkumar Krishnan, Bruce J. Palmer, P. Sadayappan |
| 2008 | ICPP | Scioto: A Framework for Global-View Task Parallelism. | James Dinan, Sriram Krishnamoorthy, D. Brian Larkins, Jarek Nieplocha, P. Sadayappan |
| 2008 | ICPP | A Duplication Based Algorithm for Optimizing Latency Under Throughput Constraints for Streaming Workflows. | Nagavijayalakshmi Vydyanathan, mit V. atalyrek, Tahsin M. Kur, P. Sadayappan, Joel H. Saltz |
| 2008 | ICS | A compiler framework for optimization of affine loop nests for gpgpus. | Muthu Manikandan Baskaran, Uday Bondhugula, Sriram Krishnamoorthy, J. Ramanujam, Atanas Rountev, P. Sadayappan |
| 2008 | PLDI | A practical automatic polyhedral parallelizer and locality optimizer. | Uday Bondhugula, Albert Hartono, J. Ramanujam, P. Sadayappan |
| 2008 | PPoPP | Automatic data movement and computation mapping for multi-level parallel architectures with explicitly managed memories. | Muthu Manikandan Baskaran, Uday Bondhugula, Sriram Krishnamoorthy, J. Ramanujam, Atanas Rountev, P. Sadayappan |
| 2008 | SC | Using overlays for efficient data transfer over shared wide-area networks. | Gaurav Khanna, mit V. atalyrek, Tahsin M. Kur, Rajkumar Kettimuthu, P. Sadayappan, Ian T. Foster, Joel H. Saltz |
| 2008 | SC | Global trees: a framework for linked data structures on distributed memory parallel systems. | D. Brian Larkins, James Dinan, Sriram Krishnamoorthy, Srinivasan Parthasarathy, Atanas Rountev, P. Sadayappan |
| 2007 | CLUSTER | Non-collective parallel I/O for global address space programming models. | Sriram Krishnamoorthy, Juan Piernas, Vinod Tipparaju, Jarek Nieplocha, P. Sadayappan |
| 2007 | EuroPar | Scheduling File Transfers for Data-Intensive Jobs on Heterogeneous Clusters. | Gaurav Khanna, mit V. atalyrek, Tahsin M. Kur, P. Sadayappan, Joel H. Saltz |
| 2007 | EuroPar | Toward Optimizing Latency Under Throughput Constraints for Application Workflows on Clusters. | Nagavijayalakshmi Vydyanathan, mit V. atalyrek, Tahsin M. Kur, P. Sadayappan, Joel H. Saltz |
| 2007 | ICPP | Analyzing and Minimizing the Impact of Opportunity Cost in QoS-aware Job Scheduling. | Mohammad Islam, Pavan Balaji, Gerald Sabin, P. Sadayappan |
| 2007 | PLDI | Effective automatic parallelization of stencil computations. | Sriram Krishnamoorthy, Muthu Manikandan Baskaran, Uday Bondhugula, J. Ramanujam, Atanas Rountev, P. Sadayappan |
| 2007 | PPoPP | Automatic mapping of nested loops to FPGAS. | Uday Bondhugula, J. Ramanujam, P. Sadayappan |
| 2007 | SC | Integrating parallel file systems with object-based storage devices. | Ananth Devulapalli, Dennis Dalessandro, Pete Wyckoff, Nawab Ali, P. Sadayappan |
| 2006 | CLUSTER | A Performance Instrumentation Framework to Characterize Computation-Communication Overlap in Message-Passing Systems. | Aniruddha G. Shet, P. Sadayappan, David E. Bernholdt, Jarek Nieplocha, Vinod Tipparaju |
| 2006 | CLUSTER | Locality Conscious Processor Allocation and Scheduling for Mixed Parallel Applications. | Nagavijayalakshmi Vydyanathan, Sriram Krishnamoorthy, Gerald Sabin, mit V. atalyrek, Tahsin M. Kur, P. Sadayappan, Joel H. Saltz |
| 2006 | FCCM | Hardware/Software Integration for FPGA-based All-Pairs Shortest-Paths. | Uday Bondhugula, Ananth Devulapalli, James Dinan, Joseph Fernando, Pete Wyckoff, Eric Stahlberg, P. Sadayappan |
| 2006 | HPDC | Task Scheduling and File Replication for Data-Intensive Jobs with Batch-shared I/O. | Gaurav Khanna, Nagavijayalakshmi Vydyanathan, mit V. atalyrek, Tahsin M. Kur, Sriram Krishnamoorthy, P. Sadayappan, Joel H. Saltz |
| 2006 | ICCS | Identifying Cost-Effective Common Subexpressions to Reduce Operation Count in Tensor Contraction Evaluations. | Albert Hartono, Qingda Lu, Xiaoyang Gao, Sriram Krishnamoorthy, Marcel Nooijen, Gerald Baumgartner, David E. Bernholdt, Venkatesh Choppella, Russell M. Pitzer, J. Ramanujam, Atanas Rountev, P. Sadayappan |
| 2006 | ICPP | An Integrated Approach for Processor Allocation and Scheduling of Mixed-Parallel Applications. | Nagavijayalakshmi Vydyanathan, Sriram Krishnamoorthy, Gerald Sabin, mit V. atalyrek, Tahsin M. Kur, P. Sadayappan, Joel H. Saltz |
| 2006 | JSSPP | A Data Locality Aware Online Scheduling Approach for I/O-Intensive Jobs with File Sharing. | Gaurav Khanna, mit V. atalyrek, Tahsin M. Kur, P. Sadayappan, Joel H. Saltz |
| 2006 | JSSPP | Moldable Parallel Job Scheduling Using Job Efficiency: An Iterative Approach. | Gerald Sabin, Matthew Lang, P. Sadayappan |
| 2006 | SC | Data management and query - Hypergraph partitioning for automatic memory hierarchy management. | Sriram Krishnamoorthy, mit V. atalyrek, Jarek Nieplocha, Atanas Rountev, P. Sadayappan |
| 2006 | SC | M12 - Overview of the global arrays parallel software development toolkit. | Jarek Nieplocha, Bruce J. Palmer, Manojkumar Krishnan, P. Sadayappan |
| 2005 | CCGRID | A hypergraph partitioning based approach for scheduling of tasks with batch-shared I/O. | Gaurav Khanna, Nagavijayalakshmi Vydyanathan, Tahsin M. Kur, mit V. atalyrek, Pete Wyckoff, Joel H. Saltz, P. Sadayappan |
| 2005 | HiPC | Data and Computation Abstractions for Dynamic and Irregular Computations. | Sriram Krishnamoorthy, Jarek Nieplocha, P. Sadayappan |
| 2005 | HPDC | Assessment and enhancement of meta-schedulers for multi-site job sharing. | Gerald Sabin, Vishvesh Sahasrabudhe, P. Sadayappan |
| 2005 | ICCS | Automated Operation Minimization of Tensor Contraction Expressions in Electronic Structure Calculations. | Albert Hartono, Alexander Sibiryakov, Marcel Nooijen, Gerald Baumgartner, David E. Bernholdt, So Hirata, Chi-Chung Lam, Russell M. Pitzer, J. Ramanujam, P. Sadayappan |
| 2005 | JSSPP | Unfairness Metrics for Space-Sharing Parallel Job Schedulers. | Gerald Sabin, P. Sadayappan |
| 2005 | PPoPP | Performance modeling and optimization of parallel out-of-core tensor contractions. | Xiaoyang Gao, Swarup Kumar Sahoo, Chi-Chung Lam, J. Ramanujam, Qingda Lu, Gerald Baumgartner, P. Sadayappan |
| 2005 | SC | Integrated Loop Optimizations for Data Locality Enhancement of Tensor Contraction Expressions. | Swarup Kumar Sahoo, Sriram Krishnamoorthy, Rajkiran Panuganti, P. Sadayappan |
| 2004 | CLUSTER | Towards provision of quality of service guarantees in job scheduling. | Mohammad Islam, Pavan Balaji, P. Sadayappan, Dhabaleswar K. Panda |
| 2004 | CLUSTER | On fairness in distributed job scheduling across multiple sites. | Gerald Sabin, Vishvesh Sahasrabudhe, P. Sadayappan |
| 2004 | HiPC | Efficient Layout Transformation for Disk-Based Multidimensional Arrays. | Sriram Krishnamoorthy, Gerald Baumgartner, Chi-Chung Lam, Jarek Nieplocha, P. Sadayappan |
| 2004 | ICPP | Job Fairness in Non-Preemptive Job Scheduling. | Gerald Sabin, Garima Kochhar, P. Sadayappan |
| 2003 | CLUSTER | Efficient Parallel Out-of-Core Matrix Transposition. | Sriram Krishnamoorthy, Gerald Baumgartner, Daniel Cociorva, Chi-Chung Lam, P. Sadayappan |
| 2003 | CLUSTER | A Robust Scheduling Strategy for Moldable Scheduling of Parallel Jobs. | Sudha Srinivasan, Sriram Krishnamoorthy, P. Sadayappan |
| 2003 | HiPC | Data Locality Optimization for Synthesis of Efficient Out-of-Core Algorithms. | Sandhya Krishnan, Sriram Krishnamoorthy, Gerald Baumgartner, Daniel Cociorva, Chi-Chung Lam, P. Sadayappan, J. Ramanujam, David E. Bernholdt, Venkatesh Choppella |
| 2003 | JSSPP | QoPS: A QoS Based Scheme for Parallel Job Scheduling. | Mohammad Islam, Pavan Balaji, P. Sadayappan, Dhabaleswar K. Panda |
| 2003 | JSSPP | Scheduling of Parallel Jobs in a Heterogeneous Multi-site Environement. | Gerald Sabin, Rajkumar Kettimuthu, Arun Rajan, P. Sadayappan |
| 2002 | CLUSTER | Selective Buddy Allocation for Scheduling Parallel Jobs on Clusters. | Vijay Subramani, Rajkumar Kettimuthu, Srividya Srinivasan, Jeanette Johnston, P. Sadayappan |
| 2002 | HiPC | Effective Selection of Partition Sizes for Moldable Scheduling of Parallel Jobs. | Srividya Srinivasan, Vijay Subramani, Rajkumar Kettimuthu, Praveen Holenarsipur, P. Sadayappan |
| 2002 | HPDC | Distributed Job Scheduling on Computational Grids Using Multiple Simultaneous Requests. | Vijay Subramani, Rajkumar Kettimuthu, Srividya Srinivasan, P. Sadayappan |
| 2002 | ICDCS | A Reliable Multicast Algorithm for Mobile Ad Hoc Networks. | Thiagaraja Gopalsamy, Mukesh Singhal, Dhabaleswar K. Panda, P. Sadayappan |
| 2002 | JSSPP | Selective Reservation Strategies for Backfill Job Scheduling. | Srividya Srinivasan, Rajkumar Kettimuthu, Vijay Subramani, P. Sadayappan |
| 2002 | PLDI | Space-Time Trade-Off Optimization for a Class of Electronic Structure Calculations. | Daniel Cociorva, Gerald Baumgartner, Chi-Chung Lam, P. Sadayappan, J. Ramanujam, Marcel Nooijen, David E. Bernholdt, Robert J. Harrison |
| 2002 | SC | A high-level approach to synthesis of high-performance codes for quantum chemistry. | Gerald Baumgartner, David E. Bernholdt, Daniel Cociorva, Robert J. Harrison, So Hirata, Chi-Chung Lam, Marcel Nooijen, Russell M. Pitzer, J. Ramanujam, P. Sadayappan |
| 2001 | HiPC | Towards Automatic Synthesis of High-Performance Codes for Electronic Structure Calculations: Data Locality Optimization. | Daniel Cociorva, J. W. Wilkins, Gerald Baumgartner, P. Sadayappan, J. Ramanujam, Marcel Nooijen, David E. Bernholdt, Robert J. Harrison |
| 2001 | ICPP | Implementing TreadMarksover VIA on Myrinet and Gigabit Ethernet: Challenges, Design Experience, and Performance Evaluation. | Mohammad Banikazemi, Jiuxing Liu, Dhabaleswar K. Panda, P. Sadayappan |
| 2001 | ICPP | NIC-Based Rate Control for Proportional Bandwidth Allocation in Myrinet Clusters. | Abhishek Gulati, Dhabaleswar K. Panda, P. Sadayappan, Pete Wyckoff |
| 2001 | ICS | Loop optimization for a class of memory-constrained computations. | Daniel Cociorva, J. W. Wilkins, Chi-Chung Lam, Gerald Baumgartner, J. Ramanujam, P. Sadayappan |
| 2000 | HiPC | Characterization and enhancement of Static Mapping Heuristics for Heterogeneous Systems. | Praveen Holenarsipur, Vladimir Yarmolenko, Jos Duato, Dhabaleswar K. Panda, P. Sadayappan |
| 1999 | HCW | Communication Modeling of Heterogeneous Networks of Workstations for Performance Characterization of Collective Operations. | Mohammad Banikazemi, Jayanthi Sampathkumar, Sandeep Prabhu, Dhabaleswar K. Panda, P. Sadayappan |
| 1999 | HiPC | Memory-Optimal Evaluation of Expression Trees Involving Large Objects. | Chi-Chung Lam, Daniel Cociorva, Gerald Baumgartner, P. Sadayappan |
| 1999 | ICPP | An Incremental Methodology for Parallelizing Legacy Stencil Codes on Message-Passing Computers. | N. S. Sundar, S. Jayanthi, P. Sadayappan, Miguel Visbal |
| 1996 | ICS | Hybrid Algorithms for Complete Exchange in 2D Meshes. | N. S. Sundar, Doddaballapur Narasimha-Murthy Jayasimha, Dhabaleswar K. Panda, P. Sadayappan |
| 1994 | ICPADS | Communication-Efficient Implementation of Block Recursive Algorithms on Distributed-Memory Machines. | Sandeep K. S. Gupta, Chua-Huang Huang, Rodney W. Johnson, P. Sadayappan |
| 1994 | ICS | An approach to communication-efficient data redistribution. | S. D. Kaushik, Chua-Huang Huang, Rodney W. Johnson, P. Sadayappan |
| 1994 | ICS | On sparse matrix reordering for parallel factorization. | Bharat Kumar, P. Sadayappan, Chua-Huang Huang |
| 1994 | SC | EXTENT: a portable programming environment for designing and implementing high-performance block recursive algorithms. | Donglai Dai, Sandeep K. S. Gupta, S. D. Kaushik, J. H. Lu, Raj Verdhan Singh, Chua-Huang Huang, P. Sadayappan, Rodney W. Johnson |
| 1994 | SPAA | Communication Efficient Matrix Multiplication on Hypercubes. | Himanshu Gupta, P. Sadayappan |
| 1993 | DAC | Architectural Synthesis of Performance-Driven Multipliers with Accumulator Interleaving. | Debabrata Ghosh, S. K. Nandy, P. Sadayappan, K. Parthasarathy |
| 1993 | ICPP | Compile-Time Characterization of Recurrent Patterns in Irregular Computations. | Kalluri Eswar, P. Sadayappan, Chua-Huang Huang |
| 1993 | ICPP | Supernodal Sparse Cholesky Facotrization on Distributed-Memory Multiprocessors. | Kalluri Eswar, P. Sadayappan, Chua-Huang Huang, V. Visvanathan |
| 1993 | ICPP | On Compiling Array Expressions for Efficient Execution on Distributed-Memory Machines. | Sandeep K. S. Gupta, S. D. Kaushik, S. Mufti, Sanjay Sharma, Chua-Huang Huang, P. Sadayappan |
| 1993 | ICPP | A Parallel Progressive Refinement Image Rendering Algorithm on a Scalable Multithreaded VLSI Processor Array. | S. K. Nandy, Ranjani Narayan, V. Visvanathan, P. Sadayappan, Prashant S. Chauhan |
| 1993 | SC | Efficient transposition algorithms for large matrices. | S. D. Kaushik, Chua-Huang Huang, John R. Johnson, Rodney W. Johnson, P. Sadayappan |
| 1992 | SC | An Algebraic Theory for Modeling Direct Interconnection Networks. | S. D. Kaushik, Sanjay Sharma, Chua-Huang Huang, Jeremy R. Johnson, Rodney W. Johnson, P. Sadayappan |
| 1991 | ICPP | Multifrontal Factorization of Sparse Matrices on Shared-Memory Multiprocessors. | Kalluri Eswar, P. Sadayappan, V. Visvanathan |
| 1991 | ICPP | Computer Graphics Rendering on a Shared Memory Multiprocessor. | Scott Whitman, P. Sadayappan |
| 1991 | PPoPP | Removal of Redundant Dependences in DOACROSS Lops with Constant Dependences. | V. Prasad Krothapalli, P. Sadayappan |
| 1991 | SC | Tiling multidimensional iteration spaces for nonshared memory machines. | J. Ramanujam, P. Sadayappan |
| 1990 | ICPP | Tiling of Iteration Spaces for Multicomputers. | J. Ramanujam, P. Sadayappan |
| 1989 | DAC | Efficient Sparse Matrix Factorization for Circuit Simulation on Vector Supercomputers. | P. Sadayappan, V. Visvanathan |
| 1989 | ICPP | Optimal Static Scheduling of Sequential Loops on Multiprocessors. | Amr Zaky, P. Sadayappan |
| 1989 | ICS | One-to-one mapping of process graphs onto a hypercube. | Fikret Eral, P. Sadayappan |
| 1989 | SC | A methodology for parallelizing programs for multicomputers and complex memory multiprocessors. | J. Ramanujam, P. Sadayappan |
| 1988 | ICCD | Comparative analysis of approaches to hardware acceleration for sparse-matrix factorization. | P. Sadayappan, V. Visvanathan |
| 1988 | ICRA | A VLSI robotics vector processor for real-time control. | Yong-Long Calvin Ling, P. Sadayappan, Karl W. Olson, David E. Orin |
| 1988 | ICS | An approach to synchronization for parallel computing. | V. Prasad Krothapalli, P. Sadayappan |
| 1988 | ICS | Parallelization and performance evaluation of circuit simulation on a shared-memory multiprocessor. | P. Sadayappan, V. Visvanathan |
| 1987 | ICPP | Mapping Finite Element Graphs onto Processor Meshes. | P. Sadayappan, Fikret Eral, Steven Martin |
| 1987 | ICS | Cluster-Partitioning Approaches to Mapping Parallel Programs onto a Hypercube. | P. Sadayappan, Fikret Eral |
| 1985 | DAC | Modeling switch-level simulation using data flow. | V. Ashok, Roger L. Costello, P. Sadayappan |