| 2013 | ASPLOS | Improving GPGPU concurrency with elastic kernels. | Sreepathi Pai, Matthew J. Thazhuthaveetil, R. Govindarajan |
| 2012 | EuroPar | CUDA-For-Clusters: A System for Efficient Execution of CUDA Kernels on Multi-core Clusters. | Raghu Prabhakar, R. Govindarajan, Matthew J. Thazhuthaveetil |
| 2009 | CGO | Software Pipelined Execution of Stream Programs on GPUs. | Abhishek Udupa, R. Govindarajan, Matthew J. Thazhuthaveetil |
| 2007 | CGO | Microarchitecture Sensitive Empirical Models for Compiler Optimizations. | Kapil Vaswani, Matthew J. Thazhuthaveetil, Y. N. Srikant, P. J. Joseph |
| 2006 | HPCA | Construction and use of linear regression models for processor performance analysis. | P. J. Joseph, Kapil Vaswani, Matthew J. Thazhuthaveetil |
| 2006 | MICRO | A Predictive Performance Model for Superscalar Processors. | P. J. Joseph, Kapil Vaswani, Matthew J. Thazhuthaveetil |
| 2005 | CGO | A Programmable Hardware Path Profiler. | Kapil Vaswani, Matthew J. Thazhuthaveetil, Y. N. Srikant |
| 2005 | HiPC | Offloading Bloom Filter Operations to Network Processor for Parallel Query Processing in Cluster of Workstations. | V. Santhosh Kumar, Matthew J. Thazhuthaveetil, R. Govindarajan |