| 2011 | GVT algorithms and discrete event dynamics on 129K+ processor cores. | Kalyan S. Perumalla, Alfred J. Park, Vinod Tipparaju |
| 2011 | Scalable clustering using multiple GPUs. | K. Wasif Mohiuddin, P. J. Narayanan |
| 2011 | Reliable and randomized data distribution strategies for large scale storage systems. | Alberto Miranda, Sascha Effert, Yangwook Kang, Ethan L. Miller, Andr Brinkmann, Toni Cortes |
| 2011 | Multi-threaded UPC runtime with network endpoints: Design alternatives and evaluation on multi-core architectures. | Miao Luo, Jithin Jose, Sayantan Sur, Dhabaleswar K. Panda |
| 2011 | Implementing a hybrid SRAM / eDRAM NUCA architecture. | Javier Lira, Carlos Molina, David M. Brooks, Antonio Gonzlez |
| 2011 | Increasing the energy efficiency of TLS systems using intermediate checkpointing. | Salman Khan, Nikolas Ioannou, Polychronis Xekalakis, Marcelo Cintra |
| 2011 | Weighted locality-sensitive scheduling for mitigating noise on multi-core clusters. | Vivek Kale, Abhinav Bhatele, William D. Gropp |
| 2011 | Partial globalization of partitioned address spaces for zero-copy communication with shared memory. | Fangzhou Jiao, Nilesh Mahajan, Jeremiah Willcock, Arun Chauhan, Andrew Lumsdaine |
| 2011 | Optimizing multicore performance with message driven execution: A case study. | Pritish Jetley, Laxmikant V. Kal |
| 2011 | Porting irregular reductions on heterogeneous CPU-GPU configurations. | Xin Huo, Vignesh T. Ravi, Gagan Agrawal |
| 2011 | Modelling and analyzing the authorization and execution of video workflows. | Ligang He, Chenlin Huang, Kenli Li, Hao Chen, Jianhua Sun, Bo Gao, Kewei Duan, Stephen A. Jarvis |
| 2011 | Compute & memory optimizations for high-quality speech recognition on low-end GPU processors. | Kshitij Gupta, John D. Owens |
| 2011 | Robust thread-level speculation. | lvaro Garca-Ygez, Diego R. Llanos Ferraris, Arturo Gonzlez-Escribano |
| 2011 | Supporting computational data model representation with high-performance I/O in parallel netCDF. | Kui Gao, Chen Jin, Alok N. Choudhary, Wei-keng Liao |
| 2011 | Highly scalable barriers for future high-performance computing clusters. | Holger Frning, Alexander Giese, Hctor Montaner, Federico Silla, Jos Duato |
| 2011 | A multiresolution data model for improving simulation I/O performance. | Andrew Foulks, R. Daniel Bergeron |
| 2011 | Parallel multiple precision division by a single precision divisor. | Niall Emmart, Charles C. Weems |
| 2011 | Adaptive memory power management techniques for HPC workloads. | Karthik Elangovan, Ivan Rodero, Manish Parashar, Francesc Guim, Isaac Hernandez |
| 2011 | Enabling CUDA acceleration within virtual machines using rCUDA. | Jos Duato, Antonio J. Pea, Federico Silla, Juan Carlos Fernndez, Rafael Mayo, Enrique S. Quintana-Ort |
| 2011 | High-level template for the task-based parallel wavefront pattern. | Antonio J. Dios, Rafael Asenjo, Angeles G. Navarro, Francisco Corbera, Emilio L. Zapata |
| 2011 | Hybrid implementation of error diffusion dithering. | Aditya Deshpande, Ishan Misra, P. J. Narayanan |
| 2011 | Coordination mechanisms for selfish multi-organization scheduling. | Johanne Cohen, Daniel Cordeiro, Denis Trystram, Frdric Wagner |
| 2011 | Maximizing throughput of jobs with multiple resource requirements. | Venkatesan T. Chakaravarthy, Sambuddha Roy, Yogish Sabharwal, Neha Sengupta |
| 2011 | A machine learning-based approach for thread mapping on transactional memory applications. | Mrcio Castro, Lus Fabrcio Wanderley Ges, Christiane Pousa Ribeiro, Murray Cole, Marcelo Cintra, Jean-Franois Mhaut |
| 2011 | STEAMEngine: Driving MapReduce provisioning in the cloud. | Michael Cardosa, Piyush Narang, Abhishek Chandra, Himabindu Pucha, Aameek Singh |