| 2013 | Exploiting data parallelism in the yConvex hypergraph algorithm for image representation using GPGPUs. | Saurabh Jha, Tejaswi Agarwal, B. Rajesh Kanna |
| 2013 | Efficient scheduling of recursive control flow on GPUs. | Xin Huo, Sriram Krishnamoorthy, Gagan Agrawal |
| 2013 | An early prototype of an autonomic performance environment for exascale. | Kevin A. Huck, Sameer Shende, Allen D. Malony, Hartmut Kaiser, Allan Porterfield, Robert J. Fowler, Ron Brightwell |
| 2013 | Network-on-chip for a partially reconfigurable FPGA system. | Justin A. Hogan, Raymond J. Weber, Brock J. LaMeres, Todd Kaiser |
| 2013 | A stencil compiler for short-vector SIMD architectures. | Thomas Henretty, Richard Veras, Franz Franchetti, Louis-Nol Pouchet, J. Ramanujam, P. Sadayappan |
| 2013 | MIC-RO: enabling efficient remote offload on heterogeneous many integrated core (MIC) clusters with InfiniBand. | Khaled Hamidouche, Sreeram Potluri, Hari Subramoni, Krishna Chaitanya Kandalla, Dhabaleswar K. Panda |
| 2013 | Toward a scalable multi-GPU eigensolver via compute-intensive kernels and efficient communication. | Azzam Haidar, Mark Gates, Stanimire Tomov, Jack J. Dongarra |
| 2013 | LibWater: heterogeneous distributed computing made easy. | Ivan Grasso, Simone Pellegrini, Biagio Cosenza, Thomas Fahringer |
| 2013 | Massively parallel loading. | Wolfgang Frings, Dong H. Ahn, Matthew P. LeGendre, Todd Gamblin, Bronis R. de Supinski, Felix Wolf |
| 2013 | Multi-layered unstructured mesh generation. | Panagiotis A. Foteinos, Daming Feng, Andrey N. Chernikov, Nikos Chrisochoides |
| 2013 | High quality real-time image-to-mesh conversion for finite element simulations. | Panagiotis A. Foteinos, Nikos Chrisochoides |
| 2013 | Conservative row activation to improve memory power efficiency. | Kun Fang, Zhichun Zhu |
| 2013 | Expressing graph algorithms using generalized active messages. | Nicholas Gerard Edmonds, Jeremiah Willcock, Andrew Lumsdaine |
| 2013 | A massively parallel domain decomposition method for large-scale DFT electronic structure calculations. | Truong Vinh Truong Duy, Taisuke Ozaki |
| 2013 | A decomposition method with minimal communication volume for parallelization of multi-dimensional FFTs. | Truong Vinh Truong Duy, Taisuke Ozaki |
| 2013 | MAD7: a memory architecture simulator targeted at design space exploration. | Hadrien A. Clarke, Antoine Trouv, Kazuaki J. Murakami |
| 2013 | FASTER run-time reconfiguration management. | Catalin Bogdan Ciobanu, Dionisios N. Pnevmatikatos, Kyprianos D. Papadimitriou, Georgi Nedeltchev Gaydadjiev |
| 2013 | Active disk meets flash: a case for intelligent SSDs. | Sangyeun Cho, Chanik Park, Hyunok Oh, Sungchan Kim, Youngmin Yi, Gregory R. Ganger |
| 2013 | Imbalance optimization in scientific workflows. | Weiwei Chen, Ewa Deelman, Rizos Sakellariou |
| 2013 | Data deduplication in a hybrid architecture for improving write performance. | Chao Chen, Jonathan Bastnagel, Yong Chen |
| 2013 | Implementing OmpSs support for regions of data in architectures with multiple address spaces. | Javier Bueno, Xavier Martorell, Rosa M. Badia, Eduard Ayguad, Jess Labarta |
| 2013 | Hobbes: composition and virtualization as the foundations of an extreme-scale OS/R. | Ron Brightwell, Ron A. Oldfield, Arthur B. Maccabe, David E. Bernholdt |
| 2013 | Business meets supercomputing: keynote talk. | Bob Blainey |
| 2013 | Improving numerical accuracy for non-negative matrix multiplication on GPUs using recursive algorithms. | Matthew Badin, Paolo D'Alberto, Lubomir Bic, Michael B. Dillencourt, Alexandru Nicolau |
| 2013 | TEAPOT: a toolset for evaluating performance, power and image quality on mobile graphics systems. | Jos-Mara Arnau, Joan-Manuel Parcerisa, Polychronis Xekalakis |