| 2009 | Solving dense linear systems on platforms with multiple hardware accelerators. | Gregorio Quintana-Ort, Francisco D. Igual, Enrique S. Quintana-Ort, Robert A. van de Geijn |
| 2009 | Multi-core demands multi-interfaces. | Yale N. Patt |
| 2009 | Techniques for efficient placement of synchronization primitives. | Alexandru Nicolau, Guangqiang Li, Arun Kejariwal |
| 2009 | Idempotent work stealing. | Maged M. Michael, Martin T. Vechev, Vijay A. Saraswat |
| 2009 | Towards concurrency refactoring for x10. | Shane Markstrum, Robert M. Fuhrer, Todd D. Millstein |
| 2009 | A compiler and runtime system for enabling data mining applications on gpus. | Wenjing Ma, Gagan Agrawal |
| 2009 | Architectural support for cilk computations on many-core architectures. | Guoping Long, Dongrui Fan, Junchao Zhang |
| 2009 | Efficient and scalable multiprocessor fair scheduling using distributed weighted round-robin. | Tong Li, Dan P. Baumberger, Scott Hahn |
| 2009 | OpenMP to GPGPU: a compiler framework for automatic translation and optimization. | Seyong Lee, Seung-Jai Min, Rudolf Eigenmann |
| 2009 | Turbocharging boosted transactions or: how i learnt to stop worrying and love longer transactions. | Chinmay Eishan Kulkarni, Osman S. Unsal, Adrin Cristal, Eduard Ayguad, Mateo Valero |
| 2009 | How much parallelism is there in irregular applications? | Milind Kulkarni, Martin Burtscher, Rajasekhar Inkulu, Keshav Pingali, Calin Cascaval |
| 2009 | Petascale computing with accelerators. | Michael Kistler, John A. Gunnels, Daniel A. Brokenshire, Brad Benton |
| 2009 | Parallelization spectroscopy: analysis of thread-level parallelism in hpc programs. | Arun Kejariwal, Calin Cascaval |
| 2009 | An efficient transactional memory algorithm for computing minimum spanning forest of sparse graphs. | Seunghwa Kang, David A. Bader |
| 2009 | Exploiting global optimizations for openmp programs in the openuh compiler. | Lei Huang, Deepak Eachempati, Marcus W. Hervey, Barbara M. Chapman |
| 2009 | Backtracking-based load balancing. | Tasuku Hiraishi, Masahiro Yasugi, Seiji Umatani, Taiichi Yuasa |
| 2009 | Opportunities beyond single-core microprocessors. | Mark D. Hill |
| 2009 | Preliminary results on nb-feb, a synchronization primitive for parallel programming. | Phuong Hoai Ha, Philippas Tsigas, Otto J. Anshus |
| 2009 | How to build programmable multi-core chips. | Jack B. Dennis |
| 2009 | Software transactional distributed shared memory. | Alokika Dash, Brian Demsky |
| 2009 | Parallel thinking. | Guy E. Blelloch |
| 2009 | Efficient, portable implementation of asynchronous multi-place programs. | Ganesh Bikshandi, Jos G. Castaos, Sreedhar B. Kodali, V. Krishna Nandivada, Igor Peshansky, Vijay A. Saraswat, Sayantan Sur, Pradeep Varma, Tong Wen |
| 2009 | Topology aware task mapping techniques: an api and case study. | Abhinav Bhatele, Eric J. Bohm, Laxmikant V. Kal |
| 2009 | Compiler-assisted dynamic scheduling for effective parallelization of loop nests on multicore processors. | Muthu Manikandan Baskaran, Nagavijayalakshmi Vydyanathan, Uday Bondhugula, J. Ramanujam, Atanas Rountev, P. Sadayappan |
| 2009 | Serialization sets: a dynamic dependence-based parallel execution model. | Matthew D. Allen, Srinath Sridharan, Gurindar S. Sohi |