Naoya Maruyama
Publication record assembled from the DBLP archive of ranked conferences.
Papers indexed
39
Venues
13
Active years
2006–2019
Best venue rank
A
Where they publish
Papers
39 indexed papers, newest first.
| Year | Venue | Title | Authors |
|---|---|---|---|
| 2019 | SC | Channel and filter parallelism for large-scale CNN training. | Nikoli Dryden, Naoya Maruyama, Tim Moon, Tom Benson, Marc Snir, Brian Van Essen |
| 2019 | SC | Preparation and optimization of a diverse workload for a large-scale heterogeneous system. | Ian Karlin, Yoonho Park, Bronis R. de Supinski, Peng Wang, Bert Still, David Beckingsale, Robert Blake, Tong Chen, Guojing Cong, Carlos H. A. Costa, Johann Dahm, Giacomo Domeniconi, Thomas Epperly, Aaron Fisher, Sara Kokkila Schumacher, Steven H. Langer, Hai Le, Eun Kyung Lee, Naoya Maruyama, Xinyu Que, David F. Richards, Bjrn Sjgreen, Jonathan Wong, Carol S. Woodward, Ulrike Meier Yang, Xiaohua Zhang, Bob Anderson, David Appelhans, Levi Barnes, Peter D. Barnes Jr., Sorin Bastea, David Bhme, Jamie A. Bramwell, James M. Brase, Jos R. Brunheroto, Barry Chen, Charway R. Cooper, Tony Degroot, Robert D. Falgout, Todd Gamblin, David J. Gardner, James N. Glosli, John A. Gunnels, Max P. Katz, Tzanio V. Kolev, I-Feng W. Kuo, Matthew P. LeGendre, Ruipeng Li, Pei-Hung Lin, Shelby Lockhart, Kathleen McCandless, Claudia Misale, Jaime H. Moreno, Rob Neely, Jarom Nelson, Rao Nimmakayala, Kathryn M. O'Brien, Kevin O'Brien, Ramesh Pankajakshan, Roger Pearce, Slaven Peles, Phil Regier, Steven C. Rennich, Martin Schulz, Howard Scott, James C. Sexton, Kathleen Shoga, Shiv Sundram, Guillaume Thomas-Collignon, Brian Van Essen, Alexey Voronin, Bob Walkup, Lu Wang, Chris Ward, Hui-Fang Wen, Daniel A. White, Christopher Young, Cyril Zeller, Edward Zywicz |
| 2018 | IJCNN | Effective Quantization Approaches for Recurrent Neural Networks. | Md. Zahangir Alom, Adam T. Moody, Naoya Maruyama, Brian C. Van Essen, Tarek M. Taha |
| 2017 | FPL | Evaluating high-level design strategies on FPGAs for high-performance computing. | Artur Podobas, Hamid Reza Zohouri, Naoya Maruyama, Satoshi Matsuoka |
| 2017 | FPL | Evaluating high-level design strategies on FPGAs for high-performance computing. | Artur Podobas, Hamid Reza Zohouri, Naoya Maruyama, Satoshi Matsuoka |
| 2017 | ICPP | Optimizations of Two Compute-Bound Scientific Kernels on the SW26010 Many-Core Processor. | James Lin, Zhigeng Xu, Akira Nukada, Naoya Maruyama, Satoshi Matsuoka |
| 2016 | HPCC | A Directive-Based Data Layout Abstraction for Performance Portability of OpenACC Applications. | Tetsuya Hoshino, Naoya Maruyama, Satoshi Matsuoka |
| 2016 | ICPADS | Tapas: An Implicitly Parallel Programming Framework for Hierarchical N-Body Algorithms. | Keisuke Fukuda, Motohiko Matsuda, Naoya Maruyama, Rio Yokota, Kenjiro Taura, Satoshi Matsuoka |
| 2016 | SC | Daino: a high-level framework for parallel and efficient AMR on GPUs. | Mohamed Wahib, Naoya Maruyama, Takayuki Aoki |
| 2016 | SC | Evaluating and optimizing OpenCL kernels for high performance computing with FPGAs. | Hamid Reza Zohouri, Naoya Maruyama, Aaron Smith, Motohiko Matsuda, Satoshi Matsuoka |
| 2015 | HPDC | Automated GPU Kernel Transformations in Large-Scale Production Stencil Applications. | Mohamed Wahib, Naoya Maruyama |
| 2015 | SC | Data-centric GPU-based adaptive mesh refinement. | Mohamed Wahib, Naoya Maruyama |
| 2014 | CCGRID | A User-Level InfiniBand-Based File System and Checkpoint Strategy for Burst Buffers. | Kento Sato, Kathryn Mohror, Adam Moody, Todd Gamblin, Bronis R. de Supinski, Naoya Maruyama, Satoshi Matsuoka |
| 2014 | SC | An OpenACC extension for data layout transformation. | Tetsuya Hoshino, Naoya Maruyama, Satoshi Matsuoka |
| 2014 | SC | Scalable Kernel Fusion for Memory-Bound GPU Applications. | Mohamed Wahib, Naoya Maruyama |
| 2013 | CCGRID | CUDA vs OpenACC: Performance Case Studies with Kernel Benchmarks and a Memory-Bound CFD Application. | Tetsuya Hoshino, Naoya Maruyama, Satoshi Matsuoka, Ryoji Takaki |
| 2013 | CLUSTER | K MapReduce: A scalable tool for data-processing and search/ensemble applications on large-scale supercomputers. | Motohiko Matsuda, Naoya Maruyama, Shin'ichiro Takizawa |
| 2013 | CLUSTER | Highly optimized full GPU-acceleration of non-hydrostatic weather model SCALE-LES. | Mohamed Wahib, Naoya Maruyama |
| 2013 | EuroPar | Topic 15: GPU and Accelerator Computing - (Introduction). | Naoya Maruyama, Leif Kobbelt, Pavan Balaji, Nikola Puzovic, Samuel Thibault, Kun Zhou |
| 2013 | ICPP | Integrating Multi-GPU Execution in an OpenACC Compiler. | Toshiya Komoda, Shinobu Miwa, Hiroshi Nakamura, Naoya Maruyama |
| 2012 | CCGRID | Design and Implementation of Portable and Efficient Non-blocking Collective Communication. | Akihiro Nomura, Yutaka Ishikawa, Naoya Maruyama, Satoshi Matsuoka |
| 2012 | CLUSTER | Hierarchical Clustering Strategies for Fault Tolerance in Large Scale HPC Systems. | Leonardo Arturo Bautista-Gomez, Thomas Ropars, Naoya Maruyama, Franck Cappello, Satoshi Matsuoka |
| 2012 | EuroPar | Scalable Reed-Solomon-Based Reliable Local Storage for HPC Applications on IaaS Clouds. | Leonardo Arturo Bautista-Gomez, Bogdan Nicolae, Naoya Maruyama, Franck Cappello, Satoshi Matsuoka |
| 2012 | EuroPar | Multi-GPU Implementation of the NICAM Atmospheric Model. | Irina Demeshko, Naoya Maruyama, Hirofumi Tomita, Satoshi Matsuoka |
| 2012 | SC | Design and modeling of a non-blocking checkpointing system. | Kento Sato, Naoya Maruyama, Kathryn Mohror, Adam Moody, Todd Gamblin, Bronis R. de Supinski, Satoshi Matsuoka |
| 2012 | SC | A Task Parallel Implementation of Fast Multipole Methods. | Kenjiro Taura, Jun Nakashima, Rio Yokota, Naoya Maruyama |
| 2011 | SC | FTI: high performance fault tolerance interface for hybrid systems. | Leonardo Arturo Bautista-Gomez, Seiji Tsuboi, Dimitri Komatitsch, Franck Cappello, Naoya Maruyama, Satoshi Matsuoka |
| 2011 | SC | Poster: fast GPU read alignment with burrows wheeler transform based index. | Aleksandr Drozd, Naoya Maruyama, Satoshi Matsuoka |
| 2011 | SC | Physis: an implicitly parallel programming model for stencil computations on large-scale GPU-accelerated supercomputers. | Naoya Maruyama, Tatsuo Nomura, Kento Sato, Satoshi Matsuoka |
| 2011 | SC | Peta-scale phase-field simulation for dendritic solidification on the TSUBAME 2.0 supercomputer. | Takashi Shimokawabe, Takayuki Aoki, Tomohiro Takaki, Toshio Endo, Akinori Yamanaka, Naoya Maruyama, Akira Nukada, Satoshi Matsuoka |
| 2011 | SYSTOR | An exact algorithm for energy-efficient acceleration of task trees on CPU/GPU architectures. | Mark Silberstein, Naoya Maruyama |
| 2010 | CCGRID | Distributed Diskless Checkpoint for Large Scale Systems. | Leonardo Arturo Bautista-Gomez, Naoya Maruyama, Franck Cappello, Satoshi Matsuoka |
| 2010 | HiPC | Low-overhead diskless checkpoint for hybrid computing systems. | Leonardo Arturo Bautista-Gomez, Akira Nukada, Naoya Maruyama, Franck Cappello, Satoshi Matsuoka |
| 2010 | SC | An 80-Fold Speedup, 15.0 TFlops Full GPU Acceleration of Non-Hydrostatic Weather Model ASUCA Production Code. | Takashi Shimokawabe, Takayuki Aoki, Chiashi Muroi, Junichi Ishida, Kohei Kawano, Toshio Endo, Akira Nukada, Naoya Maruyama, Satoshi Matsuoka |
| 2009 | CCGRID | Adaptive Resource Indexing Technique for Unstructured Peer-to-Peer Networks. | Sumeth Lerthirunwong, Naoya Maruyama, Satoshi Matsuoka |
| 2007 | CCGRID | Virtual Clusters on the Fly - Fast, Scalable, and Flexible Installation. | Hideo Nishimura, Naoya Maruyama, Satoshi Matsuoka |
| 2007 | SC | Model-based resource selection for efficient virtual cluster deployment. | Shohei Yamasaki, Naoya Maruyama, Satoshi Matsuoka |
| 2006 | ISPA | Making Wide-Area, Multi-site MPI Feasible Using Xen VM. | Masaki Tatezono, Naoya Maruyama, Satoshi Matsuoka |
| 2006 | SC | Scalable systems software - Problem diagnosis in large-scale computing environments. | Alexander V. Mirgorodskiy, Naoya Maruyama, Barton P. Miller |