| 2025 | SC | Modelling Load Imbalance In Shared Memory Multicore Systems. | Johannes Langguth, James D. Trotter, Xing Cai |
| 2025 | SC | CPU- and GPU-initiated Communication Strategies for Conjugate Gradient Methods on Large GPU Clusters. | James D. Trotter, Sinan Ekmekibasi, Dogan Sagbili, Johannes Langguth, Xing Cai, Didem Unat |
| 2022 | ICRA | DKNAS: A Practical Deep Keypoint Extraction Framework Based on Neural Architecture Search. | Li Liu, Xing Cai, Ge Li, Thomas H. Li |
| 2021 | HiPC | iPUG for Multiple Graphcore IPUs: Optimizing Performance and Scalability of Parallel Breadth-First Search. | Luk Burchard, Xing Cai, Johannes Langguth |
| 2020 | ICTAI | Towards Loss Balance and Consistent Model in Self-supervised Monocular Depth Estimation. | Chengyuan Li, Lanqing Zhang, Xing Cai, Keyao Li, Ge Li, Thomas H. Li |
| 2019 | ICCS | Combining Algorithmic Rethinking and AVX-512 Intrinsics for Efficient Simulation of Subcellular Calcium Signaling. | Chad Jarvis, Glenn Terje Lines, Johannes Langguth, Kengo Nakajima, Xing Cai |
| 2018 | ACCV | SingleGAN: Image-to-Image Translation by a Single-Generator Network Using Multiple Generative Adversarial Learning. | Xiaoming Yu, Xing Cai, Zhenqiang Ying, Thomas H. Li, Ge Li |
| 2018 | ICPADS | Memory Bandwidth Contention: Communication vs Computation Tradeoffs in Supercomputers with Multicore Architectures. | Johannes Langguth, Xing Cai, Mohammed Sourouri |
| 2016 | DSD | The EMC2 Project on Embedded Microcontrollers: Technical Progress after Two Years. | Werner Weber, Alfred Hoess, Jan van Deventer, Frank Oppenheimer, Rolf Ernst, Adam Kostrzewa, Philippe Dore, Thierry Goubier, Haris Isakovic, Norbert Druml, Egon Wuchner, Daniel Schneider, Erwin Schoitsch, Eric Armengaud, Thomas Soderqvist, Massimo Traversone, Sascha Uhrig, Juan-Carlos Perez-Cortes, Sergio Sez, Juha Kuusela, Mark van Helvoort, Xing Cai, Bjrn Nordmoen, Geir Yngve Paulsen, Hans Petter Dahle, Michael Geissel, Jrgen Salecker, Peter Tummeltshammer |
| 2016 | ICPADS | Enabling Tissue-Scale Cardiac Simulations Using Heterogeneous Computing on Tianhe-2. | Johannes Langguth, Qiang Lan, Namit Gaur, Xing Cai, Mei Wen, Chunyuan Zhang |
| 2015 | ICA3PP | Towards Detailed Tissue-Scale 3D Simulations of Electrical Activity and Calcium Handling in the Human Cardiac Ventricle. | Qiang Lan, Namit Gaur, Johannes Langguth, Xing Cai |
| 2015 | ICCS | Multi-GPU Implementations of Parallel 3D Sweeping Algorithms with Application to Geological Folding. | Ezhilmathi Krishnasamy, Mohammed Sourouri, Xing Cai |
| 2014 | EuroPar | Automated Transformation of GPU-Specific OpenCL Kernels Targeting Performance Portability on Multi-Core/Many-Core CPUs. | Dafei Huang, Mei Wen, Changqing Xun, Dong Chen, Xing Cai, Yuran Qiao, Nan Wu, Chunyuan Zhang |
| 2014 | ICA3PP | Utilizing Multiple Xeon Phi Coprocessors on One Compute Node. | Xinnan Dong, Jun Chai, Jing Yang, Mei Wen, Nan Wu, Xing Cai, Chunyuan Zhang, Zhaoyun Chen |
| 2014 | ICPADS | Heterogeneous CPU-GPU computing for the finite volume method on 3D unstructured meshes. | Johannes Langguth, Xing Cai |
| 2014 | ICPADS | Effective multi-GPU communication using multiple CUDA streams and threads. | Mohammed Sourouri, Tor Gillberg, Scott B. Baden, Xing Cai |
| 2013 | ICCS | Performance of Sediment Transport Simulations on NVIDIA's Kepler Architecture. | Huayou Su, Nan Wu, Mei Wen, Chunyuan Zhang, Xing Cai |
| 2013 | ICPADS | On the GPU-CPU Performance Portability of OpenCL for 3D Stencil Computations. | Huayou Su, Nan Wu, Mei Wen, Chunyuan Zhang, Xing Cai |
| 2013 | SC | On the GPU performance of cell-centered finite volume method over unstructured tetrahedral meshes. | Johannes Langguth, Nan Wu, Jun Chai, Xing Cai |
| 2012 | CLUSTER | Using 1000+ GPUs and 10000+ CPUs for Sedimentary Basin Simulations. | Mei Wen, Huayou Su, Wenjie Wei, Nan Wu, Xing Cai, Chunyuan Zhang |
| 2011 | ICS | Mint: realizing CUDA performance in 3D stencil methods with annotated C. | Didem Unat, Xing Cai, Scott B. Baden |
| 2010 | ISPDC | Numerical Analysis of a Dual-Sediment Transport Model Applied to Lake Okeechobee, Florida. | Stuart R. Clark, Wenjie Wei, Xing Cai |
| 2006 | HPSC | On the Efficiency of Python for High-Performance Computing: A Case Study Involving Stencil Updates for Partial Differential Equations. | Hans Petter Langtangen, Xing Cai |
| 2002 | ICCS | Parallel Iterative Methods in Modern Physical Applications. | Xing Cai, Yousef Saad, Masha Sosonkina |
| 2000 | CLUSTER | Parallel Simulation of 3D Nonlinear Acoustic Fields on a Linux-cluster. | Xing Cai, smund degrd |