| 2026 | KDD | A Multi-Stage Structural Captioning Framework for Enhancing Chinese Image-to-Video Generation in Baidu. | Lei Shen, Zhipeng Jin, Xiawei Li, Wen Tao, Shiyuan Li, Yi Yang, Cong Han, Shuanglong Li, Zhongmin Cai, Lin Liu |
| 2025 | HPCC | Remote Sensing Bathymetric Inversion Research in Shallow Seas Based on WorldView-2. | Lingyun Jiang, Cong Han, Yu Qiu, Weiwei Zhang, Jungui Zhang, Yijun Xiong, Kai Zhao, Zhongliang Zeng |
| 2025 | ICASSP | Dual-path Mamba: Short and Long-term Bidirectional Selective Structured State Space Models for Speech Separation. | Xilin Jiang, Cong Han, Nima Mesgarani |
| 2025 | ICASSP | Speech Slytherin: Examining the Performance and Efficiency of Mamba for Speech Separation, Recognition, and Synthesis. | Xilin Jiang, Yinghao Aaron Li, Adrian Nicolas Florea, Cong Han, Nima Mesgarani |
| 2025 | ICCV | UniGlyph: Unified Segmentation-Conditioned Diffusion for Precise Visual Text Synthesis. | Yuanrui Wang, Cong Han, Yafei Li, Zhipeng Jin, Xiawei Li, SiNan Du, Wen Tao, Shuanglong Li, Yi Yang, Chun Yuan, Liu Lin |
| 2025 | INDIN | Research Progress of Artificial Intelligence Technology in Livestock Pose Estimation. | Cong Han, Shilei Wei, Yangyang Guo, Xiaoping Huang |
| 2025 | KDD | Large Vison-Language Foundation Model in Baidu AIGC Image Advertising. | Zhipeng Jin, Wen Tao, Yafei Li, Yi Yang, Cong Han, Shuanglong Li, Lin Liu |
| 2025 | NAACL | StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion. | Yinghao Aaron Li, Xilin Jiang, Cong Han, Nima Mesgarani |
| 2025 | SIGIR | Retrieval-Augmented Image Captioning and Generation with Entity Concepts Enhancement for Baidu Multimodal Advertising. | Lei Shen, Kang Zhao, Zhipeng Jin, Wen Tao, Yi Yang, Cong Han, Shuanglong Li, Zhongmin Cai, Lin Liu |
| 2024 | CIKM | Scaling Vison-Language Foundation Model to 12 Billion Parameters in Baidu Dynamic Image Advertising. | Xinyu Zhao, Kang Zhao, Zhipeng Jin, Yi Yang, Wen Tao, Xiaodong Chen, Cong Han, Shuanglong Li, Lin Liu |
| 2024 | ICASSP | Unsupervised Multi-Channel Separation And Adaptation. | Cong Han, Kevin W. Wilson, Scott Wisdom, John R. Hershey |
| 2024 | ICASSP | Exploring Self-supervised Contrastive Learning of Spatial Sound Event Representation. | Xilin Jiang, Cong Han, Yinghao Aaron Li, Nima Mesgarani |
| 2024 | SIGIR | Enhancing Baidu Multimodal Advertisement with Chinese Text-to-Image Generation via Bilingual Alignment and Caption Synthesis. | Kang Zhao, Xinyu Zhao, Zhipeng Jin, Yi Yang, Wen Tao, Cong Han, Shuanglong Li, Lin Liu |
| 2023 | ICASSP | Online Binaural Speech Separation Of Moving Speakers With A Wavesplit Network. | Cong Han, Nima Mesgarani |
| 2023 | ICASSP | Phoneme-Level Bert for Enhanced Prosody of Text-To-Speech with Grapheme Predictions. | Yinghao Aaron Li, Cong Han, Xilin Jiang, Nima Mesgarani |
| 2023 | ICCV | Open-Vocabulary Semantic Segmentation with Decoupled One-Pass Network. | Cong Han, Yujie Zhong, Dengjie Li, Kai Han, Lin Ma |
| 2022 | ICASSP | Multi-Channel Speech Denoising for Machine Ears. | Cong Han, Emine Merve Kaya, Kyle Hoefer, Malcolm Slaney, Simon Carlile |
| 2022 | NAACL | Improving Conversational Recommendation Systems' Quality with Context-Aware Item Meta-Information. | Bowen Yang, Cong Han, Yu Li, Lei Zuo, Zhou Yu |
| 2021 | ICASSP | Dual-Path Modeling for Long Recording Speech Separation in Meetings. | Chenda Li, Zhuo Chen, Yi Luo, Cong Han, Tianyan Zhou, Keisuke Kinoshita, Marc Delcroix, Shinji Watanabe, Yanmin Qian |
| 2021 | ICASSP | Rethinking The Separation Layers In Speech Separation Networks. | Yi Luo, Zhuo Chen, Cong Han, Chenda Li, Tianyan Zhou, Nima Mesgarani |
| 2021 | ICASSP | Ultra-Lightweight Speech Separation Via Group Communication. | Yi Luo, Cong Han, Nima Mesgarani |
| 2021 | Interspeech | Continuous Speech Separation Using Speaker Inventory for Long Recording. | Cong Han, Yi Luo, Chenda Li, Tianyan Zhou, Keisuke Kinoshita, Shinji Watanabe, Marc Delcroix, Hakan Erdogan, John R. Hershey, Nima Mesgarani, Zhuo Chen |
| 2021 | Interspeech | Binaural Speech Separation of Moving Speakers With Preserved Spatial Cues. | Cong Han, Yi Luo, Nima Mesgarani |
| 2021 | Interspeech | Empirical Analysis of Generalized Iterative Speech Separation Networks. | Yi Luo, Cong Han, Nima Mesgarani |
| 2020 | ICASSP | Real-Time Binaural Speech Separation with Preserved Spatial Cues. | Cong Han, Yi Luo, Nima Mesgarani |
| 2019 | ASRU | FaSNet: Low-Latency Adaptive Beamforming for Multi-Microphone Audio Processing. | Yi Luo, Cong Han, Nima Mesgarani, Enea Ceolini, Shih-Chii Liu |
| 2019 | ICASSP | Online Deep Attractor Network for Real-time Single-channel Speech Separation. | Cong Han, Yi Luo, Nima Mesgarani |
| 2019 | IWCMC | A Multi-objective Service Function Chain Mapping Mechanism for IoT networks. | Cong Han, Siya Xu, Shaoyong Guo, Xuesong Qiu, Ao Xiong, Peng Yu, Kunya Guo, Dong Guo |
| 2015 | IGARSS | The National Entironmental and Geological Information System for Remote Sensing Survey and Monitoring. | Yunpeng Yan, Zhengmin He, Gang Liu, Yanzuo Wang, Cong Han |