| 2026 | AAAI | End-to-End Contrastive Language-Speech Pretraining Model for Long-Form Spoken Question Answering. | Jiliang Hu, Zuchao Li, Baoyuan Qi, Guoming Liu, Ping Wang |
| 2026 | AAAI | Scaling LLM Speculative Decoding: Non-Autoregressive Forecasting in Large-Batch Scenarios. | Luohe Shi, Zuchao Li, Lefei Zhang, Baoyuan Qi, Guoming Liu, Hai Zhao |
| 2025 | ACL | KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding. | Luohe Shi, Zuchao Li, Lefei Zhang, Baoyuan Qi, Guoming Liu, Hai Zhao |
| 2025 | ACL | SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers. | Zicong Tang, Luohe Shi, Zuchao Li, Baoyuan Qi, Guoming Liu, Lefei Zhang, Ping Wang |
| 2025 | ACL | DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression. | Yi Zhao, Zuchao Li, Hai Zhao, Baoyuan Qi, Guoming Liu |
| 2025 | EMNLP | Faster In-Context Learning for LLMs via N-Gram Trie Speculative Decoding. | Jinglin Chen, Qiwei Li, Zuchao Li, Baoyuan Qi, Guoming Liu, Haojun Ai, Hai Zhao, Ping Wang |
| 2025 | EMNLP | XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression. | Haoqi Yang, Yao Yao, Zuchao Li, Baoyuan Qi, Guoming Liu, Hai Zhao |
| 2025 | HCI | Designing for Utilitarian or Hedonic Motivation: The Impact of Appearance and Social Cues of Chatbots. | Xuanyue Feng, Guoming Liu, Man Wu |
| 2025 | ICML | What Limits Bidirectional Model's Generative Capabilities? A Uni-Bi-Directional Mixture-of-Expert Method For Bidirectional Fine-tuning. | Zuchao Li, Yonghua Hei, Qiwei Li, Lefei Zhang, Ping Wang, Hai Zhao, Baoyuan Qi, Guoming Liu |