| 2026 | AAAI | AutoSCORE: Enhancing Automated Scoring with Multi-Agent Large Language Models via Structured Component Recognition. | Yun Wang, Zhaojun Ding, Xuansheng Wu, Siyue Sun, Ninghao Liu, Xiaoming Zhai |
| 2026 | AIED | BRIDGE the Gap: Mitigating Bias Amplification in Automated Scoring of English Language Learners via Inter-group Data Augmentation. | Yun Wang, Xuansheng Wu, Jingyuan Huang, Lei Liu, Xiaoming Zhai, Ninghao Liu |
| 2026 | EACL | Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering. | Haiyan Zhao, Xuansheng Wu, Fan Yang, Bo Shen, Ninghao Liu, Mengnan Du |
| 2025 | AIED | Artificial Intelligence Bias on English Language Learners in Automatic Scoring. | Shuchen Guo, Yun Wang, Jichao Yu, Xuansheng Wu, Bilgehan Ayik, Field M. Watts, Ehsan Latif, Ninghao Liu, Lei Liu, Xiaoming Zhai |
| 2025 | EMNLP | Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders. | Dong Shu, Xuansheng Wu, Haiyan Zhao, Mengnan Du, Ninghao Liu |
| 2025 | EMNLP | A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models. | Dong Shu, Xuansheng Wu, Haiyan Zhao, Daking Rai, Ziyu Yao, Ninghao Liu, Mengnan Du |
| 2025 | ICML | Concept-Centric Token Interpretation for Vector-Quantized Generative Models. | Tianze Yang, Yucheng Shi, Mengnan Du, Xuansheng Wu, Qiaoyu Tan, Jin Sun, Ninghao Liu |
| 2025 | KDD | Self-Regularization with Sparse Autoencoders for Controllable LLM-based Classification. | Xuansheng Wu, Wenhao Yu, Xiaoming Zhai, Ninghao Liu |
| 2025 | NAACL | LMOD: A Large Multimodal Ophthalmology Dataset and Benchmark for Large Vision-Language Models. | Zhenyue Qin, Yu Yin, Dylan Campbell, Xuansheng Wu, Ke Zou, Ninghao Liu, Yih Chung Tham, Xiuzhen Zhang, Qingyu Chen |
| 2024 | ACL | InFoBench: Evaluating Instruction Following Ability in Large Language Models. | Yiwei Qin, Kaiqiang Song, Yebowen Hu, Wenlin Yao, Sangwoo Cho, Xiaoyang Wang, Xuansheng Wu, Fei Liu, Pengfei Liu, Dong Yu |
| 2024 | CIKM | Retrieval-enhanced Knowledge Editing in Language Models for Multi-Hop Question Answering. | Yucheng Shi, Qiaoyu Tan, Xuansheng Wu, Shaochen Zhong, Kaixiong Zhou, Ninghao Liu |
| 2024 | NAACL | From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning. | Xuansheng Wu, Wenlin Yao, Jianshu Chen, Xiaoman Pan, Xiaoyang Wang, Ninghao Liu, Dong Yu |
| 2024 | WWW | Could Small Language Models Serve as Recommenders? Towards Data-centric Cold-start Recommendation. | Xuansheng Wu, Huachi Zhou, Yucheng Shi, Wenlin Yao, Xiao Huang, Ninghao Liu |
| 2023 | AIED | Matching Exemplar as Next Sentence Prediction (MeNSP): Zero-Shot Prompt Learning for Automatic Scoring in Science Education. | Xuansheng Wu, Xinyu He, Tianming Liu, Ninghao Liu, Xiaoming Zhai |