| 2026 | EACL | Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation. | Jiwon Moon, Yerin Hwang, Dongryeol Lee, Taegwan Kang, Yongil Kim, Kyomin Jung |
| 2025 | ACL | LLMs can be easily Confused by Instructional Distractions. | Yerin Hwang, Yongil Kim, Jahyun Koo, Taegwan Kang, Hyunkyung Bae, Kyomin Jung |
| 2025 | EMNLP | Can You Trick the Grader? Adversarial Persuasion of LLM Judges. | Yerin Hwang, Dongryeol Lee, Taegwan Kang, Yongil Kim, Kyomin Jung |
| 2025 | EMNLP | Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation. | Yerin Hwang, Dongryeol Lee, Kyungmin Min, Taegwan Kang, Yongil Kim, Kyomin Jung |
| 2025 | NAACL | SWITCH: Studying with Teacher for Knowledge Distillation of Large Language Models. | Jahyun Koo, Yerin Hwang, Yongil Kim, Taegwan Kang, Hyunkyung Bae, Kyomin Jung |
| 2025 | NAACL | Are LLM-Judges Robust to Expressions of Uncertainty? Investigating the effect of Epistemic Markers on LLM-based Evaluation. | Dongryeol Lee, Yerin Hwang, Yongil Kim, Joonsuk Park, Kyomin Jung |
| 2024 | COLING | Kosmic: Korean Text Similarity Metric Reflecting Honorific Distinctions. | Yerin Hwang, Yongil Kim, Hyunkyung Bae, Jeesoo Bang, Hwanhee Lee, Kyomin Jung |
| 2024 | EMNLP | MP2D: An Automated Topic Shift Dialogue Generation Framework Leveraging Knowledge Graphs. | Yerin Hwang, Yongil Kim, Yunah Jang, Jeesoo Bang, Hyunkyung Bae, Kyomin Jung |
| 2023 | ACL | Injecting Comparison Skills in Task-Oriented Dialogue Systems for Database Search Results Disambiguation. | Yongil Kim, Yerin Hwang, Joongbo Shin, Hyunkyung Bae, Kyomin Jung |
| 2023 | EMNLP | Dialogizer: Context-aware Conversational-QA Dataset Generation from Textual Sources. | Yerin Hwang, Yongil Kim, Hyunkyung Bae, Hwanhee Lee, Jeesoo Bang, Kyomin Jung |
| 2023 | EMNLP | PR-MCS: Perturbation Robust Metric for MultiLingual Image Captioning. | Yongil Kim, Yerin Hwang, Hyeongu Yun, Seunghyun Yoon, Trung Bui, Kyomin Jung |