Skip to content

LLMs Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions.

Xuhao Hu, Peng Wang, Xiaoya Lu, Dongrui Liu, Xuanjing Huang, Jing Shao

VenueA*ACL
Year2026
ProceedingsACL (Findings)

Browse the full ACL paper archive.