Skip to content

LLM Safety From Within: Detecting Harmful Content with Internal Representations.

Difan Jiao, Yilun Liu, Ye Yuan, Zhenwei Tang, Linfeng Du, Haolun Wu, Ashton Anderson

VenueA*ACL
Year2026
ProceedingsACL (1)

Browse the full ACL paper archive.