LLM Safety From Within: Detecting Harmful Content with Internal Representations.
Difan Jiao, Yilun Liu, Ye Yuan, Zhenwei Tang, Linfeng Du, Haolun Wu, Ashton Anderson
Browse the full ACL paper archive.
Difan Jiao, Yilun Liu, Ye Yuan, Zhenwei Tang, Linfeng Du, Haolun Wu, Ashton Anderson
Browse the full ACL paper archive.