On Calibration of LLM-based Guard Models for Reliable Content Moderation.
Hongfu Liu, Hengguan Huang, Xiangming Gu, Hao Wang, Ye Wang
Browse the full ICLR paper archive.
Hongfu Liu, Hengguan Huang, Xiangming Gu, Hao Wang, Ye Wang
Browse the full ICLR paper archive.