Skip to content

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models.

Lillian Sun, Martin Pawelczyk, Zhenting Qi, Aounon Kumar, Himabindu Lakkaraju

VenueA*ACL
Year2026
ProceedingsACL (1)

Browse the full ACL paper archive.