Skip to content

Detecting and Understanding Vulnerabilities in Language Models via Mechanistic Interpretability.

Jorge Garca-Carrasco, Alejandro Mat, Juan Trujillo

VenueA*IJCAI
Year2024
ProceedingsIJCAI

Browse the full IJCAI paper archive.