Confident, Calibrated, or Complicit: Safety Alignment and Ideological Bias in LLM Hate Speech Detection.
Sanjeevan Selvaganapathy, Mehwish Nasim
Browse the full ACL paper archive.
Sanjeevan Selvaganapathy, Mehwish Nasim
Browse the full ACL paper archive.