Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions.
Federico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul Rttger, Dan Jurafsky, Tatsunori Hashimoto, James Zou
Browse the full ICLR paper archive.