Do as I do (Safely): Mitigating Task-Specific Fine-tuning Risks in Large Language Models.
Francisco Eiras, Aleksandar Petrov, Philip Torr, M. Pawan Kumar, Adel Bibi
Browse the full ICLR paper archive.
Francisco Eiras, Aleksandar Petrov, Philip Torr, M. Pawan Kumar, Adel Bibi
Browse the full ICLR paper archive.