Failures to Find Transferable Image Jailbreaks Between Vision-Language Models.
Rylan Schaeffer, Dan Valentine, Luke Bailey, James Chua, Cristbal Eyzaguirre, Zane Durante, Joe Benton, Brando Miranda, Henry Sleight, Tony Tong Wang, John Hughes, Rajashree Agrawal, Mrinank Sharma, Scott Emmons, Sanmi Koyejo, Ethan Perez
Browse the full ICLR paper archive.