Mantas Mazeika
Publication record assembled from the DBLP archive of ranked conferences.
Papers indexed
9
Venues
3
Active years
2019–2025
Best venue rank
A*
Where they publish
Papers
9 indexed papers, newest first.
| Year | Venue | Title | Authors |
|---|---|---|---|
| 2025 | ICLR | Tamper-Resistant Safeguards for Open-Weight LLMs. | Rishub Tamirisa, Bhrugu Bharathi, Long Phan, Andy Zhou, Alice Gatti, Tarun Suresh, Maxwell Lin, Justin Wang, Rowan Wang, Ron Arel, Andy Zou, Dawn Song, Bo Li, Dan Hendrycks, Mantas Mazeika |
| 2024 | ICML | The WMDP Benchmark: Measuring and Reducing Malicious Use with Unlearning. | Nathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue, Daniel Berrios, Alice Gatti, Justin D. Li, Ann-Kathrin Dombrowski, Shashwat Goel, Gabriel Mukobi, Nathan Helm-Burger, Rassin Lababidi, Lennart Justen, Andrew B. Liu, Michael Chen, Isabelle Barrass, Oliver Zhang, Xiaoyuan Zhu, Rishub Tamirisa, Bhrugu Bharathi, Ariel Herbert-Voss, Cort B. Breuer, Andy Zou, Mantas Mazeika, Zifan Wang, Palash Oswal, Weiran Lin, Adam A. Hunt, Justin Tienken-Harder, Kevin Y. Shih, Kemper Talley, John Guan, Ian Steneker, David Campbell, Brad Jokubaitis, Steven Basart, Stephen Fitz, Ponnurangam Kumaraguru, Kallol Krishna Karmakar, Uday Kiran Tupakula, Vijay Varadharajan, Yan Shoshitaishvili, Jimmy Ba, Kevin M. Esvelt, Alexandr Wang, Dan Hendrycks |
| 2024 | ICML | HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal. | Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David A. Forsyth, Dan Hendrycks |
| 2022 | CVPR | PixMix: Dreamlike Pictures Comprehensively Improve Safety Measures. | Dan Hendrycks, Andy Zou, Mantas Mazeika, Leonard Tang, Bo Li, Dawn Song, Jacob Steinhardt |
| 2022 | ICML | Scaling Out-of-Distribution Detection for Real-World Settings. | Dan Hendrycks, Steven Basart, Mantas Mazeika, Andy Zou, Joseph Kwon, Mohammadreza Mostajabi, Jacob Steinhardt, Dawn Song |
| 2022 | ICML | How to Steer Your Adversary: Targeted and Efficient Model Stealing Defenses with Gradient Redirection. | Mantas Mazeika, Bo Li, David A. Forsyth |
| 2021 | ICLR | Measuring Massive Multitask Language Understanding. | Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, Jacob Steinhardt |
| 2019 | ICLR | Deep Anomaly Detection with Outlier Exposure. | Dan Hendrycks, Mantas Mazeika, Thomas G. Dietterich |
| 2019 | ICML | Using Pre-Training Can Improve Model Robustness and Uncertainty. | Dan Hendrycks, Kimin Lee, Mantas Mazeika |