Skip to content

Zeroth-Order Optimization Meets Human Feedback: Provable Learning via Ranking Oracles.

Zhiwei Tang, Dmitry Rybin, Tsung-Hui Chang

VenueA*ICLR
Year2024
ProceedingsICLR

Browse the full ICLR paper archive.