Dynamic Rewarding with Prompt Optimization Enables Tuning-free Self-Alignment of Language Models.
Somanshu Singla, Zhen Wang, Tianyang Liu, Abdullah Ashfaq, Zhiting Hu, Eric P. Xing
Browse the full EMNLP paper archive.
Somanshu Singla, Zhen Wang, Tianyang Liu, Abdullah Ashfaq, Zhiting Hu, Eric P. Xing
Browse the full EMNLP paper archive.