Batch Policy Gradient Methods for Improving Neural Conversation Models.
Kirthevasan Kandasamy, Yoram Bachrach, Ryota Tomioka, Daniel Tarlow, David Carter
Browse the full ICLR paper archive.
Kirthevasan Kandasamy, Yoram Bachrach, Ryota Tomioka, Daniel Tarlow, David Carter
Browse the full ICLR paper archive.