Skip to content

Batch Policy Gradient Methods for Improving Neural Conversation Models.

Kirthevasan Kandasamy, Yoram Bachrach, Ryota Tomioka, Daniel Tarlow, David Carter

VenueA*ICLR
Year2017
ProceedingsICLR (Poster)

Browse the full ICLR paper archive.