On-line Active Reward Learning for Policy Optimisation in Spoken Dialogue Systems.
Pei-Hao Su, Milica Gasic, Nikola Mrksic, Lina Maria Rojas-Barahona, Stefan Ultes, David Vandyke, Tsung-Hsien Wen, Steve J. Young
Browse the full ACL paper archive.
Pei-Hao Su, Milica Gasic, Nikola Mrksic, Lina Maria Rojas-Barahona, Stefan Ultes, David Vandyke, Tsung-Hsien Wen, Steve J. Young
Browse the full ACL paper archive.