Dialog policy optimization for low resource setting using Self-play and Reward based Sampling.
Tharindu Madusanka, Durashi Langappuli, Thisara Welmilla, Uthayasanker Thayasivam, Sanath Jayasena
Browse the full PACLIC paper archive.
Tharindu Madusanka, Durashi Langappuli, Thisara Welmilla, Uthayasanker Thayasivam, Sanath Jayasena
Browse the full PACLIC paper archive.