Skip to content

Improving Deep Transformer with Depth-Scaled Initialization and Merged Attention.

Biao Zhang, Ivan Titov, Rico Sennrich

VenueA*EMNLP
Year2019
ProceedingsEMNLP/IJCNLP (1)

Browse the full EMNLP paper archive.