1-bit stochastic gradient descent and its application to data-parallel distributed training of speech DNNs.
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, Dong Yu
Browse the full Interspeech paper archive.
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, Dong Yu
Browse the full Interspeech paper archive.