Latent Positional Information is in the Self-Attention Variance of Transformer Language Models Without Positional Embeddings.
Ta-Chung Chi, Ting-Han Fan, Li-Wei Chen, Alexander Rudnicky, Peter J. Ramadge
Browse the full ACL paper archive.
Ta-Chung Chi, Ting-Han Fan, Li-Wei Chen, Alexander Rudnicky, Peter J. Ramadge
Browse the full ACL paper archive.