One Step of Gradient Descent is Provably the Optimal In-Context Learner with One Layer of Linear Self-Attention.
Arvind V. Mahankali, Tatsunori Hashimoto, Tengyu Ma
Browse the full ICLR paper archive.
Arvind V. Mahankali, Tatsunori Hashimoto, Tengyu Ma
Browse the full ICLR paper archive.