Skip to content

Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context.

Xiang Cheng, Yuxin Chen, Suvrit Sra

VenueA*ICML
Year2024
ProceedingsICML

Browse the full ICML paper archive.