Skip to content

Transformers Provably Learn Two-Mixture of Linear Classification via Gradient Flow.

Hongru Yang, Zhangyang Wang, Jason D. Lee, Yingbin Liang

VenueA*ICLR
Year2025
ProceedingsICLR

Browse the full ICLR paper archive.