Skip to content

SWARM Parallelism: Training Large Models Can Be Surprisingly Communication-Efficient.

Max Ryabinin, Tim Dettmers, Michael Diskin, Alexander Borzunov

VenueA*ICML
Year2023
ProceedingsICML

Browse the full ICML paper archive.