SWARM Parallelism: Training Large Models Can Be Surprisingly Communication-Efficient.
Max Ryabinin, Tim Dettmers, Michael Diskin, Alexander Borzunov
Browse the full ICML paper archive.
Max Ryabinin, Tim Dettmers, Michael Diskin, Alexander Borzunov
Browse the full ICML paper archive.