This from DeepMind:
DiLoCo: Distributed Low-Communication Training of Language Models - https://arxiv.org/pdf/2311.08105.pdf
From the first author on Twitter: "It could quite a big deal for people who don't have access to a colocated cluster of GPUs:
e.g. with DiLoCo you could train your model, with data-parallelism, across all GPU providers, looking in real-time for the cheapest price, even if pre-emptable, even across continents"