Per the technical report:
> The minimum deployment unit of the decoding stage consists of 40 nodes with 320 GPUs.
but realistically, >=671GB of VRAM to run at full precision on GPU, or >=131G VRAM to run the most heavily quantized version[1], or >=671GB RAM, and a dose of patience to run on CPU.