I'll prove it by induction: Imagine that I have a service where I "train" a model on a single image of Indiana Jones. Now you prompt it, and my model "generates" the same image. I sell you this service, and no money goes to the copyright holder of the original image. This is obviously infringment.
There's no reason why training on a billion images is any different, besides the fact that the lines are blurred by the model weights not being parseable