I have zero knowledge of AI training, but is it possible that, soon, the line between training per se and fine-tuning will get technically so narrow that some groups can feed tens (or hundreds ?) of GBs of Copyrighted datasets [1] into the best open source AI models (LLMs or otherwise) ? I guess you could call this piracy 3.0.