Is there a reason wrong data isn't considered more broadly in its context as still valuable?
Shouldn't the model effectively 1. learn to complete the incorrect thing and 2. learn the context that it's correct and incorrect? In this case the context being lazy LMArena users. And presumably, in the future, poorly filtered training data.
We seem to be able to read incorrect things and not be corrupted (well, theoretically). It's not ideal, but it seems an important component to intellectual resilience.
It seems like the model knowing the data is LMArena, or some type of un-trusted, would be sufficient to shift the prior to a reasonable place.