I had intended to caveat that: I'm sure I'm not the first person to ask about this!
> you still see improvements
This is expected if they are training their models on it, right?
> objectively-bad results
Keen to learn when this has been the case, i.e. across version increments in major models.