Makes sense. This also tracks with the research on human-AI collaboration. A single model converges to the mean of its training distribution, but adversarial multi-model setups break that pattern because each model's blind spots are different.
I wrote about why single-model AI has a structural quality ceiling and why ensemble/hybrid approaches consistently outperform: https://philippdubach.com/posts/the-impossible-backhand/