> EMA-based training dynamics like JEPA’s don’t optimize any smooth mathematical function, yet they provably converge to useful, non-collapsed representations.
All the papers say EMA avoids “representation collapse” without justifying it. Didn’t realize there were any theoretical results here.