I'm not saying those assumptions are necessarily wrong, either, just that this is a slightly arrogant simplification of a field that has already hashed these kinds of questions through and through, and decided that such simplistic views don't yet have enough experimental evidence to be cut and dry truths.
Aside from that – the paper's main contribution is essentially "I have a pet theory that lets me predict everything we've already seen LLMs do, but nothing more." This is accompanied with a simulation, which shows nothing more than that Transformers can learn an unambiguous, 18-letter/6 sentence toy language generated with a Markov chain better than an ambiguous one. This simulation does not even come close to supporting the claims and assumptions in the rest of the paper.