For example, suppose we implement short range order detection, as you mention. Then, a very well written essay about the history of various metaphors used to describe General Relativity would receive a low score, because we'd be talking about trampolines, marbles, spider webs, and maybe even balloons at the same time. Your algorithm would immediately classify this as nonsense (similar to what I have written above).
The problem is not in detecting structural integrity. The problem is in detecting information integrity.
A nonsense essay about metaphors cannot be distinguished from a reasonable essay about metaphors unless your network knows what a metaphor is, how isomorphism works, and knows how various ideas relate to that.
Long story short, you can't hire an english teacher to grade a paper you're submitting for your GR course. Such a paper would be immune from random insertions of the word "not" or changing random real-valued numbers to other numbers, etc. Any NLP-only algorithm would immediately identify a false positive here.
On the other hand, a grammar-only-grader would be reasonable, because grammar is mostly independent of topic, meaning the AI would not need domain knowledge to grade it.
To understand language, AI is going to have be a lot more versatile than simply finding order within a paragraph, which is the most naive approach.
When a human reads a paragraph, they are playing with much more than simply the words contained inside the paragraph. IIRC, the ML community isn't there yet -- not because they don't know -- but because computing isn't there yet.