https://www.nature.com/articles/s41586-024-07856-5
LLMs already discriminates against African-American English. You could argue a human grader would as well, but all tested models were more consistent in assigning negative adjectives to hypothetical speakers of that dialect.