The accents are sometimes tagged in the metadata csv files. So it’s possible to filter some of them out.
Mozilla DeepSpeech used to release checkpoints trained on Common Voice along with a few others (LibriSpeech etc). But they’ve dropped CV in the latest releases and just rely on the others.
I think the others are more standardised in terms of accents.
It’s likely possible to fine tune a model with different accents — so long as the language is the same then the model can up date the phonemes it recognises.
But accents are definitely a live issue.