Silver made the same point about probability before the 2016 election, and got into a similar Twitter argument (https://www.politico.com/story/2016/11/nate-silver-huffingto...). One guess on who proved to be correct.
But in 538's case they use the similar methods to forecast many individual contests like presidential primaries, presidential elections, senate elections, house elections, governor elections, club soccer, college football, March Madness men and women, MLB-NBA-NFL games and playoffs, etc.
Multiple-year track records over many events so you can compare their forecasts to actual results over time.
Get the data and compute your own error rate here: https://github.com/fivethirtyeight/checking-our-work-data
This isn't actually true. The probability of winning is based on aggregation of regional probabilities. The same national vote probability distribution can lead to very, very, different election probability distributions due to regional variation. The electoral college essentially guarantees that the national probability distribution is worthless for actually predicting who will be President.
If candidate B then wins, it does not mean that our analysis was "proved to be [in]correct". By itself, it doesn't actually say anything about the quality of our analysis. After all, we explicitly pointed out that this was a possibility, and it would be strange to argue "your analysis said this might happen, and then it did, so your analysis was incorrect". There's just not enough information to draw any conclusions.
How about 100 times? 10 times? 2 times?
At what point does evidence cease to 'say anything about the quality of our analysis'? The answer is never. Every datapoint can be used to update your priors according to bayesian statistics.
Especially for someone who titles them-self "Chief Scientist"