The black box can adapt to new circumstances. A lot of the algorithms used today in speech recognition were established on much smaller data sets (thousands of hours of speech); the tradeoffs made then may not apply when you have 1000x the data. The more automatic the algorithm, the more it can change.
Existing speech recognizers aren't really a "white box". You can't look at a Gaussian Mixture Model and understand what it's doing.
The more you automate the whole training process, the more.