You've got it backwards. Modern speech recognizers have a vocabulary of a million words and multi-gigabyte models. It's generally much more accurate to do speech recognition in the cloud, since you have more processing power and more RAM to hold large statistical models.
The rumor is that Apple is sending the audio to Nuance servers, i.e., they're doing cloud-based speech recognition.