Nvidia's Tegra X1 should supposedly be capable of <10ms for imagenet grade models [2]. It's fair to assume though, that this must be for trimmed down and/or 16bit models as compared to full inception models.
And finally, Sam who also facilitated building TF on the Pi is about to host a 6 weeks half theory, half practice course on TF and deep learning [3] (me thinks he deserves this plug).
[1] https://github.com/samjabrahams/tensorflow-on-raspberry-pi/t...
[2] https://youtu.be/_4tzlXPQWb8?t=53m35s
[3] https://www.thisismetis.com/deep-learning-with-tensorflow