I find this inclusion at the end very interesting:
> In our paper, we detail how we built a highly effective classifier that can distinguish between authentic speech and audio generated with Voicebox to mitigate these possible future risks.
I wonder how that classifier works. Is it also possible the original Voicebox output includes some kind of watermark, perhaps?