Three. You can also do it in Metal, which as of recently has cooperative matrix multiplication in the form of the simd_matrix type (this is similar functionality as "tensor cores" in the Nvidia world). I have no idea what the software support is, but I have seen analysis suggesting that the raw tensor multiplication throughput is larger than ANE for the high-end GPUs.