"" Very encouraging to see the steady increase of viable hardware options that can handle various AI models.
At the beginning of the year, there was only one practical option - nVidia. Now we see at-least 3 vendors providing reasonable options. Apple, AMD and Intel. We have been profiling several options and I will share some of our findings here.
The good stuff
- Apple Macs were a pleasant surprise on how easy it is to get various models running
- AMD also made impressive progress with PyTorch and a lot more models run now than even 4-5 months ago on MI2XX and Radeon
- We tried both Intel Arc and Ponte Vecchio and they were able to execute everything we have thrown at them.
- Intel Gaudi has very impressive performance on the models that work on that architecture. It's our current best option for LLM inference on select models.
- Ponte Vecchio surprised us with its performance on our custom face swap model, beating everyone including the mighty H100. We suspect that our model may be fitting largely in the Rambo cache.
The wishlist
- For training and inference of large models that don't fit in memory - nVidia is still the only practical option. Wishing that there are more options in 2024 here
- While compatibility is getting better, a ton of performance is still left on the table on Apple, AMD and Intel. Wishing that software will keep getting better and increase their HW utilization. There is still room on compatibility as well, particularly with supporting various encodings and model parameter sizes on AMD.
- Intel Gaudi looks very promising performance-wise and wishing that more models seamlessly work out of the box without Intel intervention.
- Wishing that both AMD and Intel release new gaming GPUs with more memory capacity and bandwidth.
- Wishing that Intel releases a PVC kicker with more memory capacity and bandwidth. Currently it's the best option we have to bring our artists workflow with face swap training from 3-days to a few hours. It scales linearly from 1-GPU to 16-GPUs.
- Wishing Intel support for PyTorch is as frictionless as AMD and nVdia. May be Intel should consider supporting PyTorch RocM or up-stream OneAPI support under CUDA device.
really grateful to all vendors for providing access to hardware and developer support.
Looking forward to continue filling our data center with interesting mix of architectures. ""