Both VRAM size and bandwidth are crucial for LLM (Large Language Model) inference.
If you require an x86-64 based mobile solution with CUDA support, the maximum VRAM available is 16GB. The Strix HALO is positioned as a competitor to the RTX 4070M.
"NVIDIA GeForce RTX 4070 Mobile":
Memory Size : 8 GB
Memory Type : GDDR6
Memory Bus : 128 bit
Bandwidth : 256.0 GB/s
"NVIDIA GeForce RTX 4090 Mobile" Memory Size : 16 GB
Memory Type : GDDR6
Memory Bus : 256 bit
Bandwidth : 576.0 GB/s