Does that mean that a single user has an exclusive lock on 8 of these things for the duration of the request? I'm surely missing something.
You're not missing anything, that's exactly what it means. There is a reason OpenAI severely ratelimits the number of GPT-4 prompts users can make. Every token (approximately 4 letters) requires many trillions of floating point operations.
Large language models are only possible because of decades of exponential transistor density growth.
This is why developers such as myself have opted for Apple's new unified architecture to run smaller models locally, as the GPU can access all 64GB. But these small models are still a fraction of the size of GPT-3.5-Turbo and GPT-4.
Microsoft should have the maturity, money, and any other resources needed to solve their own problems without making a public spectacle of it. It's embarrassing and childish.
I am being critical of them, but their net worth is nearly 1/10th of the debt of the entire United States of America. If they can't get it together with that level of resources, then they should be ashamed of themselves.