The purchase cost of H100 or B200 systems with comparable VRAM is a one order of magnitude higher. Although I can only guess how much lower the token/sec output of the Mac Studio will be. Probably 2-3 magnitudes lower?
While a cluster has to work with many users simultanously, and is a good investment for a company, perhaps the Mac Studio will be a good use case for a personal larger LLM deployment configuration.
Perhaps someone has the token/sec numbers for larger models running on older Mac Studios?