A practical guide to running 8x RTX PRO 6000's(gpupartner.com) |
A practical guide to running 8x RTX PRO 6000's(gpupartner.com) |
The M3 Ultra's GPU performance is around a 4070 Ti. The M5 Ultra more like a 5080. They're both amazing deals compared to Nvidia for local inference because of their massive pool of high bandwidth memory. But a single RTX PRO 6000 should be 2 or 3x the compute of an M5 Ultra.
No comparison.
None.
Memory-wise, the RTX PRO 6000 can barely hold two 1M context Qwen 3.8 27B models at 8 bit quantization at the same time. The 512GB M5 Ultra Mac Studio could hold around 14.
At such high concurrency, batch performance is usually limited more by memory bandwidth than compute. The RTX PRO 6000's memory bandwidth is just 50% faster than the M5 Ultra.
So yeah, I think if we are talking about many short context requests, sure. But if you are chewing through a backlog of coding tasks with Qwen overnight, they might actually be comparable.
Im sure in Nov when the M5 Ultra comes out we'll see a lot of interesting benchmarks.
Oh, okay, so this is an ad.
I do still think it's well written and interesting... But if anything, it's just making me more curious about the newest generation of M5 Ultra. (and less and less interested in PCI-E Gen 5 anything)
https://github.com/aikitoria/open-gpu-kernel-modules
The hardware supports it, but Nvidia disabled it if the driver detects cheaper cards.
because I would read it.
The thing is Qwen 3.8 27B can be ran on far far cheaper hardware. If you're spending the big bucks on these rigs you probably made the wrong choice if you aren't using models that require all of that VRAM.
(Edit: I'm mistaken here, the pelican didn't come from Flash on 4 cards but from the full GLM 5.3 model on 8. But the Flash model is still crazy good for its size.)
It actually runs fine at FP8 on this hardware too, with the full 1M context.
No person with real practical experience and real use cases will be using these ancient models as examples, when talking about local LLMs.
stopped reading after that. What 4k context would be usable for?
That's a solid 2200 words to spend on operating parameters and conveying state, leaving a generous 700 word window for them to decide and respond in.
When the bonsai/prism 1bit models dropped and I saw how many prompts a minute I could get from a dusty m2 mini I started hooking it up to all sorts of shit, like a traffic simulator that translates the car state/surroundings/immediate goal into text, it responds with a seqeunce of actions defined in the system prompt, which then get translated back into NPC input.
What I was hoping for here was that it would result in fucking chaos, all sorts of stupid decisions and epic car accidents. I cannot overstate my disappointment (and terror) when they were perfectly reasonable, safe drivers. I had to cut the tire grip by 75% without telling them and make them control twice as many cars to delay their ability to respond before I saw anything resembling an enjoyable traffic accident.
Sure, let me just buy $60,000 worth of GPUs to run a *quantized non-frontier model*.
For that price you could:
- put a down payment on a home in a large % of the US
- buy a brand new car in cash (possibly two!)
- take a long sabbatical and travel the world
- pay all 4 years or your child's college tuition
edit: how do you feel about paying 5% on every token's cost to stripe? it made me cancel, and sign up directly with a few providers.
For the record, I did laugh, I just didn't type it out.
> how do you feel about paying 5% on every token's cost to stripe?
Life is compromise. In a perfect world, I'd love to buy a linux native system comparable to the Mac Studio that could run something in the ballpark of Deepseek V4 flash "fast" at home.
Not only does that not exist, but even if it did I couldn't justify the cost. Between work and home my annual token spend is maybe 3k.
So... paying a 5% fee on a service that allows me to not shell out 10k on hardware at home still seems like a pretty sweet deal.
--
Plus, I still am learning a ton and playing lots. It's hard to imagine anything beating openrouter for that. So much exposure to all the latest LLMs...
The topic of the story is running 8x RTX PRO 6000s, so whining about how much they cost or what you could/should buy instead is completely off-topic.