I think everyone remembers the first time they went from "I think I can tolerate some server fans. How loud can they be?" to "I had no idea a small 12V fan could be this loud"
That's to slot into the front PCI brackets for full length cards. There's "full" size for PCI cards and most GPUs are considered "half" length cards. This is separate to Low Profile, so there are four possible slot compatibility configurations, namely HHHL, HHFL, FHFL, and FHHL. Real workstations(including Mac Pro) has slots cut in the case or a bracket to accept the card or that metal piece(which means a standardized solution to GPU sagging had existed even before PCIe was created).
Also, on fans, it's not like blower fans are quieter at all, but in case anyone encounters a situation where turning up fans seem to be only creating noises without moving air or cooling yhe cards, it might be worth remembering that radial fans have higher static pressure and are better at forcing air through.
edit: ps: people recreating this might want to know what case this is. The vast majority of ATX cases, even unnecessarily big ones, only has 7 PCI slots. Cases that can take four dual-slot cards is rare.
>and from someone that has been using Claude for a long time, I can definitely say I don't need it anymore. Not for the stuff I'm doing.
I'm curious what this implies. It suggests that you were using Claude prior to this setup, so it doesn't seem like you were limited by security or local features? I can understand that someone not wanting to send their data to Anthropic (etc.) might be willing to pay for this kind of setup to accomplish that, but what was your motivation for doing this?
If it matters, if I had enough money that I could blow $10k on 2 DGX Sparks without it being a significant cost I probably would for the hell of it, so I'm not digging on you if this is ultimately what this comes down to. But, as an investment, this doesn't seem like a very good deal.
I see this rather as an investment in improving my capabilities and knowledge of this technology, letting me play with a 'GPU cluster', vLLM and other technologies that otherwise would require me to rent GPUs on the cloud.
Offhand, these include "preferring text and my own machines for storing information that is useful to me" over "clouds" and e.g. Word docs. Also, having my own domain and email (which I pay for).
I'm thinking of, e.g. the guy that blogged for years and years and then Google just yoinked it and it was gone. Younger me was more probably more obnoxious about it, like, serves you right -- a thing I would NEVER say now -- but, still, people like me really were right all along about this sort of thing, and I daresay it would be good if everyone followed us.
Compare that to a $5k DGX than consumes much less and has the same amount of VRAM and there is a real question as whether this is worth doing at all (well aside from the cool factor).
So higher-end Claude models (I've got a subscription btw) do work fine right?
What makes you think that in six months he cannot swap DS4 Flash 0731 for another model that could be equivalent to todays' top OpenAI/Anthropic models?
Or are you going to explain in six months that, after all, the top models from Anthropic from today are unfit for use?
ROCm runs most things just fine without any ceremony or difficulty. I have a couple of the same cards as covered in the article, bought before they got expensive (I'd recommend current Radeon AI Pro 9700 cards over the V620 now, though, as they have increased in price since I bought mine), and they're at the long end of the supported chart for ROCm...nearly EOL. But, they currently work great with llama.cpp and current ROCm 7.14. They're pretty fast and stable. I also have a Strix Halo 128GB, and while the Strix Halo can run bigger models, the dedicated GPUs are quite a bit faster. And, since the best models you can run at home are probably Qwen 3.6 27B and Gemma 4 31B, and those both fit comfortably on dual 32GB GPUs, I find I use the desktop more often than the Strix Halo.
Anyway, there's very little reason to spend 3x or more for the Nvidia ecosystem these days. The only exception is the Strix Halo vs the Nvidia GB10 platform. Strix Halo was a no-brainer when it was half the price, but it's risen in price to be almost the same price as an Asus GX10. In that one instance, I think the Nvidia based Asus is a better choice. The Strix Halo is almost too slow to make use of 128GB. Some MoE models are comfortable, like Laguna S2.1, so it's probable that there will someday be an MoE model that is better then Qwen 3.6 27B or Gemma 4 31B that won't run on 64GB but will run on 128GB. I don't know of one, yet, though.
AMD also used to focus almost entirely on their data center line and not pay any attention to their lower end stuff for more advanced AI features, but that's been changing, and ROCm support has broadened to include almost everything AMD ships now, including the smaller embedded stuff. I think they've realized that as long as the only way to develop AI for AMD was to have a quarter million dollars worth of hardware, they would always fall behind a platform that can be developed for and tested on consumer hardware.
That's about 4x the price of v620, which is 450 eurobucks. Can't imagine buying 4 of those at 4x the price.
I should say, though, that I actually don't recommend buying anything right now. I wrote up my setup and made recommendations (and the main recommendation was "don't"). https://swelljoe.com/post/how-i-run-local-llms/
The V620 at $450 or the Radeon AI Pro 9700 at ~$1400 are a good deal compared to everything else right now, but buying tokens from DeepSeek is a better deal. DeepSeek V4 Flash 0731 is better than anything you can host locally, they'll serve it to you at blistering fast speeds for pennies a day, and with 1 million token context. You could host a 2-bit quantization of it on a Strix Halo or four of these V620s, and it would run at a crawl on either one. The Strix Halo gets 9-13 t/s. Four V620s would, I guess, get two-three times that. Which is still too slow for comfortable interactive agentic use, and much slower than getting it from DeepSeek, and at two bits there is measurable loss. You're paying a lot more for self-hosted and you're getting worse models.
Also, if you want four cards, you need a server-class motherboard and CPU and RAM. More money.
It's all just a bad investment. Self-hosting is a bad idea if you don't already have the hardware, unless and until memory and GPU prices come down. Even at the prices I paid (before RAMpocalypse really kicked off; $2k for the Strix Halo, ~$350 for each V620), I wouldn't recommend it if you don't have a strong urge to tinker with hardware and it'll probably never pay for itself vs. buying inference from DeepSeek directly.
Edit: The same shrouds for all the old Instinct cards also work for the V620, as they have the same dimensions and screw and cable layout. I tried several and ended up with this one: https://www.thingiverse.com/thing:7296707
It was the most trouble-free AI setup that I have done. The driver worked right out of the box with the stock Fedora kernel, ROCM can be installed from the regular repo, and most AI software has ROCM builds by now.
Ahh, one of those projects
Yes, I was the person drawing the floor plans and forgot the electric socket there too.
You can not get around this, conceptually. Local inference will keep depend on updated models for quite a while. Partly because they will contain outdated training data, partly because of demands for the improved models. And it's still not clear where this will lead us. We're still in the rosey phase where people get lured in.
Yes, bringing everything local still means you need to download things, maybe over and over. But once it is on your machine, nobody can yoink it from you just because they want more money or they don't like what you're doing with it.
We already see a lot of that with `web_search` tooling, so I imagine it would just become more essential to have tools like that.
I doubt it: PLA begins to soften at as low as 55 degrees Celsius. If there's weight on it, it's worse: the piece shall quickly deform and become unfit for its purpose.
This rig looks like it means business: I think PLA would fail.
If, like me, you cannot print ASA (say because you've got a printer that is not closed and that won't heat enough), then PETG-CF (PETG reinforced with some carbon fiber: it's got better heat deflection than plain PETG) is a safer bet than PLA for parts to put in PCs/servers/rack. Moreover PETG-CF do look really good.