That's for good NVidia H100 units.[1] There's a shortage of those. That seems to be the price after removal, cleaning, testing and refurbishing. Raw units removed from a shutdown will not be as valuable.
H100 units are available on eBay, but multiple sellers are using the same picture of a new unit in its original packaging, a bad sign.[2] Some even have pictures with the logos of a competitor.
[1] https://introl.com/blog/secondary-gpu-markets-buying-selling...
[2] https://www.ebay.com/shop/nvidia-h100-gpu?_nkw=nvidia+h100+g...
If you search ebay for server equipment like a Dell R840 with 768GB RAM, the same sort of dealers who are selling that and have thousands of feedback (at 98.5% of greater rating) are the ones I would consider much less risk.
From an annualized number on the llama 3 training report. would be interesting to see if we have a better idea given that we're already on rubin.
When I tested "ECO" mode on my AMD CPU, performance was ~97% of regular mode, and temperatures dropped 5-10C. I think ECO mode drops TDP from something like 100W->65W. Cheaper and cooler to run for basically no cost.
If you are bitcoin mining or selling GPU capacity, you are incredibly conscious of your energy bill, so it only makes sense to optimize the power draw.
It was clean, cheap, and is still going strong for daily gaming.
In its working life it was undervolted and probably cooled better than in my rig.
It's wild to think that the system now is worth at least as much as I paid for it then if not much more than that. I saw a similar one going for $2500.
I hope not though, perhaps I can pick up a H100 in a few years if they get sold on the open market.
why would anyone sign such a contract?
If you want to switch back to on prem, there's probably a way to structure acquiring hardware so it doesn't break the contract. Maybe you lease it, maybe the purchase happens through a related company, maybe there was no way for the contracted cloud to find out...
Where folks (often) get lazy is the resulting math over what the real bean-counters care about (but are too lazy to check often).
In a past life, I worked on costing models for a Cable/Fiber contract house, to help the company decide 'what was profitable to keep in house' versus 'what do we subcontract' (sometimes that could even mean we just 'rented' a machine and had a qualified operator using it, based on that employee's hourly rate and expected L2R for taxes... so many spreadsheets...)
And from from my 'I don't know all the factors for this but I've seen how people screw up the big ones' view (and frankly, I'm guessing a lot of us have seen and dealt with the same category of 'bad math' around outsourcing IT work...)
An on-prem data center means:
- You need to account for electricity costs - i.e. CA vs midwest electric rates.
- cooling and power backup capability - Smaller factor but real
- personnel cost - e.x. there's probably cases where a smaller org could be better off with 'on-site' server admins that have other roles based on local wages. Kinda case specfic but it's a case.
- whatever the 'space' holding the stuff costs
- Sardonic take :Hey, let's have another unused meeting room instead! (e.x. In the case of on-prem shops that simply fled to AWS in their migration from VMware)
- the cost of licensing whatever is running
- In defense of this, In one of my earliest IT lives, AWS handling the Oracle licensing for a DB was a *huge* win as far as making it as easy as possible to ensure whatever was going on we couldn't have the Oracle licensing folks 'ding' us on whatever infraction occurred between reviews (that could not be understood by the majority of the company, often including the accused. I was never guilty but I saw it happen to others.)
- OTOH I know lots of folks who just want to be lazy about what they have to document.
Still, IMO a lot of orgs don't do the right math around these decisions, or just buy into the 'Well trends can change' as though they can decide as an org they need to suddenly triple capacity in a month and it would be able to organically happen in the first place.Frankly, the orgs that 'might' need that either have their arch set up where they are in cloud, or they are onprem but can scale to cloud if needed in interim.
It will be perfect for stuff like GPU-accelerated query engines, "classical ML" and every other CPU-based workload that could conceivably be offloaded to GPU
You can get used 16GB P100s on AliExpress for ~$100 if you want obsolete GPUs. Allegedly new AMD BC 250s are only slightly more.
I've looked at this some but I already have a GTX1070 which is only supported upto CUDA 11.9.
That's precludes some interesting modern optimizations out of the box. I've spend a lot of LLM tokens backporting some things, but I'm really not sure the hassle is worth it.
New hardware is just better. I think in maybe 5 years when supply and demand are back in equilibrium we are going to have some killer technology for decent prices, and 15yo H100s won't look attractive.
Is the idea that previously maintaining GPU programs was expensive whereas now AI makes it cheap? If so, I could buy that line of reasoning.
Maybe relatedly, I expect (hope) the hardware manufacturers will ramp up supply in the meanwhile which would also put downward pressure on GPUs. Right now though this hardware crunch is making me sad, not even because of GPUs but also because of general memory / disk.
By doing that, you know upfront what the value of your used hardware will be at the time you decommission it. It removes a lot of the risk for buyers in a volatile market.
I don't think they need some special protection against this kind of contract.
Right, so they're not voluntary.
1. There is one supplier, so you have no choice. 2. Even if you had a choice to sign the contract, this still means that it's not the same as a trade-in, because trade-ins are always voluntary, but once you have signed the contract, a right of first refusal is not.
In general, the "you chose to sign the contract" argument is a poor justification for bad contracts. If the contract is bad, it is bad regardless of whether you chose to sign it.
As businesses are expected to be more informed and equal in the negotiations
Maybe someone could start a business buying up and rehousing these.
And precisely because it's such a huge headache to do yourself, I think a small company could make a nice business wrapping up used datacenter cards in that sort of server.
The electricity prices are relevant because if you paid $0 for your H100 and didn't use it a single time, you could buy millions of tokens in inference just on the electricity it draws while idling. If you can't keep that thing saturated through the night, you're probably underwater overnight. Likewise, it's too small to run even the frontier open source models so you need to be able to live with worse models.
Max power matters because you aren't going to run many of those H100s before you blow breakers in most houses. Newer houses in the US are 15A service to non-kitchen breakers, so 1650W (that might be peak rather than continuous, not sure). If you're plugging that into an existing run, you could maybe run 2 before you start blowing breakers? You can't just plug 4 H100s into the wall in a normal house.
Maybe I'm wrong, though. I'd be curious, it'd be neat to run my own inference for something more than what'll run on a 3080.
But I never assumed it was to be cost competitive at current token rates. I assumed it was for the same reasons people might use open source hardware. Freedom to tinker, etc.
A electrician can plop in a electric car charger for instance, that is a 40-50 amp circuit.
if you haven't heard a 5u server intended for a datacenter rack come to life it's quite the experience. Sounds like a plane taking off.
Until the thermal management kicks in, they sound like jet planes. When the thermal management starts and assesses the required cooling, it’ll throttle down the fans to reasonable levels.
That is, until the moment you push the machine to its limits. When then happens, you might get back to the same levels of the boot time, but it’ll require you to push everything to the max - CPU, memory, storage (all 24 bays) and so on. For a normal user, there is a lot of room and it’s virtually impossible, even with a dozen of Teams windows open.
If these do end up in Home Labs, it's going to be a server rack in the basement and not the server rack at the other end of the room.
Helicopter. I live in Yeovil, Somerset, UK - there's a helicopter factory just down the road. I had a IBM "AS/400" or whatever they are called now in our computer room rack for a customer and it made nearly as much noise as everything else put together. It was clearly tuned for start up noise to impress because they would fire up in sequence, rise to a crescendo and then slow down in sequence to just a din instead of painfully loud.
A switch or PC server on boot will normally run up cooling fans instantly to max as a default protection mechanism until the "OS" has started and sensors read and then the fans will slow down to deal with the actual thermal load.
If you switch off your air conn, it gets noisy, quickly. Recently in the UK we are seeing routine temperatures around 30C and we broke 200 odd year records for temperatures a few weeks back. I know its even worse elsewhere but our infrastructure is not designed for this. Here we are at the same latitude as Calgary AB!
These things get hot and are fussy about their requirements.
In North America we are on 120V, making a standard 15A outlet only 1500W max, and something like 1200W sustained. To use higher wattage appliances, we have to upgrade our outlets to 20A (2000/1600W) or up our voltage to 240V, but that carries a different set of plugs and outlets as well.
You’re limited to 24kW per leg if you ran everything only on 120V, but any really large server is going to be running off 240V anyways.
It's 1800W for short periods and 1500W sustained.
Realistically, a 15A breaker won’t trip on a 2,300 watt load for at least a few minutes, and often not until a few hours. A listed breaker is expected to trip in several seconds to 3 minutes for a 3,600 watt continuous load.
Take all those numbers and go up by 33% for a house or apartment with 20A circuits (quite common). Go up by 667% for an oven plug.
As anyone who takes care of rentals during winter knows, a typical circuit can tolerate two 1,500 W space heaters without tripping very much (although this is unsafe and a bad idea).
Most homes in the US have 200 amp service. That's 48 kW of power available to the house, though code states that you can only pull 80% of that continuously, so 160 amp/38.4 kW.
The 120V misconception comes from the fact that it is delivered as split phase on two 120V legs. Any competent electrician can run a 240V, 50-amp (40 amps/9.6 kW continuous) circuit to any room in a house. Plenty of homes have these circuits for electric ranges, EV chargers, or RV power outlets, and there's nothing stopping anyone from having one of these circuits installed in whatever room they'd like it installed in.
Heck, my old landlord and I installed one ourselves to provide an outlet for an EV charger in the garage of a house I lived in about a decade ago.
400 amp and even higher service is available, though it is fairly uncommon. Plenty of large houses will have a 400-amp service, though, which means you can double the power numbers available that I mentioned above.
We have 50 amp service running a hot tub and another 50 amp service line to charge an electric car.
The quotes are in since the difference between the phases is 230V, but in practice it is almost the same as a 40Ax400V single phase connection.
It isn't, and you're right. This is just a long-winded article by someone who thinks they've come across a deep, crucial insight.
I witnessed a bank foreclosure stemming from large, unpaid loans to a lumber mill. The bank absolutely didn't want to take possession of the operation, but had no choice when the founder decided to call it a day.
The bank had no idea what to do with finished lumber sitting in the drying ovens, let alone the entirety of mill infrastructure itself. After struggling to find a buyer, they hired the founder as a consultant to handle liquidation. The same would happen with a repossessed datacentre.
Auction prices account for that risk, you don't get to do all of the verification you do for normal purchases.
You can't get the Rubin, or even the Blackwell, so you will pay for the H100 but this won't last if fabs ramp up capacity.
What's as or more weird is how much hardware is backordered, and how much live hardware is allocated, but waiting on facilities for operation. And how many facilities are years behind at this point already... all on various credit and dept swaps between all the involved companies... it's not just a balloon, it's a house of cards balanced on a balloon.
The increase in PFLOPS/dollar has continued accelerating, a lot from process, but also a lot by simplifying the architecture- if you had placed an H100 worth of transistors on a CPU-like architecture, you wouldn’t reach the same peak performances.
but I generally agree, people put a lot of faith in the exponential leaps vs the exponential space.
You tell them we're not living on mars any time soon and they'll bring up christopher columbus.
This is all downstream of ASML who is putting out tens of EUV machines per year.
If everyone ramps up then in the best case everyone has the same sized slice of a bigger pie. So in theory it makes business sense. But the more realistic possibility is that you end up with oversupply, crash the market and everyone loses. This is what normally seems to happen with DRAM.
A key thing to understand about the gold rush is that it was not a major economic event, or at least nowhere near as big as the participants thought it would be, hence the tradegy.
The AI gold rush is different in that there actually is a mountain of "shovels" large enough to flood the global market quite severely.
Not sure what that means for the metaphor, but there you go.
Hmmn. All of this information should live in the monitoring system, in which case any frontier model will be able to get to grips with it in short order. It feels like the author doesn't really fully understand the changes brought about by the systems they are writing about.
But it can still not run truly capable coding agents at reasonable speeds right?
That doesn't make sense for the cost of quite a nice new car. If you want to spend that much, build yourself a proper workstation at least. Clearly a computer with the form-factor of an Apple TV is not going to be ideal for getting the highest performance for your buck.
Is it just wannabe vibentrepreneurs with too much spare money hyping each other up?
Sincerely, perhaps there will be a new flood of crypto farms, which I suppose would drive the prices of crypto down. Other ideas for what changes are enabled by massive supply of cheap GPUs?
You go to ebay search for a used GPU. You get a price.
Neither is used servers a new thing or used routers. There are established used server companies.
Actually, we do, people offer them to me all the time. A used box of MI300x is $257k. "There is no GPU futures market"... actually there are a few of them that people have pitched to me.
This article is a lot of words from someone who isn't actually buying or deploying compute. My point is... take it all with a grain of salt.
I was complaining that it's obviously incorrect that nobody knows what used GPUs are worth, not about latchkey.
I have NO idea why anyone is upvoting a post titled "nobody knows what a used GPU cluster is worth", that is a WILD claim.
Most of the people who shit talk GPUs don’t even understand why it relatively speaking doesn’t matter that much if I’m on ampere or Vera Rubin for full precision models. These same people are wildly misinforming investors.
To quote the ADATA ceo, “don’t talk about an AI bubble until 2040”
There are likely management service agreements from xAI proper -> SPV to cover precisely what the author talks about. Clearly, xAI could play games but without seeing the docs (which are not public), it's very difficult.
This article's basic point is right though. On the other hand, the LTV of this deal was approx 50% debt-financed (not too high; very much depends on the "V"). At 12.5%, it's not as if its being priced as a high quality asset.
Overall, substack post was too bearish. The wider point is that there's a lot of froth tied to what has now become systemically opaque - namely the circular deal flow that every hyperscaler, nvidia, neoclouds and friends are now engaged in. When the proverbial hits the fan, that stuff will be difficult to price and find few willing buyers with the competence to underwrite.
The systemic issues are the bigger concern than one specific deal imo.
slow and awkward, best market match: The auction. sell to highest bidder.
faster and more customer friendly but poor market match until a lot of units sold: The store. guess price, adjust up or down to reach sell frequency desired.
fast and good market match but takes a knowledgeable customer base: The reverse auction. Start with price too high lower it over time until it sells.
>Start with price too high lower it over time until it sells.
These are the same strategy.
No, because I am retiring next month!
1. An enthusiast had a project to get a V100 working on his PC [1]. This was a ~$10k GPU 10 years ago. It's now sold for scrap;
2. The A100 came out in 2020 and cannot run a large model like DeepSeek v4 Pro. It can run Flash. You need a 16xH100 cluster to run Pro and that's a ~4 year old GPU and AFAICT 8xB100 or 4xB200;
3. We're about to roll out R100/R200s.
I'm surprised that NVidia is moving to a 1 year product cycle (per this article) because the big question I've had is what's that going to do to existing investments in GPUs. Why? Because if 4xR100 can do the work of 32xH100 then that's a massive advantage in performance-per-Watt, which I think is going to be the only metric that ends up mattering.
In addition to raw power, new capabilities are developed and come online. For example, certain smaller, more efficient quantization methods just didn't exist on older hardware.
Oh, another thought from this: a 9% annual failure rate just goes to show you how ridiculous the idea of orbital data centers really is. Orbital DCs were always just a pump-and-dump scheme for SpaceX's IPO.
Currently it gets expensive to run models larger than ~31B locally. You start to need some pretty expensive hardware. That's going to change. I don't expect we'll be running 1T+ models on a Macbook Pro within 5 years (at reasonable inference rates) but I think people today will be shocked at what's being run locally in 5 years and that'll easily be 100-200B+ models.
[1]: https://www.hackster.io/news/hacking-a-server-grade-nvidia-g...
Which has this anecdotal data point:
>... when I left Paperspace in mid-2024, our M4000 GPUs, nine-year-old GPUs, were still consistently utilized at near-total capacity. That’s not a typo. Nine. Years. Old. Still booked, still working, still generating revenue.
Also, I won't claim to understand accounting, but in general it seems it is advantageous to accelerate depreciation schedules for high CapEx industries because they lower taxes: https://leyton.com/us/insights/articles/what-is-accelerated-...
As such you'd assume these cloud providers to want faster depreciation of their GPU assets rather than slower? I suppose in this case they do have an incentive to show bigger revenue numbers, but there seems to be a trade-off here that is not being discussed?
In my experience, AI is easier to read than this was.
> These are not catastrophic events. They are the steady state.
> There is no GPU futures market, no standardized residual value curve, and no way to lock in a forward rental rate. The premium is is the price of underwriting in the dark.
The headings are also AI like, a lot of essays before usually did not have titled sections but now they do and they all feel like these.
In addition the diagrams themselves look pretty AI generated.
And I feel that AI hardware is going to rapidly depreciate faster than even typical PC hardware
I built and operated an application that used 1-2000 L4 GPUs in production for a couple of years. Long-term GCP reservations running in a GKE cluster. At that scale we had a few GPU failures per week, and once saw 3 in a day. Nearly every failure required an engineer to manually intervene to get rid of the bad node, and then file a ticket with GCP support as they requested.
NVIDIA GPUs throw an “XID” code when they fail, which can be seen from serial port logs for GCP compute nodes (not through k8s!). If you’re lucky, it fires right as your application starts to fail, but there’s often a delay of several minutes. Even when you get an XID, by default GKE only responds to one or a few of them. You can expand the list via configuration, but the reconciliation loop is so slow that might still take 10-15 minutes during which a pod is puking errors and someone might be getting paged.
They’re working on it, and we never saw a single XID after we migrated to H100s, so the situation is improving. I imagine other clouds are even worse, though iirc Azure was leading some effort to improve k8s node problem detector to include accelerator problems so maybe they have a better story.
Training workloads are rather more sensitive to a node failure given that many modern training runs (including SFT etc) need multiple nodes where topology matters, and they might not have another e.g. 8xH100 box in the right place when one fails.
I’m mainly interested in getting some DDR4/5 and RTX5090s on the cheap :).
Without other market influences, that is a >90% expected discount when the over-provisioned market must inevitably self-correct.
If the Market follows what Samsung/SK Hynix did to the South Korean exchange this week, than the "AI" bubble will hit harder than the dot com crash.
I like the Shrek Movie correlation theory, as they always happen just before Debt-backed investors get hit hard... And the new film is due out in 2027. =3
Can you tell us more about this? Or some link
Over the past few weeks Kospi index has tumbled 25% since its June peak, resulting in a $1 trillion wipeout and its chipmaker duo have both lost at least 30% of their value. There have been days of near 10% plunges followed by sharp rebounds driven entirely by shifting confidence in whether AI spending is sustainable.
Patrick Boyle gives a summary of the situation, but not the underlying Shrek issue:
Anyone know how these get caught ultimately?
(And other things, that we know partially but not fully-- like what failure rate for current generation parts will be under this loading).
So we have big uncertainties about the revenue, moderate uncertainty about the proportion of the asset that will survive, and some uncertainty about what operating costs will be. It's difficult to turn this into a residual value.
Finally, the whole "operating the big facility" thing is not likely to be plug-and-play for a new technical team following a default. How much outage/disruption ensues?
We are increasing production of GPUs and RAM a lot.
If anything, you might have to pay to have them disposed of, they don't really have any meaningful used eBay market outside of the randos that want to do high end extreme local inference in their basement.
Also, as for RAMmageddon, the inference SBCs that all of the AI bros bought don't have DIMMs, they're not even the right chip: its all GDDR and LPDDR. The only DDR DIMMs being consumed are for regular non-inference machines that help run the business and service infrastructure behind the scenes.