I just racked some consumer GPUs in my garage (6x r9700s), cranking 24/7 pumping 20M+ tokens a day towards my goals building compilers, custom operating systems, reproducible build debugging, kernel hardening, novel confidential compute tech... harder problems than anyone I know using cloud LLMs to solve.
I have still never paid for tokens AND have session privacy.
Of course we don't NEED to pay for the cloud, but my company is doing this for me and I can't quite justify spending even 4-5K on a compute cluster in my garage, as much as I would love to.
[0]: https://docs.z.ai/devpack/overview#estimated-token-allowance
Also, when I want 1M context I have 118GB of usable vram on my Strix Halo, or I can combine 4 r9700s and have 128gb and can run 1M context models like laguna or deepseek, but in practice lots of smaller sessions is better for my workflow in most cases.
Qwen 3.8 27b is smart enough that all I want is to speed-max and paralell-max on that.
But we are not talking about the same product anymore. Your $12k homelab does not provide the same product as a $20/month subscription (to say nothing of a $200/month one). And the trade-offs of the homely may be worth it (or necessary) for you, but it may not be worth it (or even be feasible) for others.
Personally though, I would sooner trade my car for GPUs than let a third party be in control of the tools I use to do my job.
You're right people don't need the subscription, they just need a far more privileged life. One were dropping $12k instead of $20-100 a month is an equivalent financial strain.
Also factor in cloud LLM prices are subsidized by you giving up your sessions as training data with no ability to opt out.
I know that 7M a day is a pittance for my use cases, and given that you have multiple cards, it wasn't enough for your use case either.
Also, unless your sessions are end to end encrypted to a secure enclave, then they are living in plain text -somewhere- and privacy policies tend to change when money is left on the table, if blackhats do not get to the data and sell it first.
2. Use free daily tokens from opencode or similar to set it up for you with a coding agent like jcode or crush
You're very right to be wanting to use open sourced models. You're delusional for thinking that buying your GPUs directly saves any money. You are not a cloud service provider, don't act like one.
Also, $1200 for a GPUs that can produce ~7M tokens a day of Qwen 3.8 27b at 80tps is faster and cheaper than any major provider can serve a model of that class as far as I am aware. Pays for itself pretty fast.
The whole system is designed to be catastrophic. There is no rogue agent, or anything going off the rails, it is behaving exactly the way one would predict.
And they now announced that same model in their API, acknowledging it is way more difficult to monitor it. That company should not be in business
I am hoping we get an administration that will not accept his bribes so he can finally be held accountable for criminal negligence and fraud.
Discussion of Ronan Farrow’s New Yorker article (900+ comments): https://news.ycombinator.com/item?id=47659135
Disagree. Even if OpenAI didn’t exist, this problem would still exist. It is best to assume and guard against rogue AI, which is what the essay is about.
Ban OpenAI, and crown Anthropic as a golden child of safe AI development? Surely no company is "trustworthy" with this power and all future AI capabilities are dangerous, so all AI research needs to be paused everywhere, requiring the equivalent of a nuclear non-proliferation treaty between the US and China. So post-treaty we're trusting the US and Chinese governments, their military, and the billionaires closed tied to them to not to continue developing the magic thinking machine that can solve any problems, wins every battle, and is the crux of the economy? With the US track record of breaking much less significant treaties?
I'm not pro or against pausing AI development. I just don't understand what the theoretical stop button looks like.
Sam Altman can be a bad person (ask his sister about him!) and deserves to go down yada yada without us needing to try to kill the goose which is about to fling us into a golden age.
The situation is that we have to face a bet that is forced upon humanity by a few actors: in all cases we know there will be a very high cost to current society, there is some chances that we get some benefits.
You cannot think about the situation if you decide to set a probability of golden to 1, and ignore the externalities, risks, and direct negative impacts.
It’s an insane risk to go all-in with, but that’s what the AI cultists want.
Honestly you shouldn’t be close to any form of decision making if you cannot have such a basic level of analysis
I assign the probability of extremely good outcomes as being overwhelmingly likely. The amount of harm done is a necessary part of it and is acceptable in much the same way that the additional gun crime as a result of the 2nd amendment is.
Every disease uncured, every amount of work done that was unnecessary, was an affront to human dignity. I want the machine god and I wanted it yesterday.
I am finding that the world is and likely will continue to be full of carbon-chauvinists for a long time.
You dont get to wave utopic vision and loose ethical obligations.