Posts / ai
768GB of VRAM and the Beautiful Madness of the Local LLM Crowd
I’ve been down a rabbit hole on r/LocalLLaMA again. It’s my equivalent of other people’s fishing forums, I think, except instead of lures and knots it’s people arguing about VRAM bandwidth and whether a Chinese mining card can be soldered into behaving like a proper GPU. Yesterday someone posted a build: twelve secondhand CMP170HX cards, 64GB each, strung together with fibre to a second rig, giving them 768GB of usable memory. Total cost, they claimed, less than a single RTX 6000 Pro. The post title alone made me sit up: that’s the kind of number that sounds like a typo until you read the comments and realise it isn’t.
What got me wasn’t really the hardware. It’s the attitude. The original poster pre-empted the criticism before anyone could type it: don’t bother telling me the API would be cheaper, don’t bother with the electricity bill maths, don’t bother with the “it’ll take you 52 light years to break even” line. And true to form, someone in the replies summed up the whole ethos in one sentence: here at localllama, we only care about passion, we care about local setups, we couldn’t care less about how much electrons are getting fucked up in the process nor do we care about how much bills have gone away. That’s not a defence. That’s a mission statement.
I get it, honestly, more than I probably should. There’s something in the hobbyist mind, mine included, that resists the sensible answer. I could rent GPU time by the hour and get better throughput for less money than the power bill on a home rig alone. I know this. It doesn’t matter. The appeal isn’t the spreadsheet, it’s the ownership. Nobody can throttle your model, change the terms of service on you, or decide overnight that your favourite feature is now a paid tier. You built the thing in your garage out of cards that were designed for crypto mining and abandoned when that bubble moved on, and now it talks back to you. There’s a particular satisfaction in that, the same one I get finding a $40 espresso grinder at a garage sale that some barista would sell a kidney for. Different scale, same itch.
But the bit that stuck with me longer than the hardware porn was one throwaway reply, half joking: someone admitted that a lot of these builders have spent, over the years, far more chasing the next upgrade than any “ROI” argument could ever justify. I’m a mad old wizard and that shit costs dolla dolla bills, y’all. That’s honest in a way most tech spending isn’t. Nobody there is pretending this is efficient. They’re pretending it doesn’t need to be.
Which is where I start feeling the tension, because I do care about the electrons, even if the forum culture says not to. Data centres already chew through a frightening and fast-growing slice of global power, and every hobbyist rig strung together in a spare room adds its own small drop to that. I don’t think a dozen mining cards in someone’s garage is what’s cooking the planet. But the same impulse, “run it bigger, run it locally, don’t think too hard about the draw,” scales up terribly when it’s not one enthusiast but an entire industry building data centres the size of suburbs to chase the next model release. The garage tinkerer and the hyperscaler are doing the same thing for different reasons, and I don’t have a tidy way to reconcile being charmed by one and uneasy about the other. I just am.
There’s also something quietly democratic about what these people are doing, and I don’t want to lose that in the environmental hand-wringing. A handful of companies control the frontier models. The compute to train them is well out of reach for anyone without a few hundred million lying around. But running a very good, very large model on your own hardware, on your own terms, with nobody logging your prompts, that’s still within reach of a determined weirdo with a soldering iron and an eBay habit. That matters. It’s the same reason I like that Linux still exists, even though I mostly run a Mac. Somebody needs to keep the option open.
I don’t own a single GPU dedicated to any of this. My contribution to local AI enthusiasm has been running a small model on my laptop fan at full noise for twenty minutes to summarise a PDF I could have read myself. But I like that this world exists, full of people building genuinely mad rigs out of stubbornness and curiosity, arguing about tokens per second at one in the morning, openly admitting the maths doesn’t add up and doing it anyway. It’s not efficient. It’s not really sustainable if everyone did it. But it’s honest about what it is, which is more than I can say for most of the industry currently trying to sell me the same technology wrapped in words like “transformative” and “seamless.”
I don’t know where this all lands in ten years, whether it’s local rigs in garages or three companies renting us all compute by the second. Probably both, unevenly, the way most things end up. For now I’m just enjoying the spectacle of people proving, once again, that the line between hobby and obsession was never that firm to begin with.