The article talks a lot trash about AMD and ROCm but vulkan works fine too. In fact from a datacenter GPU standpoint there is an AMD option called the V620 available on US EBay that I was able to haggle to $350, with 32GB VRAM, 512GB/s bandwidth, and runs the same Qwen-3.6-27b at about 20t/s. I would argue that’s even more cost effective.
It requires a few of the same fan shenanigans this guy did but there is no need to pull specific past software versions to make it usable in LinuxThis is nice to know. Thank you.
Honestly even for the prices of around 500$ that I’m seeing it for it looks like a pretty good value to get 32gb of vram. I see it says 300w on AMD’s product page for it does it have any way of power limiting the card to get more efficiency/less heat?
I spent a lot of time researching and testing different methods for that, the only thing that worked was LACT in Linux. Using that I was able to undervolt 100mV and GPU power usage dropped about 10%. On my B450 ITX board with a Ryzen 2400GE CPU the entire system pulls 30w idle from the wall, and about 300w inferencing with VRAM filled. (330w before LACT).
My fan solution ended up being to buy the 80mm 3d-printed shroud off ebay, the fan that came with it was super loud so I switched to an arctic p8 Max, and control it with the motherboard targeting a t-sensor header with the probe attached to the backplate.
Here I was expecting a graphics demo to blow our collective minds but instead I got a story about a local LLM for cheap. It is <current year>. I should have known better.
Were I the author / tech cobbler here, I’d be concerned that too much time with an LLM, local or otherwise, might erode or dull my apparently fairly sharp reasoning and tech skills. (Clarification: Not my sharpness, theirs. I’m a potato.)
Other thoughts: For a minute I thought this whole thing was a tribute to, or a troll in the manner of, that one Redditor that always spun their stories around to being about their dad beating them with jumper cables.
Also, my old PC developed an issue like the warm reboot problem, except with the network interface. I couldn’t just restart, I had to power off and back on. I never did bother to find out whether it was early signs of hardware failure or whether it was an old hardware / newer kernel mismatch.
They cannot run graphics
The change of the meaning of the G in GPU from “graphics” to “general” is even less well documented and used than the “V” of DVD changing from “video” to “versatile”.
Indeed it only occurred to me what it must have changed to and to go looking to confirm after seeing your comment.
And frankly they ought to have changed the name to something like “MPPU” if they wanted it to stick (massively parallel).
Back in my day, we memorized logarithm tables and we liked it!
RogerSimon10! (The jumper cables guy)
Unless maybe he was the similar but later hell-in-a-cell guy? I don’t remember anymore.
Not an AI guy, but I do like using niche hardware wrong to get results cheap. Can anyone tell me what this would be like for gaming or general computing? My 1660 super was a budget pick when I got it back in '18.
A lot of data center GPUs don’t have display outputs or cooling. So you would need to figure out how to cool them.
Cooling wouldn’t be too hard, I’m currently overengineering an Xbox 360 for that so I’ve done a good deal of learning. Getting of to display video is outside my skillset though.
A cursory search says somewhere between a 3060 and 4060, which seems about right. Games that parallelize well across the GPU cores will benefit, though that benefit will be niche if it exists at all. HBM2 memory is weird for gaming.
You’d see some benefit for sure, but this really is better suited to parallel computing, given the emphasis on CUDA core counts and memory bandwidth. You may also find you run into some latency, since it requires selecting the V100 as your primary GPU but a secondary card as the video output. Integrated GPUs work well for this, since your motherboard may already have this ability. Otherwise, you could use it in tandem with a 1030 or similar to get display out.
So I’m picking up that it’s not a viable alternative for the desperate?
Probably not your best option, no
Yeah but the 0 point doing this unless you want to run AI models for some reason. These GPUs can’t do video game graphics so this isn’t a solution to the GPU shortage.
This is a bit like me writing an article about NASCAR, now I can turn left whenever I want. But I haven’t magically acquired a functional vehicle for a fraction of its value. I’ve purchased a second hand specialist product that is usually useless outside of that environment.
Why tho?

Fair enough
Ever watched bringus studios? Man played games on a drive through computer. As the old saying goes all hardware is good hardware if you know what to do.
I cant get past the new hair cut. He looks like a bond villains lacky now lol
Seems like an awful lot of trouble to save $100 not buying a 5060 Ti that also has 16GB.
448 GB/s memory vs 900 GB/s and at a cheaper price.
Good point, that should theoretically be double the decode speed, though not sure how the prefill would differ since it’s mainly compute-bound.
I had no idea you could get a 16gb card this cheap! TIL.
It’s pretty low performance, though… My card from 8 years ago nearly matches the passmark rating.
its for running ai models
The 5060 Ti does not support Nvlink though.
Sure it’s got a lot of VRAM, but the 4080 has five times the compute power.
You’re not getting a 4080 for $200 though…
Hmm… I have a 4080 from 2022. I wonder if I can get this thing too, for $200, and use them both? I hate only having 16gb VRAM.
Would be nice if I could use them both at the same time. Like, make the 4080 recognize the other device as additional VRAM
Not sure how well that without work even if you could bridge them properly to share their vram, as the latency from the other gpu will be pretty high compared to the local vram. Frame times won’t be that great is my prediction. It works ok for LLMs because they aren’t a realtime compute task like gaming graphics is. If an LLM takes an extra 0.12s to compute its result, you don’t notice, but if a gpu misses a frame deadline by 0.12s, that’s a stutter that represents less than 10 fps.
I believe that’s why earlier attempts at dual gpu systems mostly fizzled out (plus cost concerns). You can get some pure acceleration if you can fit the entire working memory onto both GPUs’ vram (so your total vram is effectively min( GPUA_VRAM, GPUB_VRAM ) rather than GPUA_VRAM + GPUB_VRAM, though that also requires all pixels to be independent of anything calculated on the other GPU, other than maybe post processing effects that could be handled on whichever GPU is handling the display.
That’s not to say that you can’t get extra performance out of multi-gpu setups that don’t just mirror their RAM, but it’s more complicated than “sum of the capabilities of each GPU”.
Yes, that’s basically what the article is about. They run the LLM across both GPUs.
But that’s a feature of llama.cpp. SLI doesn’t really exist any more, and NVlink requires a specific setup, which the 4080 is not part of (the 3090 was the last consumer one, apparently). So you couldn’t pool the VRAM.
That’s fine I just need to display pictures of your mom (they are very large) (/s)
eBay has some rad Chinese mezzanine boards for these guys too. Nvlink works and everything lol

Cool mezzanine boards, but can we talk about your dope af custom jig for offset mounting arbitrary boards?
I appreciate the kind words! Until recently, my day job was CAD monkey. I wanted to consolidate the hardware I was cobbling together, found a cheap 8 GPU mining rig and some extra 2020 extrusions to play with.
The seller for the mezzanine said it followed the mATX mounting hole pattern, (it does but is ~75% in the width and doesn’t use all the points). I have a tendency to overthink designs and kind of stalled for a bit before finally taking the plunge and whipping up these struts. I used some brass heat set inserts to accept the standoffs and everything pretty much went right together lol

Damn, even cooler than I expected 🎸 Thank you for the extra details and picture!
The last time I saw a passively cooled GPU was the '90s.
I am using PTM sheets and they idle at decent temps, though I did ziptie some high CFM fans behind them lol

82db? Wild. I guess they don’t really care that much about noise in data centers though.
If it doesn’t sound like a jet plane taking off, is it really a server at all?
Hey, my home server netbook can play
jet_plane_taking_off.opus.
They mostly don’t, but this is also not how the datacenters cool them.
A datacenter will either have an open water loop, or an all in one taking heat to a more advantagous place for a radiator to be, or at the very least better managed airflow with bigger fans and more specific air baffles.
This thing has no such luxury and has a small area and unknown broader thermal context, so screaming it is to make up for the limitations of the scenario.
No they really dont. Big ass fans running 24/7 to help the small fans running 24/7. It all blends into an easily ignored drone though just dont try to have a conversation in there.
What?
Sorry, I think I might have hearing damage from all this constant noise.
The main thing that itches me with the V100 is the fact that given that pascal is about to be EOL, a 2017 card is probably soon next
The ampere class is soon to fall out of relevance as well.
About to be? Pascal cards no longer work in TrueNAS scale due to Nvidia dropping driver support
deleted by creator
This isn’t the smart way, though.
What the homelabbers do (at least before the RAM crisis) is buy Xeon/TR/EPYC boards on the cheap, and then run gaming GPUs for hybrid inference.
This is what I do. I run MiMo 2.5 at 8-10t/s on a 7800X3D/RTX 3090/128GB CPU RAM, more with Dflash. That’s a 300B model: it’s not even in the same class as Qwen 27B, which is what the dev in OP’s article is trying to run.
And this is small-time: setups with 4-8 memory channels can run stuff like Kimi or Deepseek Pro, even faster. Or they can run smaller LLMs with quantization types that are very fast on CPUs, and get crazy speeds.
…And besides, Qwen 27B can run fine on a 4080, with the right framework. It will fit in 16GB as an exl3.
Not that this isn’t a cool hardware hacking project.
…But it’s kind of the wrong approach. It’s about 2 years out of date, as MoEs are king in LLM land now. RAM is horrendously expensive, yes, but so are most used V100s, or used 3090s.
I never expected used DC GPUs would make it to eBay.
Lots of used DC gear makes it to eBay. This is how homelabbers survive.
I know that and I use that. But recently market for switches and servers dried up. I thought they were shredding DC GPUs too. Most of them die after 3-5 years anyway.
V100s are ancient in a market obsessed with the very very latest.
The datacenters are unlikely to bother directly with eBay, but they have asset recovery companies that will take the stuff off their hands and seek buyers, including over eBay.
Hmm, there are some AMD Instict listings on ebay, after all.
Is nobody extremely tired of the AI slop articles??
The guy is explaining how to hack around GPU fans which were never meant to be run at lower speed through jumper cables and you complain about slop?
This “guy” isn’t explaining anything. The entire article is written by AI. It’s the same bland writing style everywhere now.
I don’t know whether you read the article. It was quite interesting to me and gave me some good ideas. I not going to replicate what he did, but he has shown me a methodology I could apply to do other things.
Who gives a shit? You ever read a legal document? They’ve had the same writing style for 70 years.
Actually if you’re interested in running a local llm it’s quite a good article
I thought parts felt like AI and parts felt like a British human writing it.
AI(?):
But here is the thing: this is a Volta GPU with 16GB of HBM2 memory, 5120 CUDA cores, and I picked it up for about £150 on eBay. The compute is still real. The VRAM is still real. And the memory bandwidth is where it gets genuinely surprising.
Human (?):
So I shoved some jumper wires into the connector and jammed the other ends into a spare fan header (turn your volume up)
I suspect he’s used some AI polishing tool to turn his draft into a publishable post.
It’s so clearly written with the aid of an LLM that I’m embarrassed for all the people downvoting this comment. Just because the basic outline of the information presented and the sequence of events comes from the real experience of a real person, it does not mean that the real person wrote the article themselves. Seriously, how many real people write like a bad 80s detective movie script? It’s literally a parody trope at this point.
Honestly, I just don’t care.
I am a monkey, I live on a rock, in the middle of an infinite empty space, that none of the smartest monkeys on this rock can figure out.
I am part of a species that has spent its entire history killing itself in the most horrifying and ghastly ways for the most benign reasons that the greatest war approaches us and everybody is so accepting.
The planet I live on is slowly becoming inhabitable to point where the thought of kids is pointless. And banks have used to many money glitches that currency across the globe is going into hyperinflation.
I really really cannot illustrate how much I do not give a fuck about the cadence of written words in that small space of relief I get from the inevitable destruction looming in a completely absurd and pointless existence.
word
This article ISN’T AI slop.
It has nothing to do with the subject of the article. Are you blind? It’s the exact same bland, AI regurgitated writing style as every other AI written article.
For well-known authors with a large corpus of prior writing, it’s simple to run a style match analysis, if you wish to demonstrate that they lazily used generative AI to write a particular piece.
But for all the knee-jerk “ai slop” accusations I see on almost every post now, the only justification most offer is that it’s obvious, which sounds like “I can tell by the pixels.”
I detailed several of the indications in another comment. If you’re too blind to see the very obvious indicators then oh boy. That’s what makes it AI slop. They don’t even bother trying to hide it. It’s like bad CGI, there’s no need to say “I can tell by the pixels” because it’s transparently obvious.
Yes but this isnt one of them.
It’s literally AI slop. Ignoring the subject of the article (which is whatever), the article is clearly entirely written by AI. The AI writing style is horrendous. “It’s not blah. It’s blah.”, “it’s genuinely <insert adjective of amazement>”
Its not AI writing. It reads exactly like someone casually talking. Its consice and to the point. There is no “its not x, its y”
He has been blogging dor dam near 6 years you can go read his older posts the writing is consistent.
But here is the thing: this is a Volta GPU with 16GB of HBM2 memory, 5120 CUDA cores, and I picked it up for about £150 on eBay. The compute is still real. The VRAM is still real. And the memory bandwidth is where it gets genuinely surprising.
It’s AI slop. You’re blind.
So much so…