• mierdabird@lemmy.dbzer0.com
    link
    fedilink
    English
    arrow-up
    87
    ·
    5 days ago

    The article talks a lot trash about AMD and ROCm but vulkan works fine too. In fact from a datacenter GPU standpoint there is an AMD option called the V620 available on US EBay that I was able to haggle to $350, with 32GB VRAM, 512GB/s bandwidth, and runs the same Qwen-3.6-27b at about 20t/s. I would argue that’s even more cost effective.
    It requires a few of the same fan shenanigans this guy did but there is no need to pull specific past software versions to make it usable in Linux

    • JohnWorks@sh.itjust.works
      link
      fedilink
      English
      arrow-up
      7
      ·
      4 days ago

      Honestly even for the prices of around 500$ that I’m seeing it for it looks like a pretty good value to get 32gb of vram. I see it says 300w on AMD’s product page for it does it have any way of power limiting the card to get more efficiency/less heat?

      • mierdabird@lemmy.dbzer0.com
        link
        fedilink
        English
        arrow-up
        8
        ·
        edit-2
        4 days ago

        I spent a lot of time researching and testing different methods for that, the only thing that worked was LACT in Linux. Using that I was able to undervolt 100mV and GPU power usage dropped about 10%. On my B450 ITX board with a Ryzen 2400GE CPU the entire system pulls 30w idle from the wall, and about 300w inferencing with VRAM filled. (330w before LACT).
        My fan solution ended up being to buy the 80mm 3d-printed shroud off ebay, the fan that came with it was super loud so I switched to an arctic p8 Max, and control it with the motherboard targeting a t-sensor header with the probe attached to the backplate.

  • palordrolap@fedia.io
    link
    fedilink
    arrow-up
    68
    arrow-down
    8
    ·
    5 days ago

    Here I was expecting a graphics demo to blow our collective minds but instead I got a story about a local LLM for cheap. It is <current year>. I should have known better.

    Were I the author / tech cobbler here, I’d be concerned that too much time with an LLM, local or otherwise, might erode or dull my apparently fairly sharp reasoning and tech skills. (Clarification: Not my sharpness, theirs. I’m a potato.)

    Other thoughts: For a minute I thought this whole thing was a tribute to, or a troll in the manner of, that one Redditor that always spun their stories around to being about their dad beating them with jumper cables.

    Also, my old PC developed an issue like the warm reboot problem, except with the network interface. I couldn’t just restart, I had to power off and back on. I never did bother to find out whether it was early signs of hardware failure or whether it was an old hardware / newer kernel mismatch.

      • palordrolap@fedia.io
        link
        fedilink
        arrow-up
        8
        ·
        4 days ago

        The change of the meaning of the G in GPU from “graphics” to “general” is even less well documented and used than the “V” of DVD changing from “video” to “versatile”.

        Indeed it only occurred to me what it must have changed to and to go looking to confirm after seeing your comment.

        And frankly they ought to have changed the name to something like “MPPU” if they wanted it to stick (massively parallel).

    • axmo@lemmy.ca
      link
      fedilink
      English
      arrow-up
      2
      ·
      3 days ago

      RogerSimon10! (The jumper cables guy)

      Unless maybe he was the similar but later hell-in-a-cell guy? I don’t remember anymore.

  • Postmortal_Pop@lemmy.world
    link
    fedilink
    English
    arrow-up
    10
    ·
    4 days ago

    Not an AI guy, but I do like using niche hardware wrong to get results cheap. Can anyone tell me what this would be like for gaming or general computing? My 1660 super was a budget pick when I got it back in '18.

    • tinfoilhat@lemmy.ml
      link
      fedilink
      English
      arrow-up
      4
      ·
      3 days ago

      A lot of data center GPUs don’t have display outputs or cooling. So you would need to figure out how to cool them.

      • Postmortal_Pop@lemmy.world
        link
        fedilink
        English
        arrow-up
        1
        ·
        1 day ago

        Cooling wouldn’t be too hard, I’m currently overengineering an Xbox 360 for that so I’ve done a good deal of learning. Getting of to display video is outside my skillset though.

    • Flatfire@lemmy.ca
      link
      fedilink
      English
      arrow-up
      3
      ·
      3 days ago

      A cursory search says somewhere between a 3060 and 4060, which seems about right. Games that parallelize well across the GPU cores will benefit, though that benefit will be niche if it exists at all. HBM2 memory is weird for gaming.

      You’d see some benefit for sure, but this really is better suited to parallel computing, given the emphasis on CUDA core counts and memory bandwidth. You may also find you run into some latency, since it requires selecting the V100 as your primary GPU but a secondary card as the video output. Integrated GPUs work well for this, since your motherboard may already have this ability. Otherwise, you could use it in tandem with a 1030 or similar to get display out.

  • Echo Dot@feddit.uk
    link
    fedilink
    English
    arrow-up
    17
    arrow-down
    3
    ·
    4 days ago

    Yeah but the 0 point doing this unless you want to run AI models for some reason. These GPUs can’t do video game graphics so this isn’t a solution to the GPU shortage.

    This is a bit like me writing an article about NASCAR, now I can turn left whenever I want. But I haven’t magically acquired a functional vehicle for a fraction of its value. I’ve purchased a second hand specialist product that is usually useless outside of that environment.

  • melfie@lemmy.zip
    link
    fedilink
    English
    arrow-up
    24
    ·
    5 days ago

    Seems like an awful lot of trouble to save $100 not buying a 5060 Ti that also has 16GB.

  • frongt@lemmy.zip
    link
    fedilink
    English
    arrow-up
    20
    arrow-down
    1
    ·
    5 days ago

    Sure it’s got a lot of VRAM, but the 4080 has five times the compute power.

      • partofthevoice@lemmy.zip
        link
        fedilink
        English
        arrow-up
        3
        ·
        4 days ago

        Hmm… I have a 4080 from 2022. I wonder if I can get this thing too, for $200, and use them both? I hate only having 16gb VRAM.

        Would be nice if I could use them both at the same time. Like, make the 4080 recognize the other device as additional VRAM

        • Buddahriffic@lemmy.world
          link
          fedilink
          English
          arrow-up
          4
          ·
          4 days ago

          Not sure how well that without work even if you could bridge them properly to share their vram, as the latency from the other gpu will be pretty high compared to the local vram. Frame times won’t be that great is my prediction. It works ok for LLMs because they aren’t a realtime compute task like gaming graphics is. If an LLM takes an extra 0.12s to compute its result, you don’t notice, but if a gpu misses a frame deadline by 0.12s, that’s a stutter that represents less than 10 fps.

          I believe that’s why earlier attempts at dual gpu systems mostly fizzled out (plus cost concerns). You can get some pure acceleration if you can fit the entire working memory onto both GPUs’ vram (so your total vram is effectively min( GPUA_VRAM, GPUB_VRAM ) rather than GPUA_VRAM + GPUB_VRAM, though that also requires all pixels to be independent of anything calculated on the other GPU, other than maybe post processing effects that could be handled on whichever GPU is handling the display.

          That’s not to say that you can’t get extra performance out of multi-gpu setups that don’t just mirror their RAM, but it’s more complicated than “sum of the capabilities of each GPU”.

        • frongt@lemmy.zip
          link
          fedilink
          English
          arrow-up
          6
          arrow-down
          2
          ·
          4 days ago

          Yes, that’s basically what the article is about. They run the LLM across both GPUs.

          But that’s a feature of llama.cpp. SLI doesn’t really exist any more, and NVlink requires a specific setup, which the 4080 is not part of (the 3090 was the last consumer one, apparently). So you couldn’t pool the VRAM.

    • Kairos@lemmy.today
      link
      fedilink
      English
      arrow-up
      14
      arrow-down
      2
      ·
      4 days ago

      That’s fine I just need to display pictures of your mom (they are very large) (/s)

  • pech@lemmy.world
    link
    fedilink
    English
    arrow-up
    21
    arrow-down
    1
    ·
    5 days ago

    eBay has some rad Chinese mezzanine boards for these guys too. Nvlink works and everything lol 3566 file-QvfRnBhkoKQBxmqtmrGBLV

    • Septimaeus@infosec.pub
      link
      fedilink
      English
      arrow-up
      6
      ·
      4 days ago

      Cool mezzanine boards, but can we talk about your dope af custom jig for offset mounting arbitrary boards?

      • pech@lemmy.world
        link
        fedilink
        English
        arrow-up
        4
        ·
        4 days ago

        I appreciate the kind words! Until recently, my day job was CAD monkey. I wanted to consolidate the hardware I was cobbling together, found a cheap 8 GPU mining rig and some extra 2020 extrusions to play with.

        The seller for the mezzanine said it followed the mATX mounting hole pattern, (it does but is ~75% in the width and doesn’t use all the points). I have a tendency to overthink designs and kind of stalled for a bit before finally taking the plunge and whipping up these struts. I used some brass heat set inserts to accept the standoffs and everything pretty much went right together lol

        3523

      • pech@lemmy.world
        link
        fedilink
        English
        arrow-up
        11
        arrow-down
        1
        ·
        4 days ago

        I am using PTM sheets and they idle at decent temps, though I did ziptie some high CFM fans behind them lol 3581

    • jj4211@lemmy.world
      link
      fedilink
      English
      arrow-up
      8
      ·
      4 days ago

      They mostly don’t, but this is also not how the datacenters cool them.

      A datacenter will either have an open water loop, or an all in one taking heat to a more advantagous place for a radiator to be, or at the very least better managed airflow with bigger fans and more specific air baffles.

      This thing has no such luxury and has a small area and unknown broader thermal context, so screaming it is to make up for the limitations of the scenario.

    • MorningWood@anarchist.nexus
      link
      fedilink
      English
      arrow-up
      8
      ·
      5 days ago

      No they really dont. Big ass fans running 24/7 to help the small fans running 24/7. It all blends into an easily ignored drone though just dont try to have a conversation in there.

  • First_Thunder@lemmy.zip
    link
    fedilink
    English
    arrow-up
    17
    ·
    5 days ago

    The main thing that itches me with the V100 is the fact that given that pascal is about to be EOL, a 2017 card is probably soon next

  • brucethemoose@lemmy.world
    link
    fedilink
    English
    arrow-up
    8
    arrow-down
    4
    ·
    edit-2
    4 days ago

    This isn’t the smart way, though.

    What the homelabbers do (at least before the RAM crisis) is buy Xeon/TR/EPYC boards on the cheap, and then run gaming GPUs for hybrid inference.

    This is what I do. I run MiMo 2.5 at 8-10t/s on a 7800X3D/RTX 3090/128GB CPU RAM, more with Dflash. That’s a 300B model: it’s not even in the same class as Qwen 27B, which is what the dev in OP’s article is trying to run.

    And this is small-time: setups with 4-8 memory channels can run stuff like Kimi or Deepseek Pro, even faster. Or they can run smaller LLMs with quantization types that are very fast on CPUs, and get crazy speeds.

    …And besides, Qwen 27B can run fine on a 4080, with the right framework. It will fit in 16GB as an exl3.


    Not that this isn’t a cool hardware hacking project.

    …But it’s kind of the wrong approach. It’s about 2 years out of date, as MoEs are king in LLM land now. RAM is horrendously expensive, yes, but so are most used V100s, or used 3090s.

    • 4am@lemmy.zip
      link
      fedilink
      English
      arrow-up
      19
      ·
      5 days ago

      Lots of used DC gear makes it to eBay. This is how homelabbers survive.

      • eleitl@lemmy.zip
        link
        fedilink
        English
        arrow-up
        2
        ·
        5 days ago

        I know that and I use that. But recently market for switches and servers dried up. I thought they were shredding DC GPUs too. Most of them die after 3-5 years anyway.

    • jj4211@lemmy.world
      link
      fedilink
      English
      arrow-up
      8
      ·
      4 days ago

      V100s are ancient in a market obsessed with the very very latest.

      The datacenters are unlikely to bother directly with eBay, but they have asset recovery companies that will take the stuff off their hands and seek buyers, including over eBay.

    • ranzispa@mander.xyz
      link
      fedilink
      English
      arrow-up
      35
      arrow-down
      2
      ·
      5 days ago

      The guy is explaining how to hack around GPU fans which were never meant to be run at lower speed through jumper cables and you complain about slop?

      • tyler@programming.dev
        link
        fedilink
        English
        arrow-up
        6
        arrow-down
        3
        ·
        4 days ago

        This “guy” isn’t explaining anything. The entire article is written by AI. It’s the same bland writing style everywhere now.

        • ranzispa@mander.xyz
          link
          fedilink
          English
          arrow-up
          4
          arrow-down
          1
          ·
          4 days ago

          I don’t know whether you read the article. It was quite interesting to me and gave me some good ideas. I not going to replicate what he did, but he has shown me a methodology I could apply to do other things.

    • Deebster@infosec.pub
      link
      fedilink
      English
      arrow-up
      4
      ·
      4 days ago

      I thought parts felt like AI and parts felt like a British human writing it.

      AI(?):

      But here is the thing: this is a Volta GPU with 16GB of HBM2 memory, 5120 CUDA cores, and I picked it up for about £150 on eBay. The compute is still real. The VRAM is still real. And the memory bandwidth is where it gets genuinely surprising.

      Human (?):

      So I shoved some jumper wires into the connector and jammed the other ends into a spare fan header (turn your volume up)

      I suspect he’s used some AI polishing tool to turn his draft into a publishable post.

    • eleijeep@piefed.social
      link
      fedilink
      English
      arrow-up
      8
      arrow-down
      2
      ·
      4 days ago

      It’s so clearly written with the aid of an LLM that I’m embarrassed for all the people downvoting this comment. Just because the basic outline of the information presented and the sequence of events comes from the real experience of a real person, it does not mean that the real person wrote the article themselves. Seriously, how many real people write like a bad 80s detective movie script? It’s literally a parody trope at this point.

      • IPeaceInYourFace@lemmy.world
        link
        fedilink
        English
        arrow-up
        4
        arrow-down
        4
        ·
        4 days ago

        Honestly, I just don’t care.

        I am a monkey, I live on a rock, in the middle of an infinite empty space, that none of the smartest monkeys on this rock can figure out.

        I am part of a species that has spent its entire history killing itself in the most horrifying and ghastly ways for the most benign reasons that the greatest war approaches us and everybody is so accepting.

        The planet I live on is slowly becoming inhabitable to point where the thought of kids is pointless. And banks have used to many money glitches that currency across the globe is going into hyperinflation.

        I really really cannot illustrate how much I do not give a fuck about the cadence of written words in that small space of relief I get from the inevitable destruction looming in a completely absurd and pointless existence.

      • tyler@programming.dev
        link
        fedilink
        English
        arrow-up
        1
        arrow-down
        2
        ·
        4 days ago

        It has nothing to do with the subject of the article. Are you blind? It’s the exact same bland, AI regurgitated writing style as every other AI written article.

        • Septimaeus@infosec.pub
          link
          fedilink
          English
          arrow-up
          1
          arrow-down
          1
          ·
          4 days ago

          For well-known authors with a large corpus of prior writing, it’s simple to run a style match analysis, if you wish to demonstrate that they lazily used generative AI to write a particular piece.

          But for all the knee-jerk “ai slop” accusations I see on almost every post now, the only justification most offer is that it’s obvious, which sounds like “I can tell by the pixels.”

          • tyler@programming.dev
            link
            fedilink
            English
            arrow-up
            1
            ·
            4 days ago

            I detailed several of the indications in another comment. If you’re too blind to see the very obvious indicators then oh boy. That’s what makes it AI slop. They don’t even bother trying to hide it. It’s like bad CGI, there’s no need to say “I can tell by the pixels” because it’s transparently obvious.

      • tyler@programming.dev
        link
        fedilink
        English
        arrow-up
        3
        arrow-down
        2
        ·
        4 days ago

        It’s literally AI slop. Ignoring the subject of the article (which is whatever), the article is clearly entirely written by AI. The AI writing style is horrendous. “It’s not blah. It’s blah.”, “it’s genuinely <insert adjective of amazement>”

        • Fizz@lemmy.nz
          link
          fedilink
          English
          arrow-up
          3
          arrow-down
          1
          ·
          4 days ago

          Its not AI writing. It reads exactly like someone casually talking. Its consice and to the point. There is no “its not x, its y”

          He has been blogging dor dam near 6 years you can go read his older posts the writing is consistent.

          • tyler@programming.dev
            link
            fedilink
            English
            arrow-up
            2
            ·
            3 days ago

            But here is the thing: this is a Volta GPU with 16GB of HBM2 memory, 5120 CUDA cores, and I picked it up for about £150 on eBay. The compute is still real. The VRAM is still real. And the memory bandwidth is where it gets genuinely surprising.

            It’s AI slop. You’re blind.