A computing enthusiast has repurposed a very noisy and largely obsolete enterprise GPU (with lots of VRAM) for local LLM inference purposes. They are now enjoying a system that has doubled its total VRAM quota to 32GB for just a $266 (£200) outlay. That’s a good result, especially in the midst of a RAMpocalypse.Oscar Molnar explains that a cheap Tesla V100 SXM2 with 16GB HBM2 was sourced, as was an SXM2-to-PCIe adapter, and a PWM mod for the loud-as-a-lawnmower cooler, to complete this VRAM expansion for the hefty local LLMs project. Indeed, these GPUs do look cheap right now, as I can see them listed on eBay US for under $140 each, if you don’t mind buying from China.As mentioned above, you can’t just get one of these Tesla V100 SXM2 cards with abundant VRAM and plug it into your PC. Molnar says they spent about $66 on an SXM2-to-PCIe adapter, also on eBay.You might think that was enough. However, the PC and local LLMs enthusiast baulked at the noise of “the fan from hell,” which came as standard with the Tesla V100 SXM2. That shrieking cooler was measured outputting 82dB of noise. Molnar described it as “somewhere between a garbage disposal and a lawnmower.” This may be the most complicated tweak yet, but basically the existing fan wires just needed rerouting and plugging into the motherboard PWM fan header. You could also simply purchase a “2.54mm male to PH2.0 female jumper cable” for the task. Apparently, the fan only needs to run at 10% to keep the Tesla V100 under 50C at full load.
AI enthusiast adds Nvidia Tesla V100 as loud as a lawnmower to gaming PC for $266 — 32GB of VRAM rig can run 27 billion parameter model at 32 tokens per second
Now they have a total 32GB VRAM system on a budget, for local LLM inference.
Enthusiast achieved 32GB VRAM for $266 (Tesla V100 16GB + RTX 4080), running 27-billion-parameter LLMs at 32 tokens/second locally. Used enterprise GPUs undercut cloud inference costs, reshaping capex calculus for LLM stack vs API strategy.









