The most useful thing in jamesob/local-llm is not the GPU shopping list. It is the fifteen or so BIOS settings, kernel flags, and PCIe hacks that stand between owning four RTX PRO 6000 cards and actually getting them to run a 594-billion-parameter model at a claimed 80 tokens per second. The author, James O'Beirne, frames the repo plainly: "Everything I know about running LLMs locally." What that turns out to mean is a field report on the parts of the job that no product page warns you about.
Two budgets, two very different machines
The guide splits along price. For roughly $2k, it recommends two used RTX 3090s for 48GB of combined VRAM, enough to run Qwen3.6-27B plus local speech-to-text with whisper-large-v3. The author is candid that this tier gets you "pretty far," and singles out local STT as something he actually reaches for daily, in part because he feels comfortable using it in a way he does not with a hosted equivalent. The STT runner only assumes about 11GB of VRAM, so it is the low-friction entry point in the whole repo.
The $40k tier is where the writing gets interesting. Four RTX PRO 6000 Blackwell cards give 384GB of VRAM, which the author says lands you "something pretty close to Claude Opus" via a quantized GLM-5.2 variant. Rather than pair those cards with an expensive PCIe5 and DDR5 platform, he built a last-gen EPYC Milan system almost entirely from eBay parts: an ASRock Rack ROMED8-2T board, a 16-core 7313P, and 128GB of DDR4 ECC, for a base system total of $5,587. The stated logic is to spend money on VRAM, where it counts, and cut it everywhere else. RAM prices as of July 2026, he notes, made the DDR4 route the sane one.







