I spent a night trying to fit a dense 30B model, 256K context, vision,
and speculative decoding onto one 24 GB GPU. The fastest quant lost. The
quant with the lowest perplexity lost too. What won was the
configuration that made the whole system useful, not any single number
impressive.








