I spent a night trying to fit a dense 30B model, 256K context, vision,

and speculative decoding onto one 24 GB GPU. The fastest quant lost. The

quant with the lowest perplexity lost too. What won was the

configuration that made the whole system useful, not any single number

impressive.