A 15GB model loaded fully onto the GPU and generated at 5.6 tokens per second. The theoretical ceiling was 8. Capacity and throughput are set by different resources, and the obvious fix for a tight fit — shrink the model — barely moves the one that matters. One division tells you which lever works.