Benchmarks of the same model on the same GPU across three serving stacks, then an FP8 pass on the winner. All numbers measured on our own hardware last week. Raw CSVs, the environment manifest and a one-command reproduction script exist for every figure; the script was re-run end to end after the report was written and all figures reproduced within 6 percent.

Setup

GPU: RTX PRO 6000 Blackwell, 96 GB, sm_120 (workstation Blackwell, not the datacenter parts)

Model: Qwen3-8B

Stacks: vLLM 0.27.1, SGLang 0.5.9, llama.cpp (CUDA build)