Quantization can speed up LLM inference—but results vary by hardware, format, and batch size. Learn what works, what it costs, and how to stack it with caching.