Back to Articles
FP8 KV-cache quantization transforms Intel® Arc™ Pro B70 into a significantly more capable long-context and high-concurrency inference platform by doubling effective KV-cache capacity while preserving accuracy for most tested models.
Capacity: FP8 KV delivers a deterministic 2.0× KV-cache token-capacity increase over BF16 across all ten evaluated models (1B–72B) and TP configurations (TP=1, 2, 4)
Concurrency: This directly translates into about 2.0× higher raw 4K session capacity at the same max_model_len. For example, Qwen2.5-14B-Instruct (TP=2) scales from 48.52 → 97.05 sessions, and Mistral-Small-24B-Instruct-2501 (TP=2) from 43.27 → 86.53 sessions, making FP8 KV a strong enabler for capacity-constrained deployments.
Throughput: The benefit is strongest in long-context workloads (16K–32K), where serving becomes KV-bandwidth bound. The peak gain reaches +42.3% for Qwen2.5-14B-Instruct (TP=1, 16K), with strong gains for Qwen3-8B and Llama-3.1-8B, and material improvements for 70B-class TP=4 models at 32K






