A field report on running Google's Gemma-4 family on AWS Inferentia2 (inf2), covering the
three architectural obstacles that break the vendor stack — **mixed attention heads, the
**vLLM / optimum-neuron / NxD* dead-ends, and the Neuron compiler (neuronx-cc) limits —
and the recipe that got all three model sizes serving coherently.*
Models






