A field report on running Google's Gemma-4 family on AWS Inferentia2 (inf2), covering the

three architectural obstacles that break the vendor stack — **mixed attention heads, the

**vLLM / optimum-neuron / NxD* dead-ends, and the Neuron compiler (neuronx-cc) limits —

and the recipe that got all three model sizes serving coherently.*

Models