I ported Google's biggest open Gemma-4 — the 30-billion-parameter dense model — to AWS Inferentia2.

The compiled model passed my strictest check: its output was token-for-token identical to the

CPU reference. It was also complete gibberish. Both facts were true at the same time, and the reason

is a lesson worth more than the port.

This is the sequel to porting the smaller Gemma-4 models (E2B/E4B/12B). Those were hard because of