Of the five Gemma-4 models I ported to AWS Inferentia2, the 26B-A4B was the one that was supposed to
be impossible: a Mixture of Experts. It compiled on the first try — then produced nothing. The
device output was empty. The CPU reference was perfect. My unit tests all passed. The bug was in a
place I'd have sworn couldn't have one.
This is the MoE entry in the series. The dense models (E2B/E4B/12B/31B) are hard because of






