In the last entry I got Gemma-4's 128-expert MoE running on an inf2.24xlarge and signed off with a
cliffhanger: fitting it on a 2-core box "needs fp4 — a separate expedition." This is that expedition.
It ended nowhere near where I thought it would: not with fp4, and not on the 8xlarge I was
aiming for, but on the smallest, cheapest Inferentia2 instance AWS sells — a single inf2.xlarge
with 16 GB of host RAM — running a 26B-parameter model. Here's the refinement trail, dead ends






