Serving Gemma 4 E2B on a TPU v6e-1
A Cloud TPU v6e-1 (Trillium) costs 2.25× a v5e-1 and returns 1.62–1.68× the throughput on workloads that fit in a v5e, and 2.32–2.77× on workloads that do not. Per output token that makes v6e 34–39% dearer in the first regime and 3–19% cheaper in the second — so the case for the bigger chip is narrower than the memory ratio suggests, and break-even sits at roughly 270,000 KV tokens.
v6e is not a general upgrade over v5e. It is a memory upgrade sold at a compute price: 32 GB against 16, a KV pool of 1,151,744 tokens against 321,376 (3.6×), for 1.907× the bandwidth. Where the extra memory does nothing, the workload pays 2.25× for 1.6×.
Two findings drive the rest:
There is no capacity knee at any occupancy tested. TTFT = −8542 + 265 × concurrency, R² = 0.999996, across 56% to 157% of the KV pool, with num_preemptions_total = 0 in every cell. A line fitted entirely below 100% occupancy predicts 157% to within 0.13%.






