This article compares two AWS GPU instance families for serving a small language
model, using a payload that is byte-identical on both. The older family loses
87% of decode to dtype conversion, and nothing in any log, metric or health
check says so.
The code is here:







