This article compares two AWS GPU instance families for serving a small language

model, using a payload that is byte-identical on both. The older family loses

87% of decode to dtype conversion, and nothing in any log, metric or health

check says so.

The code is here: