A field report on serving Google's Gemma 4 E2B on AWS EC2 **G5g* — a Graviton2 (aarch64)
host with an NVIDIA T4G (Turing, SM 7.5) GPU. Three obstacles: an arch list nobody
publishes for this combination, a version floor that only the newest vLLM clears, and
64 KiB of shared memory that stops the model dead. Plus the seven things I documented
wrong before I had a box.*






