A field report on serving Google's Gemma 4 E2B on AWS EC2 **G5g* — a Graviton2 (aarch64)

host with an NVIDIA T4G (Turing, SM 7.5) GPU. Three obstacles: an arch list nobody

publishes for this combination, a version floor that only the newest vLLM clears, and

64 KiB of shared memory that stops the model dead. Plus the seven things I documented

wrong before I had a box.*