This article provides a step by step deployment guide for Gemma 4 E2B onto the two cheapest
whole GPU CUDA instances AWS sells, and compares what they cost to run. A suite of Python MCP
tools is built to simplify management of the vLLM hosted deployment. Everything was measured on
What is this project trying to Do?
The question is simple: if you want a CUDA GPU on AWS as cheaply as possible, which one do you







