Amazon Web Services, the world’s largest cloud provider, has issued internal directives requiring its engineering teams to cut CPU waste and clean up underutilized EC2 instances. The reason is straightforward: demand for compute, largely driven by AI workloads, is outpacing supply, and engineers requesting new CPU capacity are now waiting days to get it.

What’s actually happening inside AWS

The directives trace back to internal meetings held in May, where leadership stressed the need to maximize existing infrastructure before expecting new hardware to arrive. Engineers have been given deadlines to identify and reduce idle EC2 instances, a category that reportedly accounts for roughly 65% of all running instances being underutilized.

The optimization playbook AWS is pushing internally includes better rightsizing (matching instance types to actual workload needs), automated shutdowns for inactive instances, and broader efficiency improvements across how teams provision and manage resources.

The multi-day delays for new CPU server capacity are the most telling symptom. AWS has historically operated with enough headroom that internal teams could spin up resources relatively quickly. When provisioning timelines stretch from hours to days, the infrastructure is running closer to the redline than anyone would prefer.