A capacity blip at GitHub resulted in an outage that lasted 7 hours and 47 minutes. The infrastructure recovered faster than the clients let it. 😬
What actually happened
GitHub experienced issues on August 17, 2026, from 13:28 to 21:15 UTC. Users were facing problems with Git operations, Actions, Issues, PRs, and Copilot. The error rate for Web and API traffic was around 20%, while for archive and raw-content downloads it was approximately 50%. The root cause of the issue was almost mundane, an Istio sidecar proxy reached its concurrency limit. Here's the nasty part. GitHub's autoscaling monitored the metrics of the host application, not the sidecar's proxy saturation. Therefore, the system detected everything as "normal" while the proxy was overloaded, and no new instances were deployed. Traffic overflowed, causing four HAProxy nodes to exceed their capacity limits, and subsequently degrading the gateway. It was a rough afternoon. However, it's fixable. What made a blip a saga was the following. ## Retries didn't help. They piled on. A hidden bug in the VS Code Copilot extension caused clients to excessively re-request auth tokens from the Copilot Token Service. Backend errors were returned, and the clients then did the "resilient" thing. They tried again. Right away. Over and over. The Token Service typically processes 7,000 to 9,000 requests per second. However, at the time of the incident, it received 70,000 to 100,000 RPS. This means 8 to 14 times the normal amount of traffic, specifically targeting infrastructure that was attempting to recover. In the blog post, GitHub CTO Vlad Fedorov provided a simple explanation: backend errors that "triggered a client-side retry loop that increased traffic during recovery."







