• Sources: GitHub blog post, GitHub status incident analysis, HN discussion
  • Summary: The company blog post states that neither August incident came from a code or configuration change and that both were capacity failures, and it puts numbers on the migration underneath. Azure serves roughly 58 percent of platform load and half of all Git operations, up from 12 percent of platform load in May, and monthly commits grew from 1.4 billion to 2.9 billion since April. GitHub's own root cause analysis on its status page names a narrower immediate cause, network saturation on Central US load balancers caused by an Istio sidecar pod reaching its concurrency limits and failing to auto scale correctly because of a misconfigured policy that watched host service limits but not sidecar limits. That analysis carries operational figures the blog post omits: peak error rates near 20 percent for web and API and near 50 percent for archive and raw content, four HAProxy nodes exhausting their flow limits, a latent retry bug in VS Code amplifying traffic roughly 10x, Copilot Token Service traffic rising from a normal 7,000 to 9,000 requests per second to 70,000 to 100,000, and scraping against codeload endpoints impeding recovery. A misconfigured autoscaling policy sits uneasily against the blog post's statement that no configuration change was involved. The named remediations are consistent retry limits, retry budgets and variable timeouts across service-to-service calls to stop retry storms, and a review of lower-priority CPU and memory alerts for components that fail under sudden spikes.
  • Why it matters: The remediations are ones any service owner can apply without GitHub's scale, and the disclosed migration numbers change how much of the platform a single cloud now carries.
  • Follow-up: The promised root cause analysis is published, so the open question narrows to whether the retry budget and variable timeout work is reported as landed.

send feedback on this story