Skip to main content

Microsoft’s GitHub Hit by Major Outage as AI-Driven Demand Strains Infrastructure

GitHub, the Microsoft Corp.-owned code hosting platform serving more than 180 million developers, is still reeling from a widespread outage on Monday that severely disrupted software development pipelines globally.

The hours-long incident was the latest in a relentless series of reliability failures for a service struggling to keep pace with an unprecedented surge in artificial intelligence (AI)-assisted coding traffic.

The disruptions began around 9:40 a.m. EDT, initially manifesting as performance degradation across core services. The issue rapidly escalated, causing error rates to spike near 20% for web interface and API traffic, while archive and raw repository content downloads suffered a staggering 50% failure rate.

Key capabilities — including GitHub Actions automated testing, webhooks, GitHub Pages, and the AI pair-programmer Copilot — were heavily compromised. Downdetector logged thousands of user reports at the peak of the disruption, with minor outage spikes simultaneously impacting Microsoft Teams.

By midday, GitHub engineers reported they had identified the problematic component and deployed corrective action. While incident reports began to subside, the root cause remained undisclosed.

The service failure arrives amid growing developer fatigue over GitHub’s operational stability.

The platform recorded eight separate degradation incidents in July alone, followed by an Aug. 6 disruption that GitHub itself labeled “unacceptable.”

Engineers have pointed to structural bottlenecks as the ecosystem buckles under massive traffic increases. Earlier this year, GitHub leadership acknowledged that AI-assisted workflows were placing significant stress on backend infrastructure. To cope with the shift in how software is built, the company embarked on an ambitious strategy to expand platform capacity thirtyfold.

The frequent downtime is prompting technical leaders to re-evaluate their reliance on single-provider ecosystems.

As GitHub increasingly centralizes source control, continuous integration, and AI generation, an infrastructure outage effectively halts the modern software assembly line.

For engineering organizations worldwide, the recurring disruptions are shifting redundancy from a secondary consideration to an operational imperative.



from DevOps.com https://ift.tt/vdjqY7J

Comments

Popular posts from this blog

AWS Adds Agentic Workspace to Kiro AI Coding Tool

Amazon Web Services (AWS) this week added an open source workspace for its Kiro artificial intelligence (AI) coding tool that enables application developers to asynchronously assign tasks to an AI agent that is capable of autonomously performing tasks, such as testing code as it is created, in a way that maintains context across multiple sessions. Darko Mesaros, a distinguished developer advocate at AWS, said the Kiro Crew workspace is also capable of creating reusable AI skills by observing the tasks developers assign to Kiro as they write code. Kiro Crew orchestrates agents using the Agent Client Protocol (ACP) to ensure every step is observable in real time as sub-agents are spawned. For example, developers can also hand off a ticket queue to Kiro Crew for it to triage issues and flag what needs their attention or ask it to investigate the root cause of an incident while a developer continues to work on another task. An Activity view shows each agent’s reasoning, every tool call,...

Waiting for a Crypto Boom in the Public Markets

Waiting for a Crypto Boom in the Public Markets By Andrew Ross Sorkin, Jason Karaian, Sarah Kessler, Michael J. de la Merced, Lauren Hirsch and Ephrat Livni from NYT Business Day https://ift.tt/3uPpd32 Virtual Currency, Armstrong, Brian (1983- ), Grupo Televisa SAB, Univision, Kelly, Jon (Editor), Mergers, Acquisitions and Divestitures, Madoff, Bernard L, Initial Public Offerings