GitHub is replacing its long‑standing Git storage stack with a design that stores repository objects in Azure Blob Storage and adds lightweight read workers in front of it. The change is driven by the massive increase in commit and push volume caused by AI coding agents that commit after every small action, a pattern that overwhelms the previous model where a push was a human‑scale event.
Why the shift matters now
In September 2026 GitHub logged 7.38 billion commits, a five‑fold jump from a year earlier, and push events rose from 0.69 billion to 3.35 billion per month. Each push triggers thousands of reads as CI pipelines clone the repository and code‑scanning tools run. The legacy storage system, called Spokes, kept five full replicas of every repo on local file servers and used a three‑phase commit protocol that forced every replica to participate in each write. That architecture made writes as slow as the slowest replica, so adding replicas to handle read load directly degraded write performance. With agents generating writes at machine speed, the write path became the primary bottleneck.
Key architectural changes
- Separation of storage and compute. Authoritative objects now live in Azure Blob Storage, which already provides durability and geo‑replication at cloud scale.
- Read workers as a caching layer. Lightweight compute nodes sit in front of Blob Storage, caching frequently requested objects. Adding or removing these workers changes read capacity without affecting write durability.
- Reduced coordination. Only reference updates (branch pointer moves) require consensus across the system. Object storage, connectivity checks, and secret scanning run independently.
- Background jobs off the serving path. Maintenance tasks such as compaction and garbage collection are handled by dedicated workers, preventing them from competing with client traffic.
Internal benchmarks reported up to a 35× increase in write throughput after the redesign. The rollout timeline and any required customer migration steps have not been disclosed.
Operational and security considerations
Practitioners should note several downstream effects:
- Write‑heavy workloads. With agents committing after each micro‑change, write latency directly impacts agent throughput. Monitoring push latency and scaling read workers will be essential to keep the pipeline moving.
- Branch‑protection bottleneck. Even if writes accelerate, required reviews and branch policies still operate at human speed. Teams may see a backlog of verification work unless they adjust review policies or require agents to attach evidence of automated checks before merging.
- Clone storms. Each push can trigger a full repository clone for CI jobs. Organizations should evaluate whether they can shift to incremental fetches, cache clones, or batch builds to reduce network and storage load.
- Cache miss handling. A compute worker failure now results in a cache miss rather than a durability event, but it still adds latency. Designing retry logic and monitoring cache hit ratios will help maintain performance.
- Security tooling placement. Secret scanning and code‑scanning now run independently of the write path, which may change the timing of detection. Ensure that scanning results are still enforced before a protected branch is updated.
Related CloudNinjas coverage: DevOps.
What This Means For Practitioners
Teams that rely on AI coding agents should audit their current Git‑centric pipelines for three pressure points: push latency, clone volume, and branch‑protection queues. Consider the following actions:
- Instrument push latency and scale read workers or cache layers before the write path becomes saturated.
- Review CI configurations to avoid full repository clones on every agent checkpoint; use shallow clones or artifact caching where possible.
- Update branch‑protection policies to require automated evidence (e.g., test results, security scans) attached to each agent‑generated commit, reducing manual review load.
- Plan for a gradual migration to the new storage model by testing against Azure Blob Storage in a staging environment and validating that background maintenance jobs no longer impact request latency.
By treating the agent‑driven commit rate as the new baseline rather than a spike, organizations can align their storage, compute, and review processes with the reality of machine‑speed development.
