Skip to content
Brian Sithu
Go back

The postmortem is already written

Cover image for The postmortem is already written

Forrester predicts AI data-center upgrades will trigger two major multiday cloud outages in 2026, and power is the binding constraint, so I wrote the postmortem now while I can still be wrong in public.

Brian Sithu

Written by

Brian Sithu

View profile

On this page

02 sections

I have read enough postmortems to recognize one before the incident happens. Forrester published its 2026 cloud predictions and the first one reads like a root cause written in advance. Forrester predicts that AI data-center upgrades will trigger two major multiday cloud outages this year. Forrester also says the high-profile AWS and Azure outages of 2025 were the preview.

The mechanism Forrester names is simple. Hyperscalers are moving investment out of legacy x86 and ARM environments and into GPU-centric data centers for AI, while the aging gear keeps serving current load and falters under growing complexity. New money goes to the new buildings. Old buildings get older. Something shared gives out in the middle.

The money behind that shift is hard to overstate. Futurum Group reports that the five largest US cloud and AI infrastructure providers have committed $660 billion to $690 billion in capital expenditure for 2026, nearly doubling 2025 levels. Amazon alone projects around $200 billion. All of them report markets that are supply-constrained rather than demand-constrained, which is vendor language for having more orders than they can fill.

The constraint is power

Here is the number that matters most to me. Futurum Group reports that Microsoft disclosed an $80 billion backlog of Azure orders it cannot fulfill because of power constraints: the grid connection, the permits, the substations. The International Energy Agency projects global data-center electricity consumption doubling between 2022 and 2026. Capacity is being committed faster than power gets built, and power does not ship any faster just because software people are in a hurry.

So I am willing to write the postmortem now. A region hits its power ceiling. The new GPU zones get capped or throttled because they are the hungriest load on the feed. Then the shared stuff starts to strain. My guess, and I am labeling it as a guess, is that gateway and auth-plane services degrade before customer compute does. Your VMs keep answering while the control plane that admits, routes, and authorizes everything starts timing out around them. That ordering is reverse-engineered from how regions are actually built, not from any vendor document, and I could have the order wrong.

My guess at the order things break: the power ceiling binds first, GPU zones throttle, the shared control plane strains, and your compute is the last to feel it

There is a second half to the Forrester prediction worth taking seriously. Forrester predicts at least 15 percent of enterprises will push private AI deployments atop private clouds this year, driven by cost, data lock-in, and exactly this kind of operational risk. If the shared control plane is the thing that breaks, running your own gateway in front of someone else’s compute stops looking paranoid. I read that as customers pricing the risk before the incident instead of after it, which is the correct order for once.

What I would change now

I run my own systems as if the guess above is right, because the cost of believing it is low and the cost of being surprised is high. Auth results get cached with generous TTLs so a slow identity plane degrades me instead of stopping me. Every call through a cloud gateway carries a timeout I chose, not the default, with a retry policy that backs off instead of piling on. Anything that must work during an outage can already run in more than one region, because failing over during the incident is how you learn your failover was decorative.

None of that requires believing my predicted order of failure. It only requires believing Forrester that the incident is coming. Two major multiday outages is a specific prediction with a falsifiable shape, and I like specific predictions. They leave the author nowhere to hide.

So here is mine, attached to theirs. The next big cloud postmortem will mention power before it mentions software, and it will describe the control plane failing while compute looked healthy. If the real postmortem says that, I called it. If it says something else, this page stays up unedited so you can check.



Previous Post
Rate limits weren't built for an agent army
Next Post
I don't think I want a thousand agents