I’ve sat through a lot of data center pipeline conversations this year, and at some point most of them arrive at the same complaint: there isn’t enough power. Then a comment from Joe Dominguez, Constellation’s CEO, kept surfacing in my notes. “We have a peak capacity concern, not an energy concern,” he said earlier this year. That one line reframed a problem I’d been thinking about the wrong way.
The US isn’t short on electricity in any aggregate sense. What’s scarce is capacity during the handful of hours each year when demand spikes all at once, a July heat wave, a January cold snap. Utilities have to build for those peak hours, then watch the same infrastructure sit half-idle the rest of the year. A data center that insists on guaranteed full power at any moment, including the worst of those peaks, is effectively asking a utility to size its whole system around that one facility’s worst case. Multiply that across every new gigawatt-scale project trying to get connected, and you get the multi-year interconnection queues everyone in this space complains about.
What took me longer to appreciate is why nobody fixed this sooner. It isn’t a technology gap. It’s an incentive gap.
Regulated utilities earn their return on the capital they build, not on the capacity they help someone else use more efficiently. A new substation or transmission line goes into the rate base and earns a guaranteed return for decades. A software layer that lets existing infrastructure serve more customers does not. So for years, the rational move for a utility has been to keep sizing for peak and let developers wait in line, even when the electrons technically exist. That’s a structural reason flexibility never got built, not a technical one.
That gap is what a new category of infrastructure software is now walking into: systems that sit between a data center and the grid and decide, in real time, how much power to draw and when. If a cluster can shed load for twenty minutes during a demand spike and make it up later, the utility no longer has to plan around that cluster’s theoretical peak. Do that across enough facilities and you’re not adding generation, you’re unlocking capacity that already existed but was never allocated to anyone.
None of this is conceptually new. Industrial demand response has existed for decades, aluminum smelters and steel mills getting paid to cut load during grid emergencies. What’s different with compute is the response time. A smelter needs hours of notice. A GPU cluster can modulate in seconds to minutes without touching the workloads running on it, which makes it a far more useful, and more valuable, flexibility resource than anything the demand response industry has worked with before.
Emerald AI is one company building this layer. It raised $150 million earlier this year at a reported $1 billion valuation. That round came on the back of pilot results in Phoenix and London, where GPU clusters reportedly shed load during grid stress events without materially slowing training or inference. I’m less interested in the specific number than in what it signals: institutional capital is now willing to price grid orchestration for data centers as its own category, distinct from power generation and distinct from data center construction.
The open question, and the one I find myself going back and forth on, is who ends up owning this layer long term. Hyperscalers already run some of the most sophisticated internal power management in the world, Google has published research on shifting compute load to follow carbon-free power availability for years now. A hyperscaler with that internal capability may simply build its own version rather than buy it. The more interesting buyers are probably neoclouds and colocation operators scaling fast without that internal expertise, which is a smaller market than the headline funding numbers suggest, but a real one.
None of this works unless utilities are willing to credit curtailable load in their planning, which means capacity markets like PJM and ERCOT actually changing how they treat flexible demand, and that moves slowly and unevenly by region. It also only works if the economics hold: the value of the freed-up capacity (faster interconnection, avoided curtailment fees) has to exceed what it costs to build and operate the orchestration layer itself. Pilot results are not the same thing as economics that hold up at scale.
Dominguez’s framing stuck with me because it points at where the real leverage is. The industry has spent two years treating this as a generation problem, more gigawatts, faster permits, more turbines. Timing is a cheaper problem to solve than supply, which is exactly why I think the coordination layer between compute and the grid ends up mattering more than it gets credit for right now.







