Introduction: What If We Move the AI Instead of the Electricity?
On August 13, 2026, two developments appeared in the same news cycle that, taken together, illustrated a fundamental change now taking place in the relationship between artificial intelligence and the electric power system. Neither development, read in isolation, would have seemed extraordinary. One was a story about cloud-computing architecture; the other was a story about utility regulation. But when the two are placed side by side, they describe the same infrastructure problem from opposite directions — and the space between them is where this paper lives.
The first development came from Amazon. Business Insider reported that Amazon was restructuring portions of its enormous e-commerce computing architecture through an internal initiative called Region Flex. The program seeks to reduce Amazon’s concentration in several very large AWS regions and to distribute services across a broader collection of locations at smaller scale. According to internal planning documents reviewed by the publication, power and datacenter-capacity constraints had become important drivers of the effort. Some Amazon retail workloads were being moved away from established hubs such as Dublin and Northern Virginia “to mitigate expansion risk due to AWS power constraints,” in the words of one 2025 planning document, and were being redistributed across additional AWS regions such as Frankfurt and Zaragoza. Company teams were planning more than one hundred software migrations in the grocery business alone, and Region Flex had been elevated to an S-Team goal — Amazon’s designation for objectives tracked by its most senior leadership. Amazon confirmed the existence of the initiative while cautioning that some internal timelines and details reported in the documents were not current.[1]
The second development came from the electricity system itself. In late July and early August 2026, PJM Interconnection — the largest grid operator in the United States, serving roughly 67 million people across a thirteen-state territory stretching from the Washington region toward Chicago — filed a pair of proposals with the Federal Energy Regulatory Commission that would have been unthinkable a decade ago. Confronting an estimated 6.8-gigawatt reliability shortfall following its most recent capacity auction, PJM proposed a one-time backstop capacity auction to fill the gap, and, separately, a framework under which new data centers and other very large electricity customers that do not bring their own power supplies could be curtailed — required to reduce or shed demand — when the system approaches emergency conditions, before ordinary consumers are subjected to outages.[2][3] The rules would apply to large loads coming online after June 1, 2027, and PJM would maintain a formal Large Load Registry to establish load-reduction priorities.[3] As Canary Media summarized the plan, PJM is now relying heavily on states and utilities to force data centers to secure their own power — or face the possibility of being cut off during grid emergencies.[4]
These two developments describe the same problem from opposite ends of the wire. PJM is asking: how can the grid deal with enormous computing loads when electricity becomes scarce? Amazon’s evolving architecture raises the reciprocal possibility: what if some of the computing load does not have to remain where the electricity is scarce? That distinction — deceptively simple, and yet with consequences that ripple through capital markets, utility regulation, state economic competition, and geopolitics — is the starting point for this paper.
For more than a century, industrial development has been organized around moving energy toward productive assets. Transmission lines move electricity from generators to cities, factories, and industrial sites. Pipelines transport natural gas. Railroads carry coal. Tankers move petroleum across oceans. The underlying logic is so deeply embedded in industrial civilization that we rarely articulate it: a factory is geographically fixed, so civilization develops increasingly sophisticated infrastructure for delivering energy to that fixed location. The steel mill does not move to the coalfield mid-shift; the aluminum smelter does not relocate when the hydro reservoir runs low. Energy follows production, because production cannot follow energy.
Artificial intelligence changes part of this equation, because the productive activity inside an AI factory is partially digital. A model-training job does not possess the same geographical permanence as a steel mill. A batch-inference workload does not necessarily have to execute in Virginia merely because the customer initiated it there. Fine-tuning can sometimes occur hours later. Data preprocessing can be distributed. Synthetic-data generation can be performed remotely. Some agentic processes can tolerate delay. Model evaluation may be performed in another region entirely. Even portions of training and inference can potentially be scheduled according to the availability of accelerators, networking, electricity, and cooling. The economic asset remains stubbornly physical — the GPUs, servers, fiber, substations, transformers, and cooling plants cannot magically move — but the work assigned to those assets can sometimes move.
Google has already demonstrated part of this principle at commercial scale. In March 2026, the company announced that it had integrated one full gigawatt of data-center demand-response capability into long-term energy contracts with five U.S. utilities: Indiana Michigan Power, the Tennessee Valley Authority, Entergy Arkansas, Minnesota Power, and DTE Energy.[5][6] Google says it can limit or shift portions of machine-learning workloads running in its data centers during periods when the grid is under stress, reducing overall facility power demand and helping utilities balance supply and demand. Michael Terrell, Google’s head of advanced energy, framed the milestone in language that would have sounded strange coming from a technology company even five years ago:
“Demand response enables our data centers to be valuable assets for the power grid.”
— Michael Terrell, Head of Advanced Energy, Google [5]
This builds on an earlier Google concept: carbon-aware computing. Beginning in 2020 and 2021, the company demonstrated that moveable computational tasks could be shifted not merely through time but between geographical locations according to the regional availability of carbon-free electricity — encoding YouTube videos, for example, in whichever data center happened to have the cleanest power that hour.[8][9] What began as a sustainability technique is now evolving into something much larger. It is becoming an infrastructure strategy — arguably the most consequential demand-side infrastructure strategy of the AI era.
I call that strategy Workload Wheeling.
This paper develops the concept in six sections. Section 1 establishes the historical and conceptual foundation, tracing the shift from an industrial equation in which energy follows production toward one in which flexible production can follow energy. Section 2 disaggregates AI electricity demand into a Workload Mobility Spectrum, because the argument collapses if every workload is naively assumed to be movable. Section 3 develops the larger economic framework: the proposition that a hyperscaler’s datacenter fleet, joined by fiber and orchestration software, begins to function as a virtual power-aware compute grid that can partially substitute — at the margin, for certain workloads, at certain times — for physical energy transportation. Section 4 assembles the empirical evidence from 2020 through August 2026: Google, Amazon, PJM, FERC, ERCOT, NVIDIA, and a rapidly maturing academic literature spanning Duke, MIT, Harvard, Lawrence Berkeley National Laboratory, and beyond. Section 5 examines the economics, politics, and geopolitics of a world in which compute can move. Section 6 distills what we have learned into seven pillars, and the conclusion explains why Workload Wheeling is likely to become an operating principle — not merely a technique — of the intelligence economy.
Why I Chose the Title “Workload Wheeling”
The word wheeling has a specific and honorable history in electricity markets, and the choice of the term here is deliberate rather than decorative. In the traditional power business, wheeling broadly refers to transporting electricity across transmission systems from one location to another. Electricity is generated in one place, crosses infrastructure owned or operated by different entities — often paying a “wheeling charge” for the privilege — and ultimately reaches the customer somewhere else. Wheeling was one of the conceptual foundations of open-access transmission in the United States: the recognition that the wires between generation and load could be shared, priced, and traversed. The entire architecture of restructured electricity markets, from FERC Order 888 onward, rests on the idea that energy can be moved across space to wherever demand happens to sit.
Workload Wheeling reverses the conceptual direction. Instead of always asking how we wheel another megawatt of electricity to the AI factory, the AI economy increasingly can ask whether we can wheel part of the computational workload to another AI factory where the megawatt already exists. That reversal is why the terminology fits, and why no existing term — demand response, load shifting, geographic load balancing — fully captures it. Demand response describes a load that shrinks on command. Load shifting describes a load that moves through time. Workload Wheeling describes a productive activity that traverses a network in search of energy, exactly as wheeled electricity traverses a network in search of load.
The symmetry can be expressed simply. The physical power grid transports electrons. The digital compute network transports workloads. Traditional wheeling searches for a transmission path between generation and load. Workload Wheeling searches for a computational path between demand and available power. The old industrial model held that energy follows production. The Workload Wheeling model holds that flexible production can follow energy.
Table 1. Traditional Wheeling versus Workload Wheeling
| Dimension | Traditional Electricity Wheeling | Workload Wheeling |
| What moves | Electrons (energy) | Computational jobs (demand) |
| Network used | Transmission and distribution wires | Fiber networks and compute orchestration |
| Direction of search | A path from generation to load | A path from demand to available power |
| Constraint being solved | Where the customer is located | Where electricity, chips, and cooling are available |
| Era of emergence | Open-access transmission (1990s) | The AI infrastructure boom (2020s) |
| Governing logic | Energy follows production | Flexible production follows energy |
This distinction becomes progressively more important as what I call the Five-Layer AI Economy grows. Layer 1 — Energy — determines where electricity is available, at what price, and with what reliability. Layer 2 — Chips — determines where computational capacity exists, and of what kind: Nvidia GPUs, Google TPUs, AWS Trainium, custom Meta accelerators. Layer 3 — Datacenters — geographically connects energy and accelerators; it is the layer where megawatts become teraflops. Layer 4 — Models — creates computational jobs that possess varying degrees of location and time flexibility: training runs, fine-tunes, evaluations, distillations. Layer 5 — Applications and Agents — continuously generates inference demand that increasingly must be scheduled across all of this infrastructure, under latency, privacy, and jurisdictional constraints that vary enormously from one use case to another.
Workload Wheeling is therefore not another electricity technology, and it should not be filed alongside batteries, small modular reactors, or transmission reconductoring. It is a coordination mechanism connecting all five layers — a control loop through which conditions in Layer 1 propagate upward into scheduling decisions in Layers 4 and 5, and through which the demands of Layers 4 and 5 propagate downward into the operational posture of Layers 1 through 3.
It also differs importantly from what might be called Compute Curtailment. Compute Curtailment asks what happens when an electricity system tells a datacenter: consume less power. It is the framing embedded in PJM’s emergency proposals, in the U.S. Department of Energy’s May 2026 emergency order allowing PJM to curtail data centers with backup generation as a last resort before rolling blackouts,[38] and in most utility demand-response tariffs. Workload Wheeling asks whether the hyperscaler can respond: we will consume less power here because we can execute some of that work somewhere else. Curtailment destroys, delays, or reduces electrical demand at one facility. Workload Wheeling attempts to preserve computational output while relocating the electrical demand that produces it. The distinction matters economically: curtailment is a cost to be minimized, while wheeling is an optimization to be exploited. One is a concession the computing industry makes to the grid; the other is a capability the computing industry builds for itself, which happens to benefit the grid. That is a much larger economic idea, and it is the idea this paper is about.

Section 1 — From Electricity Wheeling to Intelligence Wheeling: The Industrial Geography of AI Is Becoming Programmable
Every industrial era has an implicit theory of geography. The theory of the coal era was proximity: factories clustered near mines and ports because moving coal was expensive. The theory of the electrical era was transmission: once electricity could be wheeled across hundreds of miles, industry could locate near labor and markets instead of near fuel, and the twentieth-century industrial map was redrawn accordingly. The theory of the early cloud era was latency and land: data centers clustered where fiber routes intersected, where land and power were cheap, and where tax treatment was favorable — which is how a stretch of farmland in Loudoun County, Virginia became the densest concentration of computing on Earth. The question this section poses is what the theory of geography will be for the AI era, when the scarce input is no longer land or fiber but deliverable megawatts, and when the productive activity itself has become partially weightless.
The industrial economy traditionally assumes that large loads are geographically rigid. An aluminum smelter, chemical plant, refinery, or automotive factory cannot continuously teleport production among several states because electricity prices change. Interruptible tariffs for heavy industry have existed for decades, but they are blunt instruments: the smelter can pause, at real cost to its potlines, yet it cannot relocate its afternoon’s production to Oklahoma. AI factories are different in one crucial respect. A hyperscaler may operate facilities in Virginia, Ohio, Indiana, Texas, Iowa, Oregon, Arizona, Georgia, Tennessee, Finland, Spain, and dozens of additional markets. Individually, those datacenters are immovable — as fixed as any smelter. Collectively, however, the fleet creates a distributed computational system whose work assignments are software-defined.
The important economic unit may therefore gradually change from one datacenter to the hyperscaler’s entire datacenter fleet. This is critical, and it explains a persistent disconnect in current policy debates. When regulators examine a one-gigawatt datacenter, they naturally see a giant localized electricity customer — the largest single load their territory has ever contemplated. When a hyperscaler examines the same facility, it increasingly sees one node inside a much larger computational network, a network whose aggregate behavior it controls through schedulers, queues, and placement policies. Those perspectives are not equivalent, and much of the regulatory friction of 2025 and 2026 — from PJM’s capacity crisis to FERC’s show-cause orders — can be read as the collision between them.
1.1 The Old Infrastructure Equation
Traditionally, the chain of dependency runs in one direction: Generator → Transmission → Substation → Datacenter → Compute. Every additional unit of compute ultimately depends on delivering sufficient electricity to a particular physical destination, and the response to scarcity has therefore been physical expansion: construct new generation; install transformers; expand transmission; construct substations; build gas pipelines; add batteries; restart nuclear reactors; commission small modular reactors; negotiate utility tariffs; and wait — often for years — through interconnection studies. These remain essential, and nothing in this paper argues otherwise. The 2026 draft of the U.S. Department of Energy’s National Transmission Needs Study explicitly identifies load growth from data centers, expanding domestic manufacturing, and other large industrial loads as the principal contributors to rapidly rising electricity requirements and transmission congestion, and it documents that even after the country energized roughly 85,000 circuit-miles of new, upgraded, and rebuilt transmission between 2016 and 2024, congestion still added an estimated eleven billion dollars to wholesale electricity costs in 2023 alone.[19][42] Catherine Jereza, the Department’s Assistant Secretary for the Office of Electricity, put the moment plainly when the draft study was released:
“Electricity demand is accelerating faster than anything we’ve seen in decades.”
— Catherine Jereza, Assistant Secretary, Office of Electricity, U.S. Department of Energy [20]
Workload Wheeling does not eliminate this infrastructure requirement. America unquestionably needs more generation and more wires. What Workload Wheeling changes is the optimization problem: it asks which megawatts absolutely must be delivered to a particular substation at a particular hour, and which computational jobs could instead be executed somewhere else, or somewhen else, using capacity that already exists.
1.2 The New Equation
The emerging architecture looks more like a decision pipeline than a delivery chain: Grid Condition → Compute Scheduler → Workload Classification → Datacenter Selection → GPU Execution. Under this architecture, the system can ask whether a given computational job should execute in Virginia at 2 p.m., in Texas at 6 p.m., in Iowa overnight, or in another region several hours later. Suddenly, time and geography become variables in the objective function rather than constants in the constraint set. The megawatt is no longer an address; it is an argument passed to a function.
Table 2. Two Infrastructure Equations
| Old Equation | New Equation | |
| Chain | Generator → Transmission → Substation → Datacenter → Compute | Grid Condition → Scheduler → Workload Classification → Datacenter Selection → GPU Execution |
| Fixed variable | Location of demand | Deadline and constraints of the job |
| Free variable | Supply delivered to that location | Time and place of execution |
| Response to scarcity | Build more supply and wires | Re-route and re-time flexible work; build supply for the firm remainder |
| Planning horizon | Years to decades (steel in the ground) | Milliseconds to months (software policy) |
1.3 Electricity Is Difficult to Store; Computation Can Often Wait
The deepest asymmetry between the two systems is temporal. Electric grids must balance supply and demand continuously, second by second; storage helps at the margin but remains expensive at scale, which is why the grid’s entire institutional apparatus — capacity markets, reserve margins, emergency procedures — exists to guarantee balance at every instant. Computational demand, by contrast, has multiple temporal textures. Some jobs need an answer within milliseconds: the chatbot response, the fraud check, the vehicle’s perception loop. Others need an answer within seconds. Others can wait minutes. Others can wait overnight. Others only need to finish before tomorrow morning, or before the end of the quarter. The latter categories create something electricity systems rarely receive from conventional industrial loads: schedulable industrial demand — demand that can be told, by software, when and where to occur. In the language of power systems, a portion of the AI load is not merely interruptible; it is dispatchable in reverse. The grid has spent a century dispatching supply to follow demand. Workload Wheeling introduces a demand that can be dispatched to follow supply.
1.4 Three Forms of Workload Wheeling
It is useful to define three initial categories, because they involve different technical requirements, different costs, and different regulatory implications. Temporal Wheeling moves a workload from one hour to another at the same datacenter — for example, deferring a batch job from 2 p.m., when the local grid is straining under air-conditioning load, to 2 a.m., when it is not. This is the most mature form: Google’s original carbon-intelligent computing platform operated almost entirely in this mode, imposing hourly “virtual capacity curves” that delayed temporally flexible work to cleaner or cheaper hours while preserving total daily capacity.[9] Geographic Wheeling moves the workload between datacenters — Northern Virginia to Texas, Dublin to Zaragoza — and requires that data, model states, and orchestration all be portable across sites. Temporal-Geographic Wheeling optimizes both simultaneously: do not execute in Virginia at 4 p.m.; execute in Iowa six hours later, because electricity, cooling capacity, and accelerator availability are all better there and then. Recent grid-planning research models exactly this combination, representing data-center load as a mixture of temporally deferrable demand and geographically shiftable demand inside formal capacity-expansion frameworks, and finds that the two forms of flexibility are complements rather than substitutes.[37]
Table 3. Three Forms of Workload Wheeling
| Form | What Moves | Example | Primary Requirement |
| Temporal Wheeling | The hour of execution | 2 p.m. batch job → 2 a.m. at the same site | Deadline slack; queue management |
| Geographic Wheeling | The place of execution | Northern Virginia → Texas; Dublin → Zaragoza | Data/model portability; network bandwidth; jurisdictional clearance |
| Temporal-Geographic Wheeling | Both hour and place | Virginia 4 p.m. → Iowa 10 p.m. | Fleet-wide scheduler with power, chip, and network awareness |
Laboratory and applied research now provides a scientific foundation for treating these categories seriously. Lawrence Berkeley National Laboratory — which produced the definitive 2024 United States Data Center Energy Usage Report and hosted the Department of Energy’s Data Center Load Flexibility Workshop — is leading demonstrations under the REFLEX program with industry partners to enable data-center load flexibility through optimized controls, workload management, and storage integration, explicitly identifying the temporal and geographic movement of computational work as principal mechanisms through which data centers can become flexible loads.[39] EPRI’s DCFlex initiative, launched in 2024 and since taken global, is building the measurement and valuation frameworks that would let utilities treat that flexibility as a countable grid resource.[8][6] Workload Wheeling takes this technical capability and asks what happens when it stops being a demonstration and becomes an economic operating system for the AI industry.

Section 2 — Not Every AI Workload Is Equal: Creating the Workload Mobility Spectrum
Every seductive infrastructure idea eventually meets an engineer, and the engineer’s first question about Workload Wheeling is the right one: which workloads, exactly, can move? The concept collapses into hand-waving if the paper assumes that every AI job can simply be executed somewhere else. It cannot. Some workloads are highly mobile; others are practically immovable; most live somewhere in between, and their position on the spectrum is determined not by physics alone but by architecture, contract, and law. This section therefore builds what I call the Workload Mobility Spectrum — a disaggregation of AI electricity demand into categories with fundamentally different geographic and temporal properties. The disaggregation matters because regulators, utilities, and grid operators currently tend to treat “the data center” as one homogeneous block of firm load, and hyperscalers tend to market their flexibility in aggregate gigawatts without specifying what, precisely, is flexible. Both simplifications obscure the real economics.
The distinction has real empirical backing. MIT’s Center for Energy and Environmental Policy Research, in work by Christopher Knittel, Juan Ramon Senga, and Shen Wang that appeared as a 2025 working paper and was published in iScience in June 2026, modeled power systems in the Mid-Atlantic, Texas, and the western United States under varying levels of data-center flexibility. The team found that flexible data centers reduce system costs by shifting load from peak to off-peak hours — savings of up to five percent in Texas, four percent in the Mid-Atlantic, and two percent in the West — but that achieving those savings requires data centers to move more than a fifth of their consumption away from peak periods, and that the emissions consequences depend heavily on the local generation mix.[23][24] Knittel framed the design problem in a single sentence: the goal for data centers is to
“add to average usage but not the peak usage.”
— Christopher Knittel, George P. Shultz Professor, MIT Sloan School of Management [23]
Notice what this framing implies: the value of a data center to the power system depends not on how much energy it consumes but on when and where it consumes it — and when and where are precisely the variables that workload mobility controls. Knittel’s team also observed that training-oriented facilities, with their steadier and more schedulable consumption, offer more flexibility than latency-bound inference facilities — which is the Mobility Spectrum expressed in the language of power-system economics.[24]
2.1 Highly Mobile Workloads
At the mobile end of the spectrum sit workloads whose value does not depend on where or exactly when they run, so long as they complete within a generous deadline. These include synthetic-data generation; offline model evaluation; embeddings generation; data preprocessing and cleaning; some fine-tuning; checkpoint analysis; batch inference over large corpora; non-urgent internal analytics; software compilation and testing; some reinforcement-learning simulations; background agent tasks; video rendering, encoding, and transformation; and deferred enterprise AI workloads run under service agreements that promise completion by morning rather than by millisecond. Google’s original carbon-intelligent platform was applied first to exactly this class — the encoding and processing of millions of media files for YouTube, Photos, and Drive — precisely because such jobs “can technically run in many places,” subject to privacy law.[8] These jobs may tolerate meaningful geographic movement, temporal deferral of hours, or both, and they are the natural first cargo of Workload Wheeling.
2.2 Moderately Flexible Workloads
The middle of the spectrum contains workloads whose mobility depends on engineering choices rather than intrinsic properties. Examples include portions of frontier-model training, where checkpointing frequency, parallelism strategy, and cluster topology determine whether a run can be paused, migrated, or throttled without unacceptable loss; distributed inference, where routing layers can shift traffic between regions within latency budgets; enterprise copilots with tolerance measured in seconds rather than milliseconds; non-critical agent workflows; recommendation-model refreshes; search indexing; model distillation; and large-scale data transformation. A frontier training run is instructive: the run as a whole is anchored to a specific cluster for weeks, yet its power draw can be shaped — slowed, checkpointed, briefly paused — during grid events, which is exactly the capability Google contracts to utilities and that Emerald AI and NVIDIA demonstrated in live AI-factory settings, cutting facility power demand by roughly forty percent in under a minute while high-priority jobs continued at full throughput.[5][28][29] Mobility here is not free; it is designed, and the design costs money. But it can be designed, which is the essential point.
2.3 Low-Flexibility Workloads
At the immovable end sit applications that require tight latency, geographic proximity, guaranteed availability, or all three: real-time voice assistants; autonomous-vehicle decision systems; high-frequency financial systems; emergency-response systems; interactive consumer inference at the moment of the user’s request; industrial robotics; latency-sensitive gaming; defense systems; and regulated medical applications. These loads are substantially less wheelable, and honest analysis must say so. They are the firm core of AI electricity demand — the portion for which the old infrastructure equation still fully applies, and for which new generation, transmission, and local reliability investments remain irreplaceable. The policy conclusion follows directly: AI electricity demand should not be treated as one homogeneous block. Some megawatts are firm. Some are interruptible. Some are temporally flexible. Some are geographically flexible. Some possess both flexibilities at once. A regulatory regime, a tariff, or a capacity market that cannot distinguish among these categories will misprice all of them.
Table 4. The Workload Mobility Spectrum
| Mobility Class | Representative Workloads | Temporal Flexibility | Geographic Flexibility | Grid Meaning |
| Highly mobile | Synthetic-data generation; batch inference; offline evaluation; embeddings; preprocessing; media encoding; deferred enterprise jobs | Hours to days | High (subject to data law) | Schedulable, relocatable demand — prime wheeling cargo |
| Moderately flexible | Portions of frontier training; distributed inference; copilots; recommendation refreshes; indexing; distillation | Minutes to hours; power-shapeable | Architecture-dependent | Shapeable demand — demand response and partial wheeling |
| Low flexibility | Real-time assistants; autonomous systems; financial and emergency systems; interactive inference; robotics; defense; medical | Milliseconds to seconds | Low | Firm load — requires conventional infrastructure |
2.4 The Five-Variable Scheduling Problem
Once demand is disaggregated, the scheduling problem can be stated precisely. A future Workload Wheeling scheduler must consider at least five major variables simultaneously. First, power availability: is the local grid constrained, and what is the marginal price and carbon intensity of the next megawatt-hour at each candidate site? Second, accelerator availability: are the right chips — Nvidia GPUs, Google TPUs, AWS Trainium, custom Meta silicon — free at the candidate site, in the right cluster sizes and interconnect topologies? Third, network conditions: can the required data and model states be transferred efficiently, and does the transfer itself cost more energy and time than it saves? Fourth, workload urgency: does the result need to arrive in fifty milliseconds or twelve hours, and what penalty attaches to lateness? Fifth, regulatory geography: can the data legally leave the jurisdiction, and can the model legally run in the destination?
From these five, a longer tail of secondary variables appears: electricity price and its forward curve; transmission congestion; weather; water availability; cooling conditions; carbon intensity; utility demand-response requests and contracted obligations; chip type and firmware compatibility; storage capacity and data locality; security requirements; data residency; and model sovereignty. The future cloud scheduler therefore stops optimizing only for CPU + GPU + memory + latency + price, and begins optimizing CPU + GPU + memory + latency + price + megawatts + grid condition. That is a significant conceptual change — arguably the largest change in cloud resource management since the invention of the availability zone — because it imports an entire second physical system, the electric grid, into the scheduler’s objective function.
Table 5. The Power-Aware Scheduling Problem
| Variable | Question the Scheduler Asks | Time Scale |
| 1. Power availability | Is the grid at this site constrained? At what price and carbon intensity? | Minutes to hours |
| 2. Accelerator availability | Are the right chips free, in the right cluster shape? | Minutes to days |
| 3. Network conditions | Can data and model state move efficiently to the candidate site? | Seconds to hours |
| 4. Workload urgency | Fifty milliseconds or twelve hours? What is the lateness penalty? | Per job |
| 5. Regulatory geography | May the data and the model legally execute there? | Per jurisdiction; changes slowly |
2.5 The Power-Aware AI Scheduler
This is one of the most important ideas in the paper, so it deserves to be walked through concretely. Today’s cloud orchestration software — the descendants of Google’s Borg, Kubernetes and its commercial derivatives, the internal placement systems of AWS and Azure — decides where applications run largely according to computing resources, network topology, availability zones, latency, and economic cost. The next generation adds an energy layer. A request enters the cloud. The scheduler checks multiple datacenters and sees, for instance: Virginia — GPUs available, grid constrained, utility demand-response event in effect until 8 p.m. Texas — GPUs available, electricity price elevated on the real-time market. Iowa — GPUs available, grid normal, wind output strong. Oregon — GPUs available, renewable output strong, cooling favorable. The scheduler then selects the economically and operationally optimal location, subject to the job’s deadline and jurisdictional constraints, and it revisits the decision as conditions change.
At sufficient scale, this software is effectively performing energy arbitrage through compute placement. No electron is bought in one market and sold in another; instead, demand itself is teleported to the cheaper, cleaner, or less congested market. The hyperscaler captures the price spread between regions not as a trading profit but as an avoided cost — and the grid experiences it as congestion relief. That is Workload Wheeling, expressed in the native language of the cloud. FERC’s own analysis has begun to recognize the underlying physical capability. In remarks accompanying the Commission’s June 18, 2026 show-cause orders to all six regional grid operators, Commissioner David Rosner observed that modern large loads exhibit operational characteristics fundamentally unlike traditional demand, including
“the ability to quickly change their energy consumption, sometimes in seconds.”
— David Rosner, Commissioner, Federal Energy Regulatory Commission [13]
That single regulatory observation has large consequences. If AI loads can change in seconds, they are not merely loads. They begin to resemble dispatchable economic actors — entities the grid can plan around, contract with, and, under frameworks now being litigated, depend upon.

Section 3 — The Datacenter Fleet Becomes a Virtual Power-Aware Compute Grid: When Software Begins Performing the Job of Transmission
Sections 1 and 2 established that a meaningful fraction of AI demand is programmable in time and space. This section develops the larger economic framework that follows from that fact, and it advances the paper’s most provocative proposition. The proposition is not that datacenters can become flexible; that is by now well documented. The proposition is that computational networks can partially substitute for physical energy transportation — not permanently, not universally, but at the margin, for certain workloads, at certain times, and by enough to change how much physical infrastructure must be built and where. When software decides that a job will run in Iowa instead of Virginia, the software has accomplished something a transmission line would otherwise have to accomplish: it has resolved a mismatch between where energy is available and where demand wanted to occur. The line moves supply toward demand; the scheduler moves demand toward supply. Both relieve the same constraint.
3.1 Moving Electrons versus Moving Bits
Consider a stylized but realistic planning problem. Suppose Northern Virginia requires another 500 MW of AI capacity. The conventional answer is to build another 500 MW of deliverable generation and grid infrastructure into the most congested electrical region in the United States — a region where, as the Department of Energy’s 2026 Needs Study documents, transmission bottlenecks are already limiting the delivery of electricity to Dominion Energy’s rapidly growing data-center load.[19][49] But imagine that 100 MW of the computational requirement consists of flexible jobs — batch inference, evaluations, media processing, deferred enterprise work. One alternative is still to build 500 MW locally. Another, which becomes available only to an operator with a fleet and an orchestration layer, is to build 400 MW locally and make 100 MW of workload geographically movable across that fleet. The 100 MW does not disappear; another datacenter must execute the work, and the energy is consumed somewhere. But the most constrained electrical region no longer has to supply it at the most constrained moment. Workload Wheeling moves the geography of electricity consumption by moving the geography of computation.
The Duke University research that reframed this debate quantifies why such marginal flexibility has outsized value. In Rethinking Load Growth (February 2025), Tyler Norris, Tim Profeta, Dalia Patino-Echeverri, and Adam Cowie-Haskell analyzed the twenty-two largest U.S. balancing authorities — together serving ninety-five percent of national load — and introduced the concept of curtailment-enabled headroom: the amount of new load the existing system can absorb if that load can stand down during the small number of hours that define system peaks. Their central finding: if new large loads can curtail just 0.25 percent of their annual energy consumption during those hours, the existing U.S. power system could accommodate roughly 76 GW of new load — about a ten percent expansion of national peak demand, exceeding upper-end forecasts for data-center additions through the early 2030s. At 0.5 percent curtailment the figure rises to roughly 98 GW, with PJM alone able to absorb about 18 GW, MISO 15 GW, and ERCOT 10 GW.[10][11] Norris put the strategic stakes directly:
“these new mega loads can be added relatively quickly”
— Tyler Norris, Nicholas School of the Environment, Duke University; lead author of Rethinking Load Growth [12]
— provided, he added, that they embrace some degree of flexibility. Costa Samaras, director of Carnegie Mellon University’s Scott Institute for Energy Innovation, observed that the report shows load flexibility can enable data centers and other new loads to join the system without waiting for a full buildout of new capacity.[11] The nonlinearity is the whole story: a quarter of one percent of energy, surrendered at exactly the right hours, unlocks a tenth of the peak. Workload Wheeling is the mechanism by which an AI operator can surrender that quarter percent locally without surrendering the computation at all.
3.2 Virtual Transmission Capacity
This raises a provocative question that deserves careful, hedged treatment: could fiber networks and compute orchestration function, economically, as a limited form of virtual transmission capacity? A physical transmission line moves electrical production to demand. A Workload Wheeling network moves flexible demand toward electrical production. The direction is reversed, but part of the economic result can be similar: congestion at the original location is reduced, peak requirements shrink, and the investment case for the marginal wire or the marginal peaker weakens. The hedges matter. Fiber is not literally a substitute for power transmission; most electrical loads — homes, hospitals, factories — cannot migrate, and AI facilities themselves still require enormous physical electricity supplies for their firm core. A fleet cannot wheel its way out of a multi-day regional heat emergency if every region is hot at once. But at the margin, flexible computational demand can influence how much transmission and generation must be constructed to meet peak conditions — and infrastructure economics are made at the margin. Google argues precisely this: that demand flexibility reduces the need for new infrastructure designed solely to meet short-term peaks, a primary driver of electricity prices, and can accelerate interconnection of new facilities.[5][6] The MIT CEEPR results give the claim independent quantitative support at the system-planning level.[23][24]
3.3 From Datacenter to Computational Fleet
If the argument to this point is correct, the hyperscaler’s strategic asset is no longer any single facility. It is the combination: Datacenters + Power Contracts + GPUs + Fiber + Software Orchestration. This bundle is more powerful than the sum of its parts, and more powerful than owning any component separately. A company with thirty isolated AI factories owns thirty electricity problems — thirty interconnection queues, thirty local price exposures, thirty single points of regulatory failure. A company capable of dynamically routing workloads among thirty AI factories owns a portfolio of electricity options: the right, though not the obligation, to consume power wherever in the fleet it is cheapest, cleanest, and most available. Options have value precisely because the future is uncertain, and the electricity future of the late 2020s is very uncertain indeed.
This could become one of the deepest competitive moats separating the hyperscalers from smaller AI companies. Amazon, Microsoft, Google, and Meta possess global infrastructure footprints assembled over two decades; their combined 2026 capital expenditure guidance of roughly $700 to $760 billion — nearly double 2025’s already record level — is deepening those footprints at a pace with no industrial precedent.[33][34][46] OpenAI, Anthropic, and other model developers increasingly depend on multiple infrastructure partners, effectively renting slices of other companies’ fleets. NVIDIA increasingly supplies the common accelerator architecture connecting many of those facilities, and — as Section 4 details — is now shipping the software libraries that make grid-responsiveness a native property of its reference AI factory designs. The strategic objective, in every case, becomes not merely building more GPU clusters but making those clusters fungible enough that workloads can move among them. Fungibility, not raw capacity, is the scarce property.
3.4 Compute Follows Surplus
In today’s dominant model, compute capacity attracts electricity: the datacenter is sited, and the grid is then obliged — legally, in most retail regimes — to serve it. Under Workload Wheeling, the arrow acquires a second head: available electricity can attract compute. This means stranded, curtailed, or underutilized power resources acquire new option value. An electricity market with excess nighttime generation could attract deferred inference. A wind-heavy region experiencing surplus output — the kind that today is curtailed or sold at negative prices — could attract batch processing in exactly those hours. A nuclear plant with stable around-the-clock generation could anchor training. A region facing peak summer temperatures could temporarily export computational workloads rather than import emergency electricity. Electricity availability begins influencing software scheduling, and software scheduling begins influencing electricity economics. The feedback loop closes.
3.5 The Workload Wheeling Cycle
The full framework can be summarized as a continuous five-stage control loop that joins Layer 1 (Energy) to Layer 5 (Applications) of the Five-Layer AI Economy.
Table 6. The Workload Wheeling Cycle
| Stage | Name | What Happens |
| 1 | Sense | The system detects electricity availability, price, congestion, reliability signals, weather, and facility conditions across the fleet. |
| 2 | Classify | AI workloads are classified according to latency, geography, privacy, sovereignty, and deadline requirements — their position on the Mobility Spectrum. |
| 3 | Route | The scheduler selects the optimal datacenter, accelerator pool, and execution window for each movable job. |
| 4 | Execute | The job runs where the combined compute-energy conditions are most favorable, with power-shaping available during grid events. |
| 5 | Rebalance | The system continuously updates workload placement as grid and compute conditions change, returning to Stage 1. |
Nothing in this loop is science fiction. Stage 1 exists today in the market-data and telemetry feeds every large operator already consumes. Stage 2 exists in embryonic form wherever cloud providers distinguish spot, batch, and on-demand service classes. Stage 3 is what Google’s carbon-intelligent platform, Emerald AI’s Conductor, and NVIDIA’s DSX Flex library each implement in their own domains. Stages 4 and 5 are ordinary cluster operations. What does not yet exist — and what would mark the maturation of Workload Wheeling from technique to operating system — is the integration of all five stages, across whole fleets, under contracts that make the resulting flexibility visible, verifiable, and compensable to the power system. Section 4 assembles the evidence that this integration has already begun.

Section 4 — The Early Evidence: Google, Amazon, PJM, FERC, ERCOT, NVIDIA, and the Emerging Flexible-Compute Economy
A conceptual framework earns its keep only if the world is already behaving as the framework predicts. This section therefore turns from theory to evidence, and the evidence of 2025 and 2026 is unusually rich: a hyperscaler operating grid-contracted workload flexibility at gigawatt scale; a second hyperscaler redesigning its application architecture around power scarcity; the largest U.S. grid operator proposing curtailment rules for large loads; the federal regulator ordering every organized market to modernize its treatment of them; the Texas grid rebuilding its interconnection process around locational scarcity; the dominant chipmaker embedding grid-responsiveness in its reference designs; and an academic literature, running from 2020 through mid-2026, that has moved from proof-of-concept to national-scale quantification. Each piece looks like a separate story about cloud architecture, demand response, utility regulation, transmission planning, or AI infrastructure. Together they are one story.
4.1 Google: The Clearest Early Example
Google provides the strongest existing proof that software-defined computational flexibility can interact directly, contractually, and at scale with the electricity system. The company announced on March 19, 2026 that it had integrated a total of one gigawatt of demand-response capacity into long-term energy contracts with five U.S. utilities — Indiana Michigan Power, the Tennessee Valley Authority, Entergy Arkansas, Minnesota Power, and DTE Energy — with each contract treating demand response as a defined capacity resource that grid planners factor into reliability assessments.[5][6] The mechanism is workload manipulation: Google limits or shifts portions of machine-learning workloads running in its data centers, decreasing overall facility power demand during hours of system stress.[5][7] One gigawatt matters because it is not a laboratory demonstration; it is roughly the output of a large natural-gas plant, or the consumption of about 750,000 homes, and it begins to resemble utility-scale infrastructure — except that it is made of software and scheduling policy rather than turbines.[6]
The lineage of this capability runs directly through Google’s earlier carbon-aware systems. The 2021 Carbon-Intelligent Compute Management platform, described in a paper by Ana Radovanović and colleagues that later appeared in IEEE Transactions on Power Systems, delayed temporally flexible workloads against day-ahead carbon and demand forecasts across Google’s entire fleet.[9] The company then extended the system geographically, announcing in 2021 that it could shift moveable compute tasks between data centers based on regional hourly carbon-free-energy availability.[8] Radovanović had described the trajectory from the start — the plan, she said in 2020, was to
“shift load in both time and location”
— Ana Radovanović, Technical Lead for Carbon-Intelligent Computing, Google [43]
to maximize the reduction in grid-level emissions. The progression is significant: Carbon Optimization → Demand Response → Grid Flexibility → Workload Wheeling. The original objective was sustainability. The emerging objective is capacity. The same scheduler that once chased clean electrons now negotiates interconnection timelines — Google explicitly credits its demand-response capability with allowing large loads to be interconnected more quickly and with reducing the need to build new transmission and power plants.[7]
4.2 Amazon: Region Flex and the Power-Constrained Cloud
Amazon provides a different, and in some respects more radical, case. Region Flex was not created as an electricity-market program, and Amazon earns no demand-response revenue from it. Nevertheless, power availability and datacenter-capacity limitations have reportedly become important forces shaping its evolution: internal documents describe retail teams leaving Dublin explicitly to mitigate expansion risk from AWS power constraints, more than one hundred software migrations planned in the grocery business alone, services redistributed to Frankfurt and Zaragoza at infrastructure costs ten to fifteen percent higher, and roughly ninety million dollars in one-time migration spending budgeted for 2025.[1] This is arguably even more significant than Google’s program, because it shows electricity constraints beginning to influence application architecture itself. Amazon’s services historically benefited from concentration in giant AWS hubs — concentration delivered economies of scale, operational simplicity, and low intra-region latency. Now electricity scarcity makes concentration a liability, and the architecture responds by becoming more geographically distributed, even at a measurable cost premium. The implication is profound: grid scarcity can rewrite software architecture. When the power system cannot flex, the software does.
It is worth noting what Amazon’s scale makes possible here. The company guided roughly $200 billion in 2026 capital expenditure — the largest figure of any hyperscaler — and reported in its Q2 2026 earnings that AWS grew 36.7 percent year over year, its fastest rate in eighteen quarters, with its AI and in-house chip businesses each reaching annual revenue run rates of about $25 billion.[35][33] A fleet growing that quickly, under binding power constraints in its oldest hubs, has every incentive to make regional distribution — and eventually power-aware placement — a first-class architectural principle rather than an emergency measure.
The Dublin story deserves a moment of its own, because it is the clearest existing case of a whole national grid disciplining a hyperscale region — and of the cloud responding by redistributing work rather than by waiting. Ireland became one of the world’s great datacenter concentrations during the 2010s, and by the early 2020s the consequences were visible in national statistics: datacenter power consumption grew 31 percent in a single year and reached roughly 18 percent of all electricity used in the country, with the International Energy Agency warning the share could approach a third of national consumption, and EirGrid, the transmission operator, acknowledging constraints on what it terms large energy users.[47] The operational reality filtered down to individual cloud customers, some of whom reported being unable to obtain GPU capacity in the AWS Dublin region at all — the facilities were, in one user’s memorable phrase, “maxed out power-wise” — and being pointed instead toward spare capacity in Sweden and elsewhere in the European Union.[47] Read alongside the Region Flex documents, the sequence is a complete miniature of this paper’s thesis: a grid reached its limit; the limit propagated upward into the cloud’s commercial offering; and the cloud’s answer was not merely to queue for megawatts but to move the work — first ad hoc, customer by customer, and now systematically, migration plan by migration plan. Dublin is what the end state of an unwheelable architecture looks like. Region Flex is what the beginning of a wheelable one looks like.
4.3 PJM: From Passive Customer to Controllable Large Load
PJM’s 2026 proposals demonstrate how grid institutions increasingly view data centers: no longer merely as another class of utility customer, but as controllable components of emergency grid management. The sequence is instructive. In its capacity auction for the delivery year beginning mid-2028, PJM failed to secure enough capacity to meet its twenty percent reserve requirement, leaving a 6.8-GW shortfall as prices hit their cap; over four auctions, surging data-center-driven demand had added an estimated $29.4 billion to capacity costs, according to PJM’s independent market monitor.[2][48] The board’s response, released July 27, 2026 and filed at FERC in early August, paired a one-time backstop capacity auction with an Interim Resource Adequacy Service under which new large loads that do not bring their own generation by June 1, 2027, and have not otherwise secured supply, will be subject to curtailment before pre-emergency load management is deployed — that is, before ordinary customers are touched.[2][3][41] A Large Load Registry will catalog every such facility, its service area, and whether it brings its own supply, providing what PJM calls the critical data transparency needed to establish load-reduction priorities.[3] Julia Hoos of Aurora Energy Research captured the institutional shift — PJM, historically reluctant to intrude on state jurisdiction, is now
“making a definitive request to the states to accomplish what it needs.”
— Julia Hoos, Head of USA East, Aurora Energy Research [40]
PJM forecasts that data centers and other large loads will add 30 to 34 GW of new demand by the early 2030s and as much as 70 GW by 2038.[40] The direction of institutional travel is unmistakable, and it had a preview: in May 2026, under a Department of Energy emergency order issued during unseasonable heat, PJM was authorized to curtail data centers with backup generation as a last resort before rolling blackouts — the first time the tool had been placed formally in the operator’s hands.[38] For the Workload Wheeling thesis, the salient point is this: a data center that can wheel work away from PJM during those hours converts a compliance burden into a routine scheduling event. The regulatory environment is, in effect, creating the price signal that makes mobility valuable.
4.4 FERC: The Federal Frame Arrives
On June 18, 2026, the Federal Energy Regulatory Commission issued six show-cause orders under Section 206 of the Federal Power Act to all six jurisdictional grid operators — PJM, MISO, SPP, CAISO, ISO New England, and NYISO — preliminarily finding that existing tariffs may be unjust and unreasonable because they do not adequately address the integration of large loads and co-located loads, and directing responsive filings within sixty days.[14][15] The orders, organized by Commissioner Rosner around four pillars — protecting consumers, safeguarding reliability, enhancing transparency, and fostering innovation — target speculative interconnection requests with escalating readiness requirements, direct cost-recovery agreements so large loads pay their share, and, critically for this paper, recognize that load flexibility can avoid inefficient and costly transmission buildout.[13][14][15] Rosner’s remarks also noted that literal co-location of load and generation is not the only way to achieve faster, more efficient grid connections — an observation that opens conceptual space for exactly the kind of fleet-level, software-mediated flexibility this paper describes.[13] When the federal regulator writes flexibility into the legal architecture of interconnection, workload mobility stops being a private optimization and becomes a bankable regulatory asset.
4.5 ERCOT: Location Starts to Matter Institutionally
Texas provides another piece, and in some ways the most structurally interesting one. Facing an interconnection queue of more than 438,000 MW of large-load requests — roughly eighty-nine percent of it data centers — the Public Utility Commission of Texas approved, on June 18, 2026, ERCOT’s new Batch Zero process, effective July 11, 2026.[16][17][18][45] Batch Zero replaces project-by-project study with a single system-wide evaluation of all qualifying loads of 75 MW and above: ERCOT studies the cohort together, determines what the transmission grid can actually support, allocates megawatts to each project for years one through six, and produces a coordinated transmission construction plan, with binding capacity allocations expected by spring 2027.[17][18] The framework also formally credits on-site and co-located generation, allowing it to offset the transmission capacity a project requires.[17] The institutional message reinforces the central logic of Workload Wheeling: a megawatt is not equally available everywhere. A 500-MW datacenter in one county and a 500-MW datacenter somewhere else are not equivalent to the power system, and Texas will now say so explicitly, in advance, with numbers. Location therefore has value — and if software can move some demand between locations after energization, the operator holds an option the interconnection study never priced.
4.6 NVIDIA, Emerald AI, and the Portability Question
NVIDIA enters the framework from a different direction: not as a load, but as the architect of the machinery that makes loads programmable. At CERAWeek in March 2026, NVIDIA and Emerald AI announced a collaboration with AES, Constellation, Invenergy, NextEra Energy, Nscale, and Vistra to advance a class of AI factories that connect to the grid faster and operate as flexible energy assets, built on the Vera Rubin DSX reference design and its DSX Flex software library for connecting AI factories to power-grid services.[27] The lineage behind that announcement is a sequence of live field demonstrations: the original Phoenix trial with Oracle Cloud Infrastructure, NVIDIA, EPRI, and Salt River Project, in which a 256-GPU cluster cut its power draw by twenty-five percent for three full hours during a real summer peak event while every workload stayed within its service-level guarantees;[50][51] and companion work with National Grid in the United Kingdom, where software slowed a live, grid-connected data center’s power-hungry chips within moments of a simulated national demand spike, with flexible jobs deferred while priority jobs continued.[52][28] NVIDIA’s case-study material reports demand reductions of roughly forty percent in under a minute, and the partners argue such flexibility can help unlock up to 100 GW of latent U.S. grid capacity — a figure that consciously echoes the Duke headroom analysis.[27][29] Emerald AI’s founder, Varun Sivaram, states the design philosophy in terms this paper would endorse: AI factories are
“too valuable to be treated as either passive loads or permanent islands”
— Varun Sivaram, Founder and CEO, Emerald AI [27]
— and, with orchestration, they
“The workloads that drive AI factory energy use can now be flexible.”
— Varun Sivaram, Founder and CEO, Emerald AI, in NVIDIA’s account of the demonstrations [28]
The deeper significance for Workload Wheeling is homogenization. The more uniform AI accelerator infrastructure becomes across clouds, new clouds, and hyperscalers — common chips, common interconnects, common container and checkpoint formats, common grid-interface libraries — the easier computational portability becomes. The future battle may therefore involve more than raw GPU performance. Infrastructure providers may compete on checkpoint portability; accelerator interoperability; orchestration quality; network fabrics; storage synchronization; distributed inference; scheduler intelligence; and power awareness. The winner in the AI infrastructure economy may not simply operate the fastest datacenter. It may operate the most movable compute fleet.
4.7 The Academic Literature, 2020–2026: From Proof-of-Concept to Planning Tool
Beneath the corporate and regulatory developments runs a maturing research literature, and its trajectory over six years mirrors the thesis of this paper. The first wave, 2020–2022, established feasibility: Radovanović and colleagues documented Google’s production system for time-shifting flexible work against carbon forecasts,[9] while Wiesner, Thamsen, and co-authors quantified, in “Let’s Wait Awhile” (2021), how much carbon temporal workload shifting could save in public clouds and situated it within LBNL’s earlier field taxonomy of data-center demand-response strategies — postponing load, migrating load, idling equipment, adjusting cooling.[44] The second wave, 2024–2025, quantified system value: the Duke Rethinking Load Growth study introduced curtailment-enabled headroom and put national numbers — 76 to 126 GW — on what small, well-timed flexibility is worth;[10][11] the MIT CEEPR team embedded flexible data centers in capacity-expansion models of three U.S. regions and measured cost savings of two to five percent alongside regionally contingent emissions effects;[24][23] and Harvard’s Electricity Law Initiative, in Eliza Martin and Ari Peskoe’s widely cited Extracting Profits from the Public, documented the ratemaking side of the boom — nearly fifty regulatory proceedings in which utility contracts risk socializing Big Tech’s power costs — warning that
“data centers are upending this long-standing model”
— Eliza Martin and Ari Peskoe, Harvard Law School Environmental and Energy Law Program [30]
of shared-cost utility service.[30] The third wave, 2026, is operational and spatial: University of Alberta researchers Yize Chen and Xiaogui Zheng, in “To Defer or To Shift?” (April 2026), modeled temporal deferral and geographic shifting jointly inside grid capacity-expansion planning and found flexible AI load can reduce grid investment and operational costs by three to twenty-one percent depending on location, flexibility range, and system conditions — while also showing diminishing returns to ever-longer deferral windows, a sober correction to naive optimism;[37] parallel 2026 work on security-constrained unit commitment, frequency support, and market-driven scheduling has begun to ask not whether wheeling-class flexibility works, but how the grid should dispatch, compensate, and defend against it.[37] MIT’s energy-systems researchers frame the destination clearly. Deepjyoti Deka of the MIT Energy Initiative, whose team studies multi-company flexible workload adjustment, notes the core spatial problem Workload Wheeling addresses — a grid operator may have generation available elsewhere, but
“the wires may not have sufficient capacity to carry the electricity”
— Deepjyoti Deka, Research Scientist, MIT Energy Initiative [25]
to where it is wanted. When the wires cannot carry the electricity to the computation, the remaining degree of freedom is to carry the computation to the electricity. And the stakes of getting this right are set by the sheer scale of the load: as MIT’s Noman Bashir observes, a generative-AI training cluster can draw
“seven or eight times more energy than a typical computing workload.”
— Noman Bashir, Computing and Climate Impact Fellow, MIT [26]
The International Energy Agency supplies the global frame. Its landmark Energy and AI report projects data-center electricity consumption more than doubling from roughly 415 TWh in 2024 to around 945 TWh by 2030 — more than Japan’s entire present consumption — with AI-optimized facilities quadrupling and data centers accounting for almost half of U.S. electricity-demand growth over the period; its 2026 update holds the central trajectory while noting that bottlenecks are forcing a scramble for exactly the solutions this paper describes, from flexible data centers to storage.[21][22] Fatih Birol, the IEA’s Executive Director, has framed both halves of the story. At the report’s launch:
“AI is one of the biggest stories in the energy world today.”
— Dr. Fatih Birol, Executive Director, International Energy Agency [21]
And a year later, surveying the 2026 landscape, he noted that the agency was early in recognizing that
“there is no AI without energy”
— Dr. Fatih Birol, Executive Director, International Energy Agency [22]
— and that AI, still an energy taker, is becoming an energy maker, driving forward next-generation nuclear, flexible data centres, and long-duration storage.[22] The literature, the companies, the regulators, and the international institutions have converged on the same object from four directions. What none of them has yet fully named is the unifying mechanism. That mechanism is Workload Wheeling.

Section 5 — The Economics, Politics, and Geopolitics of Workload Wheeling: If Compute Can Move, Everything Around the Datacenter Changes
Infrastructure concepts rarely stay inside their engineering boundaries, and Workload Wheeling will not stay inside its scheduler. If a meaningful fraction of the largest new industrial load in modern history becomes movable in time and space, the movement reprices assets, redistributes bargaining power, and redraws maps. This section works outward from the utility contract to the statehouse to the international system, asking at each level the same question: who gains, who pays, and what new institutions become necessary when demand itself becomes a portfolio? The section is deliberately speculative in places, because the phenomena are young; but each speculation is anchored to a 2025–2026 development already on the public record.
5.1 A New Bargain Between Utilities and Hyperscalers
Utilities historically operate under a simple covenant: customers consume whenever they decide to, and the utility plans, builds, and charges accordingly. The obligation to serve is unconditional, and the cost of that unconditionality is embedded in every rate. Flexible AI introduces a different possible contract: the grid provides faster or cheaper access in exchange for the datacenter agreeing to become flexible. Google’s utility agreements are the prototype — demand-response provisions that the company says help new data centers connect more quickly to local grids, embedded directly in long-term energy contracts and counted by planners as capacity.[5][6] Federal policy has been converging on the same bargain: the White House’s July 2025 AI Action Plan, “Winning the Race,” devotes an entire pillar to developing a grid that matches the pace of AI innovation,[31] and its March 2026 Ratepayer Protection Pledge points the same direction from the contracting side: participating AI companies commit to negotiate separate rate structures, to pay for the infrastructure built to serve them whether they use it or not, and to coordinate with grid operators to make backup resources available during emergencies.[32] A future tariff architecture could therefore formally differentiate two products. Firm AI Load: electricity must remain available except under extraordinary emergencies, priced to carry the full cost of the peak infrastructure it requires. Flexible AI Load: the operator agrees that specified workloads can be shifted, reduced, or relocated during defined system conditions, priced at a discount that reflects the capacity, transmission, and reserve costs it avoids. The spread between those two prices is, in effect, the market price of workload mobility — and once a price exists, an industry will organize itself to capture it.
Table 7. A Two-Product Tariff for AI Load
| Attribute | Firm AI Load | Flexible AI Load |
| Availability commitment | Near-unconditional service | Service subject to defined flexibility events |
| Grid planning treatment | Full peak-coincident capacity obligation | Reduced capacity obligation; counted partly as a resource |
| Interconnection | Standard queue; full network upgrades | Accelerated access reflecting curtailment-enabled headroom |
| Price | Full embedded cost of peak infrastructure | Discounted; shares avoided capacity and transmission value |
| Analog in current practice | Traditional large industrial tariff | Google’s 2026 demand-response contracts; PJM IRAS-class service |
5.2 Workload Flexibility Becomes an Infrastructure Asset
A one-gigawatt datacenter possessing 150 MW of wheelable demand is economically a different object from a completely inflexible one-gigawatt facility, even though the two look identical from the road. That difference will not remain informal. Flexibility can be expected to affect interconnection priority — ERCOT’s Batch Zero already credits on-site generation against transmission requirements, and crediting verified workload flexibility is the natural next step;[18] utility tariffs and capacity payments — PJM’s emerging framework effectively penalizes large loads that bring nothing to the system and spares those that do;[2][3] financing and insurance — a facility with contracted flexibility revenue and lower curtailment risk is a different credit; power-purchase agreements; site selection; and grid-upgrade cost allocations, which FERC now insists must be tied to the costs large loads actually impose.[14] The market, in short, will begin placing a value on how movable a megawatt of AI demand is. Harvard’s ratepayer research supplies the political economy behind this repricing: when flexibility is absent, the costs of serving inflexible giants migrate — through the subjectivity and complexity of ratemaking — onto the public, a dynamic that Martin and Peskoe documented across nearly fifty proceedings and that has made data-center rate design one of the most contested questions in American utility regulation.[30] Verified flexibility is one of the few mechanisms that can genuinely shrink the costs being fought over, rather than merely reallocating them.
5.3 State Competition Changes
Governors have spent the AI boom competing for physical datacenter investment — the construction jobs, the property tax base, the announcement photograph. Workload Wheeling creates an additional and subtler competition: not merely where companies construct AI factories, but where hyperscalers prefer to execute discretionary AI workloads once the fleet exists. States with abundant generation; reliable and uncongested transmission; predictable permitting; favorable electricity tariffs; plentiful water or water-efficient cooling; strong fiber; and attractive tax treatment could become destinations for wheelable compute — earning energy sales and utilization on infrastructure others financed. The datacenter becomes analogous to a port. Some ports carry more physical trade than their hinterlands would predict because geography, infrastructure, and price make routing through them attractive; Singapore and Rotterdam are entrepôts because flows choose them. Some states may similarly become computational entrepôts — while others discover that hosting the building is not the same as hosting the work. A state that wins the ribbon-cutting but prices its peak-hour electricity badly may find its shiny facility running at low utilization on August afternoons, its flexible workloads quietly executing in a neighboring interconnection. Conversely, the DOE Needs Study’s congestion maps — Virginia and Texas leading current data-center demand, with Arizona and Oregon among the fastest growers — begin to read like an early routing table for this competition.[19][42]
5.4 Workload Wheeling Meets the Five-Layer AI Economy
The framework developed here reveals why the Five-Layer AI Economy should increasingly be analyzed as one integrated system rather than five stacked industries. Layer 1 — Energy — determines available electrical capacity. Layer 2 — Chips — defines accelerator capability and compatibility, and therefore the technical feasibility of moving work between sites. Layer 3 — Datacenters — provides the physical execution locations and their grid interfaces. Layer 4 — Models — creates workloads with different mobility characteristics. Layer 5 — Applications and Agents — determines latency, jurisdiction, and reliability requirements, and therefore fixes each workload’s position on the Mobility Spectrum. Workload Wheeling is the control loop circulating through all five: a power shortage in Layer 1 can cause software in Layer 5 to relocate computation through Layers 4, 3, and 2, and the aggregate of those relocations flows back down as a changed demand profile that Layer 1’s planners must model. That is a completely different industrial system from conventional cloud computing, in which the energy layer was an invisible utility bill. It is a cyber-physical economy in which the layers negotiate with one another continuously, in machine time.
Table 8. The Five-Layer AI Economy and the Wheeling Control Loop
| Layer | Domain | Role in Workload Wheeling |
| Layer 1 | Energy | Sets where and when megawatts are available, at what price, carbon intensity, and reliability |
| Layer 2 | Chips | Sets where compatible accelerators exist; homogenization increases portability |
| Layer 3 | Datacenters | Physical nodes joining energy to accelerators; each node’s grid contract defines its flexibility |
| Layer 4 | Models | Generates jobs with intrinsic mobility properties (training, tuning, evaluation, batch inference) |
| Layer 5 | Applications & Agents | Imposes latency, privacy, and jurisdiction constraints; originates the demand to be scheduled |
5.5 The Geopolitical Boundary
There is an important limit, and it is political rather than technical. Compute may be digitally mobile, but sovereignty is not. A U.S. workload may not be permitted to move to China under any circumstances. European personal data faces restrictions on where it can be processed, and European policymakers are constructing sovereign-cloud requirements precisely to keep certain computations inside certain borders. Defense workloads have strict location and clearance requirements. Export-controlled accelerators exist only in approved jurisdictions, so the chips themselves define a legal geography before any scheduler runs. Models trained on sensitive data may face national-security restrictions on where their weights may reside. Workload Wheeling will therefore operate inside geopolitical containers. The United States may eventually constitute one wheeling domain — enormous, internally liquid, spanning four time zones and three interconnections. Europe another, with stricter data-residency partitions inside it. China another, walled by both sides. Sovereign AI programs — in the Gulf, in India, in Southeast Asia — may establish smaller national domains whose whole premise is that certain intelligence workloads shall not wheel out. The paradoxical result: the global cloud becomes more geographically flexible inside political blocs while becoming less flexible across them. The map of the intelligence economy will look less like the borderless internet of the 2000s and more like the trading blocs of the twentieth century — free movement within, controlled crossings between.
5.6 New Systemic Risks
Flexibility also produces risk, and a candid treatment must dwell on it. If thousands of megawatts of AI demand come under the control of automated schedulers, synchronized decisions begin to matter to grid stability in a way individual decisions never did. Imagine that electricity prices spike in Virginia; thousands of AI workloads move toward Texas; five minutes later ERCOT experiences an unexpected load surge that its operators never scheduled and its models never forecast. Or a software error misinterprets a market signal and moves excessive compute simultaneously — the power-system equivalent of the 2010 flash crash, except the falling object is frequency rather than price. Or attackers compromise the orchestration layer itself, gaining the ability to swing gigawatts of national demand with a configuration change. Recent power-systems research is already modeling exactly these pathologies — price-chasing datacenter fleets inducing oscillations and grid-security violations that must be mitigated in market design.[37] FERC’s Rosner has likewise flagged that loads able to change consumption in seconds create voltage and frequency phenomena the interconnection rules never anticipated.[13][14] The workload-routing layer, in other words, is becoming critical infrastructure, and it will need to be governed as such: cybersecurity standards for power-aware schedulers analogous to NERC CIP for grid control systems; rate limits and ramp-rate obligations on coordinated fleet movements; disclosure regimes so operators can see the routing layer’s aggregate posture; and perhaps, eventually, circuit breakers. The more successfully compute becomes electrically aware, the more urgently electricity institutions must become computationally aware.
5.7 The Political Question for 2026 and Beyond
All of this converges on a practical reframing for policymakers. Politicians should stop asking only how many datacenters their state should approve, and begin asking what kind of AI load they are approving. Is it firm? Interruptible? Temporally movable? Geographically movable? Supported by onsite generation? Capable of contracted demand response? Eligible for emergency relocation of its workloads? ERCOT’s Batch Zero and PJM’s Large Load Registry are the first institutional acknowledgments that these questions have answers and that the answers differ by facility.[17][3] The distinctions could determine whether AI expansion raises electricity prices for households — the outcome the Harvard ratepayer literature warns of and that has already animated moratorium politics from Ohio to the Mid-Atlantic[30][40] — or becomes part of the solution to electricity scarcity, as the Duke headroom analysis and Google’s contracts suggest it can be.[10][5] The same gigawatt can be either, and the deciding variable is mobility.

Section 6 — What Have We Learned? Seven Pillars of Workload Wheeling
A long argument deserves a compact summary, but a compact summary should not be a mere list; each pillar below is a claim with evidence behind it and consequences ahead of it. The first five correspond to the load-bearing walls of the analysis; the sixth and seventh are the honest caveats and governance obligations without which the structure would be marketing rather than analysis.
Pillar 1 — Compute Demand Is Becoming Programmable
Traditional industrial electricity demand is tied to physical production: the kiln fires when the plant runs. AI creates something categorically different — a large industrial demand of which a meaningful portion can be scheduled through software, in time and in space. The location and timing of electricity consumption become programmable variables rather than facts of geography. This is the foundational principle behind Workload Wheeling, and it is already operating in production at Google, demonstrably at gigawatt contract scale.[5][9]
Pillar 2 — The Hyperscaler Fleet Matters More Than the Individual Datacenter
A single AI datacenter is geographically fixed. A network of twenty, fifty, or one hundred datacenters is not operationally fixed in the same way. The more workloads that can move across that fleet, the more the operator can treat electricity capacity as a portfolio of options rather than a set of local obligations. The future strategic asset is the bundle — Power + Chips + Datacenters + Fiber + Workload Orchestration — and Amazon’s Region Flex shows the bundle being actively re-architected under power constraint, at a knowing cost premium, because the portfolio is worth more than the concentration savings it replaces.[1]
Pillar 3 — Flexible Compute Can Become a Grid Resource
Datacenters have traditionally been described as threats to electricity reliability, and PJM’s 2026 emergency proposals show how real the threat framing has become.[2][3] But the characterization is incomplete. A sufficiently flexible AI factory can reduce demand, delay demand, or relocate demand — and utilities are now counting that capability as capacity in formal reliability assessments.[6] Google’s one-gigawatt demand-response milestone, the NVIDIA–Emerald AI demonstrations with their forty-percent-in-under-a-minute reductions, and EPRI’s DCFlex valuation work together mark the transition from theory to commercial-scale deployment.[5][28][29] The AI datacenter can be simultaneously a gigantic electricity consumer and a controllable grid resource; which face it shows depends on contract design.
Pillar 4 — The Cheapest Megawatt May Be the One the Workload Does Not Need Locally
America unquestionably needs more generation and transmission; DOE’s 2026 Needs Study makes clear that the grid must expand substantially to accommodate AI and other large-load growth, and nothing about flexibility repeals that requirement.[19] But infrastructure expansion and computational flexibility are not mutually exclusive — they are complements, and the Duke analysis quantifies the complementarity: a quarter of one percent of annual curtailment unlocks roughly 76 GW of headroom on infrastructure that already exists.[10] The strategic planning question for every operator and regulator becomes: which megawatts absolutely must be delivered to this datacenter, and which computational jobs could instead be executed somewhere else? Answering it well saves capital, shortens interconnection queues, and improves reliability. Answering it badly builds peakers for jobs that could have waited until midnight.
Pillar 5 — AI Infrastructure Is Becoming a Cyber-Physical Network
The AI economy is no longer simply software running on computers, nor simply GPUs consuming electricity. It is becoming a cyber-physical system in which electricity conditions affect software decisions; software decisions affect grid demand; grid prices affect workload placement; workload placement affects datacenter investment; and datacenter investment affects regional power planning. Layer 1 and Layer 5 are beginning to communicate directly, in machine time, through the wheeling control loop. That feedback loop — not any single facility, chip, or model — may prove to be the defining infrastructure development of the next phase of the AI economy.
Pillar 6 — Mobility Has Limits, and Honest Accounting Requires Naming Them
Workload Wheeling is a marginal instrument, and margins can be exhausted. The firm core of AI demand — interactive inference, autonomous systems, regulated workloads — cannot wheel, and it is growing. Geographic shifting presupposes network bandwidth, data-residency clearance, and destination headroom that may not exist when most needed; correlated weather can stress many regions at once; and the University of Alberta results show diminishing returns to longer deferral windows and benefits that vary from three to twenty-one percent depending on where the flexibility sits on the network — location matters even for the cure.[37] The MIT CEEPR findings add an environmental caution: in low-renewable regions, flexibility that chases cheap off-peak power can increase emissions even as it lowers costs.[23][24] Anyone selling Workload Wheeling as a substitute for building generation and transmission is overselling it. It is a multiplier on infrastructure, not an alternative to infrastructure.
Pillar 7 — The Orchestration Layer Must Be Governed as Critical Infrastructure
Once schedulers command gigawatts, scheduler governance is grid governance. The institutions that certified turbines and relays must learn to certify placement algorithms and fleet ramp behavior; cybersecurity regimes must extend to the routing layer; market rules must anticipate synchronized machine responses to price signals; and transparency obligations — registries like PJM’s, batch studies like ERCOT’s, cost-recovery agreements like FERC’s — must give the physical system visibility into the digital system that now steers a growing share of its demand.[3][17][14] The reward for building these institutions well is substantial: a power system that gains, in flexible compute, the largest new source of demand-side dispatchability since the invention of the interruptible tariff. The penalty for building them badly is a new category of systemic risk wired directly into both the grid and the intelligence economy at once.

Conclusion: Why “Workload Wheeling” Fits the Emerging AI Economy
The first phase of the AI infrastructure boom was dominated by a straightforward question: where can we find enough electricity to power all these GPUs? The question produced a straightforward, colossal answer — the roughly $700 to $760 billion the four largest hyperscalers will spend on capital expenditure in 2026 alone, nearly double the prior year, much of it chasing power as much as silicon.[33][34][36] That question remains enormously important, and nothing in this paper diminishes it. But it is no longer sufficient, because a second question has emerged alongside it: when electricity cannot reach the AI workload quickly enough, can the AI workload move toward the electricity?
The evidence assembled here suggests the second question is already being answered in the affirmative, piece by piece, by institutions that do not yet use a common name for what they are doing. Amazon’s Region Flex initiative shows power constraints beginning to reshape the geographic architecture of computing itself.[1] Google’s demand-response agreements demonstrate that machine-learning workloads can be shifted or limited at contracted, gigawatt scale to accommodate grid conditions.[5] FERC is examining a generation of enormous loads that can change electricity consumption at speeds traditional industrial customers rarely could, and is rewriting interconnection law around them.[13][14] ERCOT is evaluating giant loads collectively, according to where the power system can actually accommodate them.[17] PJM is constructing emergency procedures that would require the largest electricity users to react before households are interrupted.[2][3] NVIDIA is shipping grid-responsiveness as a native property of its AI factory reference designs.[27] The research community — Duke, MIT, Harvard, LBNL, EPRI, the IEA, and a fast-thickening arXiv literature — has moved from asking whether computational flexibility is real to measuring what it is worth and how it should be dispatched.[10][23][30][39][21][37]
Individually, these developments can look unrelated: one is cloud architecture, one is demand response, one is utility regulation, one is transmission planning, one is chip strategy, one is scholarship. Together they reveal something more consequential. Compute is becoming electrically aware — and electricity systems are becoming computationally aware — at the same moment, from both directions, under the pressure of the same scarcity.
That is why Workload Wheeling fits this moment. Wheeling was historically about transporting energy across geography so that production could remain where it was. Workload Wheeling describes an industrial economy in which the productive activity itself becomes partially transportable. The power plant does not move. The transmission line does not move. The datacenter does not move. The GPUs do not move. The intelligence workload moves. And that seemingly small architectural distinction could reshape where AI factories are built, how utilities price them, how hyperscalers design networks, how states compete for infrastructure, how regulators define large loads, and ultimately how efficiently the Five-Layer AI Economy converts scarce electricity into artificial intelligence.
The old cloud asked: where is the nearest available computer? The AI cloud increasingly asks: where is the best available combination of GPU, network, electricity, and time? When that question is answered automatically, continuously, and safely across hundreds of facilities and millions of computational jobs — under tariffs that price flexibility, regulations that verify it, and governance that constrains it — Workload Wheeling will no longer be a theoretical concept or a working paper’s coinage. It will be what it is already becoming in fragments across Mountain View, Seattle, Valley Forge, Austin, and Santa Clara: an operating principle of the intelligence economy.

Footnotes and Endnotes:
[1] Eugene Kim, Business Insider (reported and summarized by Paul Drecksler, Shopifreaks), “Amazon Is Spreading Its E-Commerce Systems Across More AWS Regions as Power Constraints Squeeze Its Biggest Data Center Hubs” (Region Flex), August 2026. https://www.shopifreaks.com/amazon-is-spreading-its-e-commerce-systems-across-more-aws-regions-as-power-constraints-squeeze-its-biggest-data-center-hubs/
[2] Ethan Howland, Utility Dive, “PJM Board Proposes Backstop Capacity Auction, Data Center Curtailment Plans,” July 28, 2026. https://www.utilitydive.com/news/pjm-board-backstop-capacity-auction-data-center-curtailment/826347/
[3] Ethan Howland, Utility Dive, “PJM Files Backstop Auction Plan at FERC to Meet Capacity Shortfall” (Interim Resource Adequacy Service; Large Load Registry), August 2026. https://www.utilitydive.com/news/pjm-backstop-capacity-auction-ferc-data-centers/826792/
[4] Jeff St. John, Canary Media, “PJM’s Big New Data Center Plan: Make the States Figure It Out,” August 2026. https://www.canarymedia.com/articles/data-centers/pjm-data-center-plan
[5] Michael Terrell, Google (The Keyword), “We’ve Signed 1 GW of Data Center Demand Response with Utility Partners,” March 19, 2026. https://blog.google/innovation-and-ai/infrastructure-and-cloud/global-network/demand-response-data-center-milestone/
[6] mGrid, “Google Secures 1 GW Demand Response From 5 U.S. Utilities” (Indiana Michigan Power, TVA, Entergy Arkansas, Minnesota Power, DTE Energy), March 20, 2026. https://mgrid.org/2026/03/20/google-1gw-demand-response-utility-contracts/
[7] Reuters, “Google Agrees to Curb Power Use for AI Data Centers to Ease Strain on US Grid When Demand Surges.” https://de.tradingview.com/news/reuters.com,2025:newsml_L1N3TT1CX:0-google-agrees-to-curb-power-use-for-ai-data-centers-to-ease-strain-on-us-grid-when-demand-surges
[8] Google (The Keyword), “Using Location to Reduce Our Computing Carbon Footprint” (carbon-intelligent computing moves flexible tasks between data centers based on regional hourly carbon-free energy availability), May 2021. https://blog.google/company-news/outreach-and-initiatives/sustainability/carbon-aware-computing-location/
[9] Ana Radovanović et al., “Carbon-Aware Computing for Datacenters,” arXiv:2106.11750; IEEE Transactions on Power Systems, Vol. 38 (2023), pp. 1270–1280. https://arxiv.org/abs/2106.11750
[10] Tyler H. Norris, Tim Profeta, Dalia Patiño-Echeverri, and Adam Cowie-Haskell, “Rethinking Load Growth: Assessing the Potential for Integration of Large Flexible Loads in US Power Systems,” Nicholas Institute for Energy, Environment & Sustainability, Duke University, February 2025. https://nicholasinstitute.duke.edu/publications/rethinking-load-growth
[11] Ethan Howland, Utility Dive, “Existing US Grid Can Handle ‘Significant’ New Flexible Load: Report” (including comment by Costa Samaras, Carnegie Mellon University), February 2025. https://www.utilitydive.com/news/us-grid-headroom-flexible-load-data-center-ai-ev-duke-report/739767/
[12] Inside Climate News, “Flexibility Will Go a Long Way Toward Managing the Grid of the Near Future, Researchers Say” (quoting Tyler Norris, Duke University), February 2025. https://insideclimatenews.org/news/11022025/grid-flexibility-ai-data-centers/
[13] Federal Energy Regulatory Commission, “Commissioner Rosner’s Remarks on the Large Load Show Cause Orders, E-7 to E-12,” June 18, 2026 Open Meeting. https://www.ferc.gov/news-events/news/commissioner-rosners-remarks-large-load-show-cause-orders-e-7-e-12-june-18-2026
[14] Ethan Howland, Utility Dive, “6 Takeaways From FERC’s Data Center Interconnection Decision,” June 22, 2026. https://www.utilitydive.com/news/ferc-doe-data-center-interconnection/823360/
[15] Holland & Knight, “FERC Advances New Oversight Framework for Large Loads and Co-Located Generation,” June 23, 2026. https://www.hklaw.com/en/insights/publications/2026/06/ferc-advances-new-oversight-framework-for-large-loads
[16] ERCOT, “Large Load Integration” (Batch Zero process documentation, PGRR145/NPRR1325), 2026. https://www.ercot.com/services/rq/large-load-integration
[17] ERCOT, Market Notice M-B062326-01, “Implementation of the Batch Zero Process” (PUCT approval of PGRR145 and NPRR1325, effective July 11, 2026), June 23, 2026. https://www.ercot.com/services/comm/mkt_notices/M-B062326-01
[18] Sheppard Mullin, “PUCT Approves ‘Batch Zero’ Framework That Will Provide a Pathway for Large Load to Connect to the Grid,” June 23, 2026. https://www.sheppard.com/insights/blogs/puct-approves-batch-zero-framework-that-will-provide-a-pathway-for-large-load-to-connect-to-the-grid-and-receive-electricity
[19] U.S. Department of Energy, Office of Electricity, “National Transmission Needs Study” (2026 draft for consultation and public comment, released July 9, 2026), 2026. https://www.energy.gov/oe/national-transmission-needs-study
[20] U.S. Department of Energy, Office of Electricity, “DOE’s Office of Electricity Publishes 2026 Draft National Transmission Needs Study to Strengthen America’s Grid” (statement of Assistant Secretary Catherine Jereza), July 9, 2026. https://www.energy.gov/oe/articles/does-office-electricity-publishes-2026-draft-national-transmission-needs-study
[21] International Energy Agency, “AI Is Set to Drive Surging Electricity Demand from Data Centres” (Energy and AI special report; remarks of Executive Director Fatih Birol), April 2025. https://www.iea.org/news/ai-is-set-to-drive-surging-electricity-demand-from-data-centres-while-offering-the-potential-to-transform-how-the-energy-sector-works
[22] International Energy Agency, “Data Centre Electricity Use Surged in 2025, Even with Tightening Bottlenecks Driving a Scramble for Solutions” (remarks of Fatih Birol), April 2026. https://www.iea.org/news/data-centre-electricity-use-surged-in-2025-even-with-tightening-bottlenecks-driving-a-scramble-for-solutions
[23] MIT News, “How Data Centers Can Better Manage Energy Use” (Christopher Knittel, Juan Ramon Senga, Shen Wang; iScience, 2026), June 26, 2026. https://news.mit.edu/2026/how-data-centers-can-better-manage-energy-use-0626
[24] Christopher R. Knittel, Juan Ramon L. Senga, and Shen Wang, MIT Center for Energy and Environmental Policy Research, “Flexible Data Centers and the Grid: Lower Costs, Higher Emissions?” Working Paper 2025-14, July 2025. https://ceepr.mit.edu/wp-content/uploads/2025/07/MIT-CEEPR-WP-2025-14-Brief.pdf
[25] MIT Energy Initiative, “The Multi-Faceted Challenge of Powering AI” (quoting research scientist Deepjyoti Deka), January 2025. https://energy.mit.edu/news/the-multi-faceted-challenge-of-powering-ai/
[26] MIT News, “Explained: Generative AI’s Environmental Impact” (quoting Noman Bashir, MIT Climate and Sustainability Consortium / CSAIL), January 2025. https://news.mit.edu/2025/explained-generative-ai-environmental-impact-0117
[27] NVIDIA Newsroom, “NVIDIA and Emerald AI Join Leading Energy Companies to Pioneer Flexible AI Factories as Grid Assets” (CERAWeek 2026; quoting Varun Sivaram), March 23, 2026. https://nvidianews.nvidia.com/news/nvidia-and-emerald-ai-join-leading-energy-companies-to-pioneer-flexible-ai-factories-as-grid-assets
[28] NVIDIA Blog, “How AI Factories Can Help Relieve Grid Stress” (Emerald Conductor demonstrations; statement of Varun Sivaram), 2025–2026. https://blogs.nvidia.com/blog/ai-factories-flexible-power-use
[29] NVIDIA, “How Emerald AI Makes AI Factories Power-Flexible” (case study: grid-flexible AI factories cut power demand roughly 40% in under a minute; NVIDIA Vera Rubin DSX; up to 100 GW of untapped capacity), 2026. https://www.nvidia.com/en-us/case-studies/emerald-ai/
[30] Eliza Martin and Ari Peskoe, Harvard Law School Environmental and Energy Law Program, “Extracting Profits from the Public: How Utility Ratepayers Are Paying for Big Tech’s Power,” March 2025. https://eelp.law.harvard.edu/extracting-profits-from-the-public-how-utility-ratepayers-are-paying-for-big-techs-power/
[31] The White House, “Winning the Race: America’s AI Action Plan,” July 23, 2025. https://www.whitehouse.gov/wp-content/uploads/2025/07/Americas-AI-Action-Plan.pdf
[32] The White House, “Fact Sheet: President Donald J. Trump Advances Energy Affordability with the Ratepayer Protection Pledge,” March 2026. https://www.whitehouse.gov/fact-sheets/2026/03/fact-sheet-president-donald-j-trump-advances-energy-affordability-with-the-ratepayer-protection-pledge/
[33] Ari Levy et al., CNBC, “Amazon, Meta and Microsoft Face Skeptical Investors This Week After Google Report Sparked Sell-Off” (Q2 2026 hyperscaler earnings and capex), July 28, 2026. https://www.cnbc.com/2026/07/28/hyperscalers-face-higher-capex-scrutiny-after-alphabet-report-panned.html
[34] Value Add VC, “AI Spending Tracker 2026: $725B by Big Tech” (compilation of company earnings calls and guidance through Q2 2026; Alphabet ceiling raised to $205B at Q2 2026 earnings; Meta raised guidance twice), August 2026. https://valueaddvc.com/ai-spending
[35] Amazon.com, Inc., “Amazon Q2 2026 Earnings Report” (AWS grew 36.7% year over year, fastest in 18 quarters, to a $169 billion annualized run rate; AI and chips businesses each eclipsed $25 billion run rates; statement of Andy Jassy), July 30, 2026. https://www.aboutamazon.com/news/company-news/amazon-earnings-q2-2026-report
[36] AL Capital Advisory, “AI Capex Cycle 2026: The $775–800B Hyperscaler Buildout” (Q2 CY2026 earnings disclosures: Google Cloud +82%, AWS +37%, Azure +43%), August 2026. https://alcapitaladvisory.com/research/intelligence/ai-infrastructure.html
[37] Yize Chen and Xiaogui Zheng, University of Alberta, “To Defer or To Shift? The Role of AI Data Center Flexibility on Grid Interconnection,” arXiv:2604.05376, April 8, 2026. https://arxiv.org/pdf/2604.05376
[38] Ethan Howland, Utility Dive, “PJM Gets Emergency Approval to Curtail Data Centers, Large Loads During Hot Weather” (U.S. DOE emergency order), May 19, 2026. https://www.utilitydive.com/news/pjm-doe-emergency-order-curtail-data-centers/820571/
[39] Lawrence Berkeley National Laboratory, Energy Technologies Area, “Driving Energy Innovation in Data Centers” (2024 U.S. Data Center Energy Usage Report; DOE Data Center Load Flexibility Workshop; REFLEX program), 2024–2026. https://eta.lbl.gov/data-centers
[40] Jeff St. John, Canary Media (republished by Ohio Capital Journal), “PJM’s Big New Data Center Plan: Make the States Figure It Out” (quoting Julia Hoos, Aurora Energy Research), August 10, 2026. https://ohiocapitaljournal.com/2026/08/10/pjms-big-new-data-center-plan-make-the-states-figure-it-out/
[41] Network World, “AI Data Centers in the US May Face Power Cuts Under PJM Reliability Proposal,” July 2026. https://www.networkworld.com/article/4202800/ai-data-centers-in-the-us-may-face-power-cuts-under-pjm-reliability-proposal.html
[42] Shane Snider, Data Center Knowledge, “DOE: AI Data Centers Are Transforming America’s Transmission Map” (85,000 circuit-miles energized 2016–2024; ~$11 billion in 2023 congestion costs; Virginia and Texas lead current data-center demand with Arizona and Oregon among the fastest growers), July 2026. https://www.datacenterknowledge.com/energy-power-supply/doe-says-ai-data-centers-are-rewriting-america-s-transmission-map
[43] Ana Radovanović, Google (The Keyword), “Our Data Centers Now Work Harder When the Sun Shines and Wind Blows” (carbon-intelligent computing platform; plan to shift load in both time and location), April 2020. https://blog.google/inside-google/infrastructure/data-centers-work-harder-sun-shines-wind-blows/
[44] Philipp Wiesner, Ilja Behnke, Dominik Scheinert, Kordian Gontarska, and Lauritz Thamsen, “Let’s Wait Awhile: How Temporal Workload Shifting Can Reduce Carbon Emissions in the Cloud,” arXiv:2110.13234 (ACM/IFIP Middleware 2021). https://arxiv.org/pdf/2110.13234
[45] Foley & Lardner LLP, “ERCOT’s Proposed ‘Batch Zero’ Process: What Developers of Large Loads Need to Know,” March 19, 2026. https://www.foley.com/p/102mnfa/ercots-proposed-batch-zero-process-what-developers-of-large-loads-need-to-kno/
[46] Futurum Group, “AI Capex 2026: The $690B Infrastructure Sprint,” February 12, 2026. https://futurumgroup.com/insights/ai-capex-2026-the-690b-infrastructure-sprint/
[47] Dan Robinson, The Register, “AWS Resource Restrictions Point to Datacenter Power Issues” (AWS users report GPU capacity in Dublin “maxed out power-wise,” with customers pointed to spare capacity in Sweden and other parts of the EU), April 9, 2024. https://www.theregister.com/on-prem/2024/04/09/aws-resource-restrictions-point-to-datacenter-power-issues/456885
[48] The Logical Insight, “PJM Plans Backstop Auction for 6.8 GW Shortfall, Curtailment for Data Centers” (Joseph Bowring, Monitoring Analytics: data-center load increased capacity costs by $29.4 billion over the last four auctions), August 1, 2026. https://www.thelogicalinsight.com/energy/pjm-plans-backstop-auction-for-68-gw-shortfall-curtailment-f–94c1b5df-4d8f-4932-a679-e91f72a66c14
[49] Southwest Alliance for Contractor Compliance and Accountability (SWACCA), “DOE Transmission Study Highlights Growing Data Center Demand” (PJM transmission bottlenecks limiting delivery to Dominion Energy’s data-center load; Virginia, Texas, Arizona, Oregon), July 2026. https://swacca.org/doe-transmission-study-highlights-growing-data-center-demand/
[50] Ayse K. Coskun et al. (Emerald AI, Boston University, with Oracle Cloud Infrastructure, NVIDIA, EPRI, and Salt River Project), field-demonstration paper on transforming AI data centers into flexible grid resources via the Emerald Conductor platform — 25% cluster power reduction sustained for three hours in Phoenix while maintaining AI quality-of-service guarantees, arXiv:2507.00909, July 2025. https://arxiv.org/abs/2507.00909
[51] American Public Power Association, “SRP Participates in Artificial Intelligence Data Center Demonstration” (EPRI DCFlex Phoenix demonstration with Oracle Cloud Infrastructure, NVIDIA, and Salt River Project), July 2025. https://www.publicpower.org/periodical/article/srp-participates-artificial-intelligence-data-center-demonstration
[52] MIT Technology Review, “Want to Get a Data Center Online Quickly? Give It Some Flex.” (National Grid demonstration slowing chips at a London data center during a simulated demand spike, December 2025; quoting Ayse Coskun, chief scientist, Emerald AI), June 16, 2026. https://www.technologyreview.com/2026/06/16/1138591/data-center-online-quickly-electric-grid-flex/



