Introduction: The Day a GPU-Hour Became a Financial Instrument

On October 5, 2026, pending completion of regulatory review, something genuinely unusual is scheduled to appear on the screens of derivatives traders. It will not be a barrel of West Texas Intermediate crude, a bushel of corn, an ounce of gold, a Treasury note, or a megawatt-hour of electricity delivered to a named hub. It will be computing power. CME Group, the world’s largest derivatives marketplace, and Silicon Data, a GPU market-intelligence firm backed by the global trading house DRW, plan to list two contracts on NYMEX: the Silicon Data H100 Rental Index Futures and the Silicon Data B200 Rental Index Futures, cash-settled instruments designed to track indexes measuring the hourly rental cost of two generations of Nvidia graphics processing units, with each contract representing approximately one month’s worth of rent for the underlying accelerator [1][7]. CME has prepared its trading infrastructure to recognize an entirely new unit of measure—GPU-H, or GPU Hours—and has framed the product explicitly as a hedging and investment vehicle for businesses that need to manage the cost of compute, the processing power and hardware infrastructure that machines require to train and run artificial-intelligence models [1].

The symbolism is difficult to overstate, and the principals involved have made no attempt to understate it. When CME and Silicon Data first announced their partnership in May 2026, CME Chairman and Chief Executive Officer Terry Duffy reached for the most loaded comparison available in the history of commodity markets [2]:

“As the backbone of the digital economy, compute is the new oil of the 21st century.”

— Terry Duffy, Chairman and CEO, CME Group [2]

Carmen Li, the former DRW trader who founded Silicon Data and now serves as its Chief Executive Officer, described the problem her benchmarks are meant to solve in terms that any student of opaque, bilateral, over-the-counter markets will recognize immediately. For years, she observed, two companies buying identical GPU capacity could pay wildly different prices with no way of knowing who received the better deal; the futures market finally gives them a benchmark to check against [3]:

“A public, tradable reference price for the resource every AI system runs on.”

— Carmen Li, CEO, Silicon Data [3]

CME is not alone, and that fact matters enormously for how this market will develop. One week after CME’s initial May announcement, Intercontinental Exchange—the owner of the New York Stock Exchange and the operator of the Brent crude benchmark complex—announced plans with Ornn, a financial-infrastructure company building capital markets for artificial intelligence, to develop a suite of U.S.-dollar-denominated, cash-settled GPU compute futures based on the Ornn Compute Price Index, a benchmark constructed exclusively from printed transactions rather than surveys or scraped quotes, covering Nvidia’s H100, H200, and B200 accelerators as well as the RTX 5090, with additional hardware types to follow as the market matures [4][5]. Trabue Bland, ICE’s Senior Vice President of Futures Markets, justified the initiative in language that could have been lifted from the founding documents of any twentieth-century commodity exchange, declaring that the compute market is [4]:

“In desperate need of a globally accepted pricing mechanism and risk management tool.”

— Trabue Bland, SVP of Futures Markets, Intercontinental Exchange [4]

Nor do the two giants exhaust the field. Architect Financial Technologies announced in January 2026 that it would develop GPU and DRAM perpetual futures referencing Ornn’s benchmark data, meaning that three separate exchange-level initiatives to financialize compute appeared within roughly four months of one another [10]. Ornn’s index is already distributed on the Bloomberg Terminal and has attracted more than four hundred datacenter operators, investors, and artificial-intelligence companies to its platform, while observers have compared the moment to the early 1980s, when competing exchanges raced to establish the benchmark contracts for crude oil and natural gas that still govern global energy finance four decades later [6]. Even the White House has blessed the project: the administration’s AI Action Plan of July 2025 explicitly recommended ensuring access to large-scale computing power for startups and academics by [41]:

“Improving the financial market for compute.”

— America’s AI Action Plan, The White House, July 2025 [41]

At first glance, all of this looks like the natural maturation of artificial-intelligence infrastructure, and in one important sense it is. Airlines hedge jet fuel. Farmers hedge grain. Utilities hedge natural gas and electricity. Mining companies hedge metals. If artificial-intelligence laboratories, hyperscalers, neoclouds, datacenter developers, lenders, and ordinary enterprises are going to spend on the order of three-quarters of a trillion dollars annually building and renting computational capacity—and the earnings reports of 2026 confirm that they are—then they too need mechanisms for managing price risk. A frontier-model company that expects to require millions of GPU-hours twelve months from now has every rational reason to want protection against a sudden increase in compute prices. Conversely, a datacenter operator that has borrowed billions of dollars against racks of accelerators has an equally rational reason to want protection against a future collapse in rental rates. The hedging logic is impeccable, the demand is real, and the volatility that justifies the product is empirically documented: one-year H100 rental contracts rose nearly forty percent in five months, from roughly $1.70 per GPU-hour in October 2025 to approximately $2.35 per GPU-hour by March 2026, according to SemiAnalysis data, while Silicon Data’s own index recorded a ten percent jump in on-demand H100 rates in a single four-week window between December 9, 2025 and January 6, 2026—the largest short-term move the firm had tracked since mid-2025—even as A100 and B200 pricing stayed essentially flat over the same period [8][9].

But then comes the problem, and it is a problem that no contract specification can legislate away.

The simplest illustration requires nothing more than two numbers. Across the market in August 2026, on-demand pricing for a single Nvidia H100 80GB accelerator spanned roughly $2.19 to $11.06 per GPU-hour depending on provider, contract terms, and region—a fivefold spread for hardware carrying the same part number [8]. Zoom out further and the dispersion becomes almost surreal: H100 listings have been observed ranging from $0.72 to $15.14 per GPU-hour within a single day, a twenty-one-fold spread across just twenty-four marketplaces, and independent pricing trackers place the full observable band between $1.38 and $12.29 per GPU-hour once hyperscaler on-demand instances and marketplace spot capacity are both included [10][12]. A processor might carry the same Nvidia designation on both ends of that spread, yet its rental price differs by an order of magnitude. And beneath the provider-level fragmentation lies a deeper geographic problem that will be familiar to anyone who has traded electricity or natural gas: an H100-hour available in Texas is not necessarily equivalent, in economic usefulness, to an H100-hour available in Northern Virginia, Frankfurt, Riyadh, or Tokyo, because compute cannot be loaded onto a tanker and shipped from where it is cheap to where it is expensive.

That discrepancy points toward the central argument of this paper, which can be stated in a single sentence: a financial exchange can standardize a contract far more easily than the physical economy can standardize compute.

Two H100s may share the same chip architecture while existing inside completely different economic systems. One might sit inside a tightly interconnected cluster containing tens of thousands of GPUs linked by high-bandwidth optical networking; another might operate inside a small, bandwidth-starved deployment where the same silicon delivers a fraction of the effective training throughput. One datacenter may enjoy inexpensive and reliable electricity under a decade-long power purchase agreement; another may face congestion charges, curtailment risk, or a utility tariff that has been repriced upward by an angry state commission. One facility may sit adjacent to the data, the customers, and the software stack that a particular application requires; another may be separated from them by thousands of miles and tens of milliseconds. One provider may guarantee availability, enterprise-grade security, and integration with hundreds of adjacent cloud services; another may sell opportunistic, interruptible surplus. One GPU-hour may be immediately usable for a frontier-model training run; another, nominally identical GPU-hour may be almost useless to that particular workload. The twenty-one-fold intraday price spread is not evidence of an irrational market. It is evidence that the market is pricing twenty-one different products that happen to share a part number [10].

This is the paradox at the heart of the emerging compute market, and it is the paradox from which everything else in this paper follows: artificial-intelligence compute is becoming financially standardized before it has become economically fungible—and there are structural reasons to believe it never fully will.

That paradox is what I call Compute Friction.

Compute Friction describes the persistent difference between the standardized unit that financial markets would like compute to become and the heterogeneous physical, geographic, electrical, networking, software, regulatory, and operational reality through which compute is actually delivered. It is the collection of obstacles preventing one GPU-hour from being economically interchangeable with another GPU-hour, and—as the later sections of this paper will argue—it is also the collection of signals through which the artificial-intelligence economy reveals where its real constraints are located.

Within this broader framework, the paper introduces a narrower and more operational concept: GPU Basis. In traditional commodity markets, basis describes the difference between a local or physical price and the benchmark futures price—the discount that North Dakota crude suffers against Cushing delivery, or the premium that a congested load zone pays over the hub price of electricity. A directly analogous concept can now be applied to artificial-intelligence infrastructure:


GPU Basis = Effective Delivered GPU-Hour Value − Benchmark GPU-Hour Price


But effective delivered value must be understood far more broadly than the advertised rental rate. It incorporates hardware generation, cluster topology, interconnect performance, achieved utilization, cloud environment, geographic location, electricity economics, latency, availability, data-transfer costs, software compatibility, contractual restrictions, service-level guarantees, and regulatory exposure—every characteristic that affects what a customer can actually accomplish with that GPU-hour. The futures contract, in other words, will be the beginning of compute price discovery rather than the end of it. If this market develops as its architects intend, an entirely new financial vocabulary will follow: regional compute differentials, hardware-generation spreads, cloud premiums, latency premiums, topology premiums, utilization discounts, sovereign-compute premiums, power-adjusted compute prices, and eventually forward curves showing what markets expect different classes of artificial-intelligence infrastructure to cost years into the future. Indeed, the first pieces of that vocabulary already exist: Silicon Data now publishes separate neo-cloud and hyperscaler readings of its H100 index precisely because the two segments trade so differently, and it has documented the hyperscaler premium compressing from roughly 250 percent to about 120 percent across the seven regions it tracks as hyperscaler rates fell by 26 to 54 percent while non-hyperscaler rates recovered [9][11].

The important development is consequently much larger than the phrase GPU futures suggests. Artificial intelligence is beginning to construct a financial market around one of the most important productive resources of the twenty-first century—and it is doing so at a moment when that resource remains stubbornly, irreducibly heterogeneous. The emergence of compute futures does not eliminate the differences among artificial-intelligence infrastructure. It makes those differences financially visible, measurable, and—for the first time—tradable. That is the story this paper sets out to tell.


Why I Chose the Title “Compute Friction”

I chose Compute Friction because the most important economic story unfolding here is not simply that compute can now serve as the underlying reference for a futures contract; announcements of new derivatives products arrive every year, and most of them matter to no one but their product managers. The deeper story is that financial markets are attempting to commoditize something whose usefulness remains extraordinarily heterogeneous, and the attempt itself will generate an entire economics of the gap. A standardized H100 or B200 contract can establish a reference price, but it cannot make electricity prices identical across regions, eliminate the speed of light, equalize cluster topology, dissolve data-sovereignty rules, standardize cloud service bundles, guarantee utilization, or teleport a GPU from Virginia to Texas, Europe, the Middle East, or Asia. Those persistent differences are the friction separating financial compute from physically delivered compute, and friction—as every engineer and every commodity trader knows—is not an incidental nuisance. It is where the energy goes, and it is where the information is.

The title also fits the longer-term direction of the artificial-intelligence economy because Compute Friction survives individual generations of hardware, which is precisely what a durable analytical framework must do. H100s will become legacy accelerators; the market is already watching them age, with legacy-generation GPU pricing sliding from $1.77 to $1.14 per hour over a recent twelve-month window as older SKUs lose their high-end hyperscaler anchors [14]. B200s will eventually be surpassed by Rubin-generation architectures, which Nvidia’s own chief executive has already framed as part of a half-trillion-dollar demand pipeline [17]. AMD accelerators, Google TPUs, Amazon Trainium, Microsoft Maia, custom OpenAI silicon, and successive families of Chinese accelerators will keep redrawing the hardware landscape. But differences in location, power, networking, software, regulation, reliability, capacity, and workload suitability will remain, because they are rooted in physics, geography, and law rather than in any particular chip. Compute Friction is therefore not a paper about one futures launch. It is a framework for understanding why the emerging global market for artificial intelligence may develop a common financial benchmark without ever developing a truly uniform physical commodity—and why the distance between those two things will become one of the most closely watched numbers in the world economy.


Section 1: From GPU-Hour to Financial Asset — The Birth of the Compute Derivatives Market


1.1 October 5, 2026: The Institutionalization of the GPU-Hour

Every commodity market has a founding date that later generations treat as a hinge, even when the early volumes were trivial: the 1983 launch of NYMEX crude oil futures, the 1990 launch of natural gas futures at Henry Hub, the first electricity contracts of the mid-1990s. October 5, 2026 is designed to be such a date for compute. The contracts themselves are, on the surface, modest instruments: two cash-settled monthly futures listed on NYMEX under CME Group’s rules, one tracking a Silicon Data index of hourly Nvidia H100 rental costs and one tracking the equivalent index for the Blackwell-generation B200, with each contract representing a month’s worth of rent for the underlying accelerator [1][7]. The underlying indexes are constructed from roughly 150,000 daily verified pricing records drawn from fifty to one hundred platforms—hyperscalers, neoclouds, and marketplaces—across forty to fifty countries, with every observation standardized for rental-term length, cluster scale, and interconnect so that the index reflects underlying compute cost rather than the headline rate on any single provider’s website [10][11].

The significance, however, is institutional as much as financial, and it is worth dwelling on why. Once a productive resource has benchmarks, futures curves, clearing mechanisms, margin and collateral treatment, market makers, and a community of financial participants, businesses can begin organizing investment decisions around expectations of its future price rather than around bilateral negotiation and rumor. Lenders can underwrite against a public curve. Auditors can mark positions to an observable settlement. Regulators can monitor open interest. Journalists can report a number that everyone recognizes. The resource acquires what economists would call common knowledge of price, and common knowledge changes behavior even among firms that never touch a futures contract. The transition can be written as a simple chain, and every link in it has now been forged for compute: GPU → GPU-hour → benchmark index → futures contract → financial asset. CME’s decision to formally introduce GPU-H as a unit of measure inside its trading infrastructure is the small technical detail that captures the whole transformation, in the same way that the standardization of the barrel—forty-two gallons, a unit inherited from Pennsylvania cooperage—once captured the industrialization of oil [1].


1.2 Why Artificial Intelligence Suddenly Needs a Futures Market

A futures market is an expensive institution to build and a difficult one to sustain; most proposed contracts die of illiquidity. The reason two major exchanges are nonetheless racing to list compute is that the underlying exposure has become enormous, volatile, and structurally unhedgeable by other means, and the scale of that exposure is documented in the corporate earnings record through the second quarter of 2026 with a clarity that borders on the alarming. The four largest hyperscalers—Microsoft, Alphabet, Amazon, and Meta—deployed approximately $301 billion of capital expenditure in the first half of 2026 alone, and their updated guidance points to combined annual capital expenditure of roughly $732.5 billion for the year, a figure more than 150 percent higher than forecasts issued only two years earlier—and once Oracle’s program is included, the five-company total climbs toward and beyond $690 billion [18][20]. Statista places the four-company total near $760 billion for 2026 against $413 billion in 2025; Goldman Sachs now projects $5.3 trillion of combined hyperscaler capital expenditure from fiscal 2025 through fiscal 2030 and a baseline of $7.6 trillion of aggregate AI-related investment between 2026 and 2031 across compute, datacenters, and power [19][20][21]. Oracle has guided to roughly $70 billion of capital expenditure in its fiscal 2027, and Amazon’s chief executive told investors that capacity constraints are likely to persist through 2027, with projected 2028 demand already informing infrastructure planning [23]. Nvidia’s second quarter of fiscal 2027, reported in August 2026, delivered $96 billion of revenue with $89 billion from the data-center segment and guidance of $108 billion for the following quarter—numbers that make a single chip company’s quarterly data-center revenue larger than the annual revenue of most Fortune 100 industrial firms [16].


Table 1. The Exposure Behind the Hedge: Hyperscaler Capital Expenditure, 2025 vs. 2026

Company2025 Capex (approx.)2026 Guidance (approx.)Notes from the 2026 Earnings Round
Amazon~$125B~$200BCapacity constraints expected to persist through 2027; 2028 demand already informing planning [23]
Microsoft~$88B~$190B (tracking)Apparent guidance reduction reflects an accounting change, not lower investment [19][21][22]
Alphabet~$91B$175–185BRaised 2026 guidance with Q2 2026 results; shares fell 7% on the announcement [19][20]
Meta~$72B$115–135B (raised toward $125–145B)Guidance lifted citing memory-chip prices and added datacenter costs; external cloud sales “on the table” [22]
Oracle~$21B~$50B (FY26) → ~$70B (FY27)Clearest forward capex outlook of the group [18][23]
Four-company total~$410–413B~$725–760BUp roughly 77% year over year; H1 2026 alone: $301B deployed [19][20][21]

Sources: company earnings reports and guidance through Q2 2026, as compiled in [18][19][20][21][23]. Figures are approximate calendar-year totals; Microsoft figure reflects analyst tracking of calendar-year spend.


Against that backdrop, compute-price volatility ceases to be a procurement inconvenience and becomes a balance-sheet event, and it becomes one differently for every class of participant. For an artificial-intelligence laboratory, compute is the dominant input cost, and an unhedged forty-percent move in rental rates—of exactly the kind observed between October 2025 and March 2026—can vaporize the margin assumptions underneath an entire product roadmap [8]. For a neocloud, compute is inventory-like productive capacity whose market value fluctuates daily against a mountain of fixed-rate debt. For a datacenter developer, forward compute prices determine future tenant economics and therefore the viability of projects that take three to five years to energize. For a lender, compute prices govern collateral values and debt-service capacity on tens of billions of dollars of GPU-backed credit. For Nvidia and other accelerator manufacturers, rental economics feed directly into customers’ willingness to purchase the next hardware generation, since no operator will buy chips whose rentable output cannot cover their financing. And for investors at large, forward compute prices offer something the market has never had: a continuously traded, forward-looking signal of expected artificial-intelligence demand, distinct from equity multiples and analyst surveys. The futures market therefore sits at the intersection of several very different balance sheets, which is exactly the condition under which derivatives markets historically achieve liquidity—hedgers with opposite natural exposures meeting speculators willing to warehouse the residual risk.


1.3 CME Versus ICE: The Competition to Define the Reference Price of Compute

The parallel efforts of CME with Silicon Data and ICE with Ornn deserve to be read as more than routine product competition, because the two initiatives embody genuinely different philosophies of what a compute price is. Silicon Data’s index family is built for breadth: approximately 150,000 daily verified pricing records across as many as one hundred platforms in up to fifty countries, spanning on-demand, spot, and reserved lease types, normalized for term, cluster scale, and interconnect [10][11]. Ornn’s Compute Price Index is built for depth: it is constructed exclusively from printed transactions—negotiated levels between datacenters and compute buyers—on the explicit theory that posted list prices often diverge from what large buyers actually pay, and that scraped indices can misstate where the market actually clears [5][10]. Both philosophies are defensible; both have distinguished ancestors in energy and rates benchmarking; and they can produce materially different numbers for what is nominally the same GPU-hour.

A crucial question therefore emerges, and it is a question with consequences measured eventually in billions of dollars of financial exposure: who gets to define the reference price of artificial intelligence? The competition concerns far more than exchange market share. A benchmark weighted toward hyperscaler list pricing will read structurally higher than one weighted toward neocloud transactions, because hyperscaler capacity has carried premiums of 120 to 250 percent over neocloud floors in recent tracking, and hyperscalers typically carry newly launched accelerators at three to five times neocloud rates during the first year of a generation’s life [9][14]. A benchmark built from long-term contracted rates will move differently from one built from spot listings, as demonstrated vividly when one-year H100 commitments trended upward through a window in which the on-demand median stayed roughly flat [14]. Whichever methodology attracts liquidity first will shape how the world understands the cost of intelligence—how lenders haircut GPU collateral, how developers underwrite datacenters, how policymakers judge whether compute is scarce or abundant. The history of LIBOR, of oil price reporting agencies, and of power-market index manipulation all counsel that benchmark design is never a neutral technical exercise, a theme to which Section 5 returns at length.


1.4 What Compute Futures Could Actually Hedge

It is worth being concrete about the hedging architecture this market makes possible, because the abstractions can obscure how directly these instruments map onto existing pain. An artificial-intelligence laboratory planning a training campaign for the first half of 2027 could buy strips of monthly H100 or B200 futures today, converting an uncertain future rental bill into a known cost and insulating its model-development budget from the kind of ten-percent monthly spikes the spot market has already produced [9]. A cloud provider holding a large unhedged inventory of accelerators could sell futures against expected rental income, protecting revenue against the price decay that reliably arrives as each hardware generation ages toward legacy status [14]. A datacenter lender could require borrowers to hedge a portion of projected GPU revenue as a condition of credit, exactly as project-finance lenders in the power sector require hedges on merchant electricity exposure, thereby stabilizing the debt-service coverage ratios on which multi-billion-dollar facilities depend. An enterprise planning a major inference deployment could lock in a portion of its anticipated compute expenditure the way transportation firms lock in fuel. An infrastructure investor could read the forward curve directly into a project-finance model, replacing consultant guesswork with a market-cleared expectation. The consequence, if liquidity develops, could be substantial: infrastructure whose future revenue was previously difficult to quantify becomes easier to underwrite, and easier underwriting lowers the cost of capital for the entire buildout—an effect the White House AI Action Plan anticipated when it framed a financial market for compute as a mechanism for broadening access to computing power itself [41].


1.5 Why Financialization Does Not Equal Commoditization

And yet the entire remainder of this paper turns on a distinction that the celebratory language around these launches tends to blur: financial standardization is not physical fungibility. Oil markets work as well as they do because barrels meeting a specified grade and delivery point can genuinely be treated as close substitutes; West Texas Intermediate at Cushing is a real, deliverable, interchangeable thing, and where physical oil differs—in sulfur content, density, or location—the market prices the difference as a basis relationship rather than pretending it away. GPU-hours cannot be treated so easily, because the characteristics that determine their economic value are not reducible to a grade specification. A compute futures contract can produce a standardized financial exposure without producing standardized computing capability, and the gap between the two is not a temporary imperfection that liquidity will erode. It is a structural feature rooted in cluster topology, electricity economics, geography, latency, software ecosystems, service quality, and law. That structural gap is Compute Friction, and Section 2 dissects it layer by layer.


Section 2: The Anatomy of Compute Friction

Compute Friction is not one thing. It is a stack of at least ten distinguishable frictions, each with its own physics, its own economics, and its own implications for how far the futures benchmark can travel before it stops describing reality. This section takes them in turn, and then assembles them into the GPU Basis framework that the rest of the paper deploys. The guiding discipline throughout is borrowed from energy economics: whenever a market participant says the price of compute, the correct response is to ask which compute, delivered where, over what network, under what contract, powered by whose electrons, and usable by which software. Every one of those questions identifies a friction, and every friction is a candidate component of basis.


2.1 Hardware Friction: An H100 Is Not a B200—and a GPU Is Not Always the Right Accelerator

The most visible friction is generational. Different accelerator generations embody different memory capacities, memory bandwidth, interconnect standards, energy efficiency, software support, and therefore fundamentally different workload economics; an H100 purchased for roughly $25,000 to $40,000 depending on form factor and configuration is not a discount B200 any more than a ten-year-old aircraft is a discount new one [13]. The market already prices this ruthlessly. Over a recent twelve-week window of daily index data, the H200 rental index rose 14.4 percent while the H100 rose 7.0 percent, and the premium between the two chips swung from 5 percent down to 0.4 percent and then up to a 16.2 percent peak—a live, continuously repriced hardware-generation spread of exactly the kind that crack spreads and quality differentials represent in oil [9]. Meanwhile on-demand rates for the newest generation of accelerators—B200, B300, MI300X, RTX 5090—roughly doubled over the past year even as mainstream cards held a tight band, and legacy-generation pricing slid from $1.77 to $1.14 per hour as older SKUs lost their hyperscaler anchors [14].

The eventual consequence is a family of forward curves rather than a single one: H100, H200, B200, later Nvidia architectures such as the Rubin generation, AMD accelerators, Google TPUs, AWS Trainium, Microsoft Maia, custom silicon from the model laboratories themselves, and the maturing Chinese accelerator families that export controls have called into existence. ICE’s announced contract suite, which contemplates H100, H200, B200, and RTX 5090 references with additional GPU types to follow, is an early institutional acknowledgment of this multiplicity [5]. For hedgers, the multiplicity means hardware basis risk: a company may hedge against one benchmark while actually consuming another accelerator type, and the correlation between the two—like the correlation between jet fuel and the heating-oil contracts airlines once used to hedge it—will be strong, imperfect, and occasionally treacherous. Silicon Data’s own analysis concedes the point directly, noting that an application running custom configurations in a specific datacenter does not hedge cleanly with a standardized H100 contract and predicting that a mature basis market around the standardized contract is poised to develop [9].


2.2 Topology Friction: A GPU Is Only as Valuable as the Cluster Around It

Frontier artificial intelligence increasingly depends on systems rather than isolated processors, and this is the friction that most sharply separates compute from every commodity that preceded it. A barrel of oil does not become more valuable because it sits next to other barrels; a GPU emphatically does. The economic value of an accelerator depends on the high-speed interconnects that link it to its neighbors, the switches and optical networking that stitch racks into pods and pods into clusters, the memory architecture and storage systems that feed it data, the CPU coordination that schedules it, the rack architecture and cooling that determine its sustained clock speeds, the orchestration software that assigns it work, the total cluster size that bounds the largest model it can help train, and the failure rates that determine how often a ten-thousand-GPU job must checkpoint and restart. Ten thousand well-connected accelerators possess fundamentally different economic value from ten thousand poorly interconnected accelerators, even when the twenty thousand chips are individually identical.

The implication deserves to be stated as a principle, because it will eventually have to be encoded in benchmark methodologies and lending covenants alike: GPU quantity is not equivalent to effective compute capacity. Silicon Data’s indexes already standardize observations for cluster scale and interconnect precisely because unadjusted comparisons would be meaningless, which is itself an admission that the raw GPU-hour is not a uniform unit [11]. A futures contract settles against the standardized abstraction; a training run executes against the physical topology; and the difference between the two—call it the topology premium—is pure basis. As models scale and the industry’s unit of account migrates from the chip to the rack to the multi-building campus, topology friction will grow rather than shrink, because the share of system value residing in the network fabric between chips keeps rising with every hardware generation.


2.3 Cloud Friction: Hyperscalers, New Clouds, Private Clusters, and Surplus Capacity

The third friction is the provider layer, and it is the one for which the empirical record is now overwhelming. The observed spread between hyperscaler and non-hyperscaler H100 pricing has been one of the defining facts of the rental market: hyperscaler rates fell 26 to 54 percent by region after the first quarter of 2025 while non-hyperscaler rates recovered, compressing the hyperscaler premium from roughly 250 percent to about 120 percent across the seven regions Silicon Data tracks—which is to say that even after dramatic compression, an enterprise renting through a major cloud still pays more than double the neocloud rate for nominally identical silicon [9]. Independent trackers tell the same story from different angles: hyperscaler on-demand H100 instances have been listed as high as $11.68 to $12.29 per hour while marketplace and boutique clouds offered the same chip for $1.38 to $1.49, and hyperscalers typically carry each newly launched accelerator generation at three to five times neocloud floors during its first year [12][14].

It would be a serious analytical mistake to read this spread as pure inefficiency waiting to be arbitraged away, and the persistence of the spread is itself the evidence. Hyperscaler pricing bundles global private networks, managed data services, identity and security infrastructure, enterprise agreements and compliance certifications, software ecosystems, availability guarantees, technical support, and integration with hundreds of adjacent cloud services; frontier laboratories and regulated enterprises demonstrably continue to pay the premium because guaranteed capacity, integrated tooling, and enterprise support are worth that much to them [13]. A smaller provider may offer cheaper raw compute bundled with less of everything else. The same GPU-hour is therefore really two distinct products—closer to the difference between firm and interruptible power, or between contracted pipeline capacity and spot cargoes, than to a simple retail markup. Which of the two products the futures benchmark represents is a methodological choice with billion-dollar consequences, and it is why Silicon Data publishes the hyperscaler and neo-cloud readings separately rather than pretending a single number describes both [11].


Table 2. One Part Number, Many Markets: Observed Nvidia H100 Rental Pricing, 2026

Segment / ChannelObserved Rate (per GPU-hour)Evidence
Marketplace / interruptible spotas low as $0.72–$1.49Lowest intraday marketplace listings and boutique-cloud floors [10][12]
New-cloud composite index (SDH100RT)≈ $2.53Silicon Data standardized like-for-like benchmark reading [11]
One-year committed contracts$1.70 → $2.35 in five monthsSemiAnalysis series, October 2025 to March 2026, a ~40% rise [8]
Managed / boutique on-demand$3.29–$4.29Runpod Secure, Crusoe, Lambda list rates, mid-2026 [13]
Hyperscaler on-demandup to $11.68–$12.29Highest hyperscaler instance listings; 3–5x neocloud floors for new generations [12][14]
Full intraday extreme$0.72–$15.14 (≈21x)Single-day spread across 24 marketplaces [10]

The dispersion is not irrational pricing; it is the market pricing distinct bundles of networking, guarantees, location, and service quality that share a part number [10][12][13].


2.4 Geographic Friction: Texas Compute Is Not Virginia Compute

Location matters, and it matters through channels that compound rather than cancel. A GPU-hour in Northern Virginia exists inside the densest datacenter and fiber agglomeration on earth, adjacent to major cloud regions, enterprises, and government customers—but it also exists inside a grid under visible strain, in a state where nearly three-quarters of surveyed voters blame datacenters for rising electricity costs and where studies project that electricity generation costs could rise as much as 57 percent by the end of the decade under high-demand scenarios [28][30]. A GPU-hour in Texas benefits from a different electricity market design, faster interconnection, abundant land, and a permissive development regime, at the price of exposure to ERCOT scarcity pricing and weather events. Arizona, Ohio, Indiana, Pennsylvania, Michigan, Nevada, and Oregon each present their own combination of electricity supply, transmission congestion, water availability, taxes and incentives, fiber connectivity, permitting speed, climate, land, labor, regulation, and political acceptance—and the political variable is no longer secondary, with seven in ten Americans now opposing an artificial-intelligence datacenter near their homes and 78 percent worried that datacenters will raise their bills [28][30].

Internationally the differences widen further. Compute located in Europe, the Middle East, China, Japan, Korea, India, or Southeast Asia faces different energy costs, different sovereignty requirements, different chip-access regimes under U.S. export controls, different latency to end markets, and different regulatory rules on data, privacy, and model deployment. The result is that compute develops something structurally identical to the locational basis of electricity markets: a persistent, measurable price surface across geography, driven by constraints that cannot be shipped away. The only difference is that electricity markets spent decades building the nodal pricing infrastructure to make that surface visible, while the compute market is only now acquiring its first hub price.


2.5 Power Friction: Every GPU-Hour Contains an Electricity Market

Every accelerator ultimately converts electricity into computation, which means every compute price silently embeds an electricity position, and the scale of that embedded position has become one of the central facts of the world energy economy. The International Energy Agency reports that global datacenter electricity demand grew 17 percent in 2025—with consumption by AI-focused datacenters surging 50 percent—against total global electricity demand growth of only about 3 percent, and projects that datacenter demand will more than double to approximately 945 terawatt-hours by 2030, driven almost entirely by artificial intelligence [24][25][33]. Lawrence Berkeley National Laboratory projects that United States datacenter demand alone will grow from 176 terawatt-hours in 2023—about 4.4 percent of national consumption—to between 325 and 580 terawatt-hours, or 6.7 to 12 percent of national consumption, by 2028 [27]. The capital expenditure of the five largest technology companies surged past $400 billion in 2025 and is set to jump a further 75 percent in 2026, and the IEA notes—remarkably—that power consumption per artificial-intelligence task is simultaneously declining by at least an order of magnitude annually, a pace of efficiency improvement it believes may be unprecedented in energy history, without that efficiency doing anything to slow aggregate demand growth [26].

A compute price consequently contains exposure to wholesale electricity, utility tariffs, natural-gas prices, nuclear availability, renewable generation profiles, transmission capacity, backup generation, battery storage, grid congestion, interconnection queues, and curtailment regimes. The pass-through to households has already become a first-order political fact: PJM, the largest grid operator in the United States, attributes most of a projected $6.3 billion increase in consumer electricity costs over the next three years to datacenter demand; its independent market monitor has estimated more than $20 billion in recent capacity-market costs tied to datacenters; utilities requested a record $31 billion of rate increases in 2025 and another $18.6 billion in the first half of 2026 alone; and peer-reviewed modeling puts the possible rise in national average wholesale electricity costs at 6 to 29 percent by 2030 [29][30][31]. Ari Peskoe, director of the Electricity Law Initiative at Harvard Law School and one of the closest academic observers of this collision, summarized the distributional reality of those regional costs in four words [29]:

“Everybody pays for that.”

— Ari Peskoe, Director, Electricity Law Initiative, Harvard Law School [29]

Steve Clemmer of the Union of Concerned Scientists, whose organization estimates total datacenter-linked electricity costs of $886 billion to $978 billion by 2050, drew the same conclusion about even partial cost pass-through, calling it [32]:

“A pretty significant increase for customers.”

— Steve Clemmer, Director of Energy Research, Union of Concerned Scientists [32]

This is why the Five-Layer AI Economy that structures Section 3 begins with energy: Layer 1 can reprice Layer 2 even when the chip itself never changes. A regional power shortage widens local compute basis without a single transistor differing; a new nuclear unit or transmission line narrows it. Compute derivatives are therefore, indirectly but unavoidably, energy derivatives wearing a silicon costume, and sophisticated participants will trade them as such—hedging the electricity leg separately, arbitraging compute forwards against power forwards, and eventually demanding power-adjusted compute indexes.


2.6 Latency Friction: Distance Has an Economic Price

A theoretically cheap GPU may be commercially expensive if it sits too far from the user, the dataset, the agent, or the adjacent infrastructure on which a workload depends, because the speed of light is the one input cost no amount of capital expenditure can reduce. Training workloads tolerate distance gracefully; a pre-training run cares about throughput, not round-trip time, which is why training clusters migrate toward cheap power in remote geographies. Real-time inference does not tolerate distance, and the emerging agentic layer tolerates it least of all: autonomous systems interacting with users, factories, robots, financial exchanges, vehicles, and industrial equipment place a premium on response time that maps directly into a premium on proximity. The economic consequence is a bifurcation of the compute market along latency lines—remote, power-seeking training compute on one side and metro-adjacent, latency-bound inference compute on the other—with different cost structures, different competitive dynamics, and ultimately different prices for the same silicon. The futures market will initially ignore this distinction because its indexes must, but the physical market already prices it, and the gap between the two is yet another recurring component of GPU Basis: an explicit or implicit latency premium that will become more visible as inference and agentic workloads come to dominate total demand.


2.7 Software Friction: CUDA, APIs, Frameworks, and Switching Costs

Physical hardware does not operate independently of software, and software is where the compute market’s most durable moat lies. A company whose training stack, kernels, tooling, and institutional knowledge are woven into Nvidia’s CUDA ecosystem cannot substitute an alternative accelerator merely because it rents for less per hour; the switching cost—rewriting, revalidating, retraining staff, absorbing schedule risk on models whose launch windows are worth billions—can dwarf any plausible rental saving. Software portability therefore creates a friction that operates like a currency-conversion fee between hardware ecosystems, taxing every attempted arbitrage between them. This is also where Nvidia’s advantage extends beyond semiconductor performance into financial benchmark formation itself: the first regulated compute futures reference Nvidia hardware exclusively—H100 and B200 at CME, an Nvidia-dominated suite at ICE—which means the financial market’s very definition of compute is being written in Nvidia’s units [1][5]. Section 4 develops the consequences; the point here is simply that software friction is what makes hardware basis sticky. In a world of frictionless software portability, price gaps between accelerator families would arbitrage quickly toward performance parity. In the world that actually exists, they can persist for years.


2.8 Utilization Friction: Installed Compute Is Not Productive Compute

The existence of an accelerator guarantees nothing about its productive use. Idle capacity, scheduling gaps, fragmented clusters, insufficient customer demand, networking bottlenecks, maintenance windows, thermal throttling, failure-and-checkpoint overhead on large jobs, and plain workload incompatibility all drive wedges between installed GPU-hours and productive GPU-hours—and the wedge is where fortunes are made and lost, because a fleet financed at eighty-percent utilization assumptions and running at fifty is a slow-motion default regardless of the rental price printed on the screen. The distinction deserves the same canonical status in compute economics that capacity factor holds in power economics: installed GPU-hours are not productive GPU-hours. A mature compute market will eventually need measures based not merely on machine availability but on useful computational output—tokens processed, training FLOPs delivered, jobs completed within service windows—and the first index provider to crack a credible utilization-adjusted or output-based benchmark will have built something closer to the true unit of the intelligence economy than any rental-price index can be. Until then, utilization risk remains almost entirely unhedgeable through the listed contracts, sitting squarely inside GPU Basis, and it is worth noting that the same wedge is at the center of the fiercest accounting controversy in the sector, examined in Section 4: the fight over how quickly a GPU’s earning power decays is, at bottom, a fight about the future utilization and pricing of aging fleets [34][35].


2.9 Service-Level Friction: Reliability Has a Price

Enterprise customers pay considerably more for guaranteed capacity than opportunistic customers pay for interruptible capacity, and the spread between those two prices is neither noise nor gouging; it is the market pricing reliability as a distinct good, exactly as electricity markets have always done. The natural taxonomy, borrowed consciously from energy, is already assembling itself in contract structures across the industry: firm compute, backed by availability guarantees and penalties; interruptible compute, sold cheap because it can be reclaimed; reserved compute, committed for one to three years at rates that—as 2026 demonstrated—can trend upward even while spot drifts sideways [14]; spot compute, the residual sold by the hour; surplus compute, the opportunistic overflow of fleets built for internal workloads, a category that may grow dramatically if Meta proceeds with reported plans to sell excess capacity externally [22]; sovereign compute, delivered inside a specified jurisdiction under specified law at a premium that is fundamentally regulatory; and latency-guaranteed compute, engineered proximity sold as a service level. Each category commands a different price, each maps onto a different customer, and no single futures contract can represent them all. The listed benchmark will inevitably settle closest to one category—most plausibly something between spot and short-term reserved—leaving every other category to trade at a basis to it, which is not a defect of the design but the design itself: benchmarks exist so that differences can be priced against them.


2.10 GPU Basis: Measuring the Gap

The ten frictions above are not a list of complaints about an imperfect market; they are the decomposition of a measurable quantity. Within the Compute Friction framework, GPU Basis is the subordinate analytical tool that turns qualitative heterogeneity into a number that can be tracked, charted, traded, and regulated. Its simplified form is the difference already stated in the Introduction:


GPU Basis = Effective Delivered GPU-Hour Price − Benchmark GPU-Hour Price


Its developed form is an additive decomposition across the frictions this section has catalogued:

Quality-Adjusted GPU Basis = Hardware Differential + Topology Differential + Cloud Differential + Geographic Differential + Power Differential + Latency Differential + Software Differential + Utilization Differential + Service-Level Differential + Regulatory Differential


Table 3. Decomposition of Quality-Adjusted GPU Basis

Basis ComponentUnderlying FrictionIllustrative Evidence
Hardware differentialGeneration, memory, architectureH200 premium over H100 swinging between 0.4% and 16.2% over twelve weeks [9]
Topology differentialInterconnect, cluster scale, fabricIndexes must normalize for cluster scale and interconnect to compare listings at all [11]
Cloud differentialProvider class and service bundleHyperscaler premium compressing from ~250% to ~120% [9]
Geographic differentialRegional supply, fiber, policyHyperscaler rates falling 26–54% by region while other segments recovered [9]
Power differentialElectricity cost, congestion, reliabilityProjected wholesale power cost increases of 6–29% by 2030; Virginia scenarios to 57% [30]
Latency differentialPhysical distance to workloadTraining tolerates remoteness; real-time inference and agents do not (Section 2.6)
Software differentialCUDA ecosystem switching costsEcosystem lock-in taxes every cross-hardware arbitrage (Section 2.7)
Utilization differentialAchieved vs. installed hoursInstalled GPU-hours ≠ productive GPU-hours (Section 2.8)
Service-level differentialFirm vs. interruptible capacityOne-year committed rates rising while on-demand stayed flat [14]
Regulatory differentialExport controls, sovereignty, lawJurisdictional chip-access regimes create international scarcity premiums (Section 5.2)

Each differential corresponds to one friction in Sections 2.1–2.9 and is, in principle, separately measurable and eventually separately tradable.


The exact econometric methodology can and will evolve—hedonic regression across provider listings, matched-pair comparisons within index constituent data, and eventually the market’s own traded spreads will all contribute—and the early ingredients already exist in the segment-level indexes, regional breakdowns, and generation spreads that benchmark providers publish today [9][11]. The important conceptual contribution does not depend on any particular estimator. It is the recognition that the futures benchmark and the economic value received by an individual user will rarely be identical, that the difference is structured rather than random, and that the structure is information. When Texas basis widens, the market is reporting an electricity story. When the hyperscaler differential compresses—as it did, from roughly 250 percent to 120 percent, during 2025 and 2026—the market is reporting a story about enterprise trust, capacity release, and competitive convergence [9]. Learning to read GPU Basis will be the compute market’s equivalent of learning to read crack spreads and locational marginal prices, and Section 6 will argue it may ultimately matter more than the benchmark itself.


Section 3: Compute Friction Through the Five-Layer AI Economy

The frictions of Section 2 do not float freely; they live at specific layers of the artificial-intelligence production stack, and they propagate between layers with the same directional logic by which feedstock costs propagate through any processing industry. The Five-Layer AI Economy—energy at Layer 1, chips at Layer 2, datacenters at Layer 3, models at Layer 4, applications and agents at Layer 5—provides the map. This section walks the stack from bottom to top, showing where the new financial machinery attaches and how compute-price risk, once financialized, travels.


3.1 Layer 1 — Energy: The Hidden Input Inside Every Compute Contract

The futures contract is denominated in GPU-hours, but beneath every GPU-hour lies electricity, and the transmission mechanism from power markets to compute prices is direct enough to be traced line by line. Regional power scarcity raises the operating cost and constrains the expansion of local datacenters; constrained local supply of rentable accelerators raises local compute prices; and local compute basis widens against the global benchmark even though the global accelerator market has not moved at all. The converse holds symmetrically: a region that adds abundant nuclear, natural-gas, renewable, or battery capacity, or that clears its interconnection queue faster than its neighbors, improves its compute economics and watches its basis narrow. The empirical preconditions for this mechanism are all in place and documented—AI-focused datacenter electricity consumption growing 50 percent in a single year, national datacenter shares of consumption heading toward 6.7 to 12 percent in the United States by 2028, wholesale cost projections up 6 to 29 percent by 2030 with Virginia scenarios reaching 57 percent, and more than $40 billion of capacity-market and transmission costs already attributed to datacenter demand in a single regional grid [25][27][29][30][31]. What the futures market adds is a visible instrument through which these energy facts become compute prices in real time. Compute derivatives therefore become connected, hedge by hedge and arbitrage by arbitrage, to energy derivatives; the sophisticated desk of 2028 will run power and compute books side by side, and the power-adjusted compute price will join the vocabulary of Section 1 as a standard analytical object.


3.2 Layer 2 — Chips: Nvidia as the Emerging Reference Grade

Nvidia’s position in this architecture is unlike anything in the history of commodity benchmarks, because the reference grade of the new market is not a natural substance but a proprietary product line in active development by a single company. The first major compute futures are centered on Nvidia accelerator economics—H100 and B200 at CME, an Nvidia-weighted suite at ICE—at the very moment Nvidia’s data-center franchise has reached a scale that gives the reference grade macroeconomic significance in its own right: $89 billion of data-center revenue in a single quarter, guidance of $108 billion for the next, and a chief executive publicly framing demand in half-trillion-dollar units [16][17]. Jensen Huang’s own characterization of market conditions, delivered with Nvidia’s third-quarter fiscal 2026 results, doubles as a description of why hedging instruments found demand [15]:

“Blackwell sales are off the charts, and cloud GPUs are sold out.”

— Jensen Huang, Founder and CEO, Nvidia [15]

And at the GTC 2026 keynote he quantified the pipeline in terms no commodity producer has ever been able to use about its own product [17]:

“We saw $500B in GPU demand last year for Blackwell and Rubin.”

— Jensen Huang, CEO, Nvidia, GTC 2026 [17]

This creates a provocative structural possibility: Nvidia may come to occupy three roles simultaneously—manufacturer of the physical asset, proprietor of the software platform that determines the asset’s usability, and issuer, in effect, of the reference grade against which the financial market prices compute. No oil major ever defined the chemistry of Brent; no miner ever controlled the specification of COMEX gold. The closest analogy is not a commodity producer at all but a central bank, an institution whose liabilities become the unit of account for an entire economy—which is why the broader argument that Nvidia is becoming a Central Bank of AI, setting the terms on which computational liquidity enters the system, gains rather than loses force from the arrival of futures denominated in its hardware. Section 4 examines the reinforcing loops this status creates and the custom-silicon challenge that could eventually fragment it.


3.3 Layer 3 — Datacenters: From Megawatts to Financially Hedgeable Compute Factories

Datacenters are the refineries of this economy: they transform electricity and hardware into sellable computational capacity, and the moment their output acquires a forward curve, their financing begins migrating toward the project-finance template that transformed power generation a generation ago. A developer underwriting a campus in 2027 will be able to build a model whose every major line references an observable market: future GPU rental rates from the forward curve, expected utilization from fleet telemetry benchmarks, regional basis from the emerging locational data, hardware depreciation from the traded generation spreads, hedge costs from the options market that will follow the futures, tenant credit from conventional analysis, electricity from the power curve, and replacement cycles from Nvidia’s published roadmap. This is precisely how commodity-processing infrastructure is financed everywhere else, and the financial system is not waiting for the futures to launch before moving in this direction. Incremental annual debt across the five major hyperscalers rose from 9 percent of capital expenditure in fiscal 2024 to 32 percent on a trailing basis by mid-2026 as the buildout outran even the world’s deepest corporate cash flows, and the financing mix now spans bonds, leases, tenant prepayments, and partnership capital [22]. One layer down, the neocloud sector already borrows directly against compute: CoreWeave alone raised more than $30 billion of debt and equity in 2026, much of it GPU-collateralized or secured by a small number of hyperscaler contracts, including a single $8.5 billion facility backed by a Meta contract worth up to $14.2 billion [36]. Every one of those structures embeds an implicit view on future compute prices; the futures market simply makes the view explicit, hedgeable, and—for the first time—auditable against a market settlement. AI datacenters are becoming commodity-processing infrastructure from a financing perspective, and the forward curve is the document that completes the conversion.


3.4 Layer 4 — Models: Compute Price Becomes a Model-Economics Variable

Frontier-model laboratories consume compute on a scale that makes its price a strategic variable rather than a cost line, and a liquid futures market changes their decision architecture in a specific way: it allows them to separate technological decisions from a portion of price risk. Today, the decision to train a next-generation model bundles an engineering bet with an unhedged commodity bet—the laboratory commits to consuming an enormous quantity of compute at whatever prices prevail across the training window, with documented spot volatility of ten percent in a month and forty percent in five months [8][9]. With a forward curve, the two bets can be unbundled: the engineering schedule proceeds on its own logic while the treasury desk locks the compute leg, exactly as an airline separates route planning from fuel-price exposure. Model economics then acquires a canonical expression—tokens generated multiplied by revenue per token, minus compute cost, minus hedge cost—in which each term is observable and the middle term can be fixed in advance. The strategic consequence cuts both ways. Hedged laboratories gain budget certainty and become more attractive to lenders and investors, lowering their cost of capital; but the same forward curve disciplines them, because a curve in steep contango is the market publicly pricing the scarcity of their principal input, and burn-rate projections that once lived in private spreadsheets become checkable against a public number. Compute financialization, in other words, does not merely serve the model layer. It audits it.


3.5 Layer 5 — Applications and Agents: Compute Risk Travels Downstream

The largest long-run effect may materialize at the top of the stack, farthest from the trading screens. As autonomous agents execute growing volumes of inference on behalf of ordinary businesses—handling customer interactions, monitoring supply chains, writing and testing software, negotiating routine transactions—those businesses acquire persistent, structural compute exposure of exactly the kind transportation firms have to fuel. A corporation running millions of agentic tasks per day will discover that its operating margin moves with the inference price, and it will respond the way every fuel-exposed industry eventually responded: first with contractual pass-throughs, then with procurement hedging, and finally with financial hedging through the very instruments this paper describes. The IEA’s observation that per-task energy consumption is falling by an order of magnitude annually while aggregate demand nonetheless compounds captures the dynamic perfectly—efficiency gains at the task level are being overwhelmed by volume growth at the deployment level, which is the classic signature of a technology becoming infrastructure [26]. When that happens, a financial innovation that began with two contracts on NYMEX referencing Nvidia rental rates will have reached the operating statements of companies that have never owned a GPU and never will. Compute risk, like energy risk before it, will have become a general condition of doing business—and the depth of the futures market will grow accordingly, because the deepest hedging markets are always the ones whose underlying touches everyone.


Section 4: The Financialization of Artificial Intelligence

Financialization is a word with a history, and the history counsels neither celebration nor panic. The academic literature on the financialization of commodities—Cheng and Xiong’s survey work, Tang and Xiong on index investment, Basak and Pavlova’s equilibrium modeling, Singleton on investor flows in the 2008 oil cycle—documents a consistent pattern: the arrival of financial capital in a physical market improves price discovery and risk transfer while simultaneously importing new correlations, new leverage, and new channels of contagion, and recent theoretical work has begun extending precisely this apparatus to compute and even to inference-token derivatives [42]. This section examines what that pattern implies when the underlying is the productive substrate of machine intelligence, taking in turn the forward curve as macroeconomic signal, the mechanics of compute arbitrage, the collateral revolution in infrastructure lending, the unresolved question of what the benchmark actually prices, Nvidia’s benchmark advantage, and the custom-silicon challenge to it.


4.1 Compute Forward Curves as a New Signal of AI Expectations

If the CME and ICE contracts achieve liquidity across multiple listed months, the world will acquire something it has conspicuously lacked through three years of furious debate about artificial intelligence: a continuously traded, money-backed, forward-looking consensus on the price of the technology’s binding input. The shape of that curve will be read the way oil and rates curves are read. A steep upward slope—contango, in the market’s vocabulary—signals expected scarcity: demand compounding faster than datacenters can energize, exactly the condition Nvidia’s sold-out order books and Amazon’s capacity constraints through 2027 describe today [15][23]. A flattening or inverted curve signals expected relief: hardware generations arriving on schedule, efficiency gains compounding, or—the possibility that hangs over the entire buildout—demand disappointing the capacity constructed for it. The curve will be quoted in earnings calls, cited in Federal Reserve financial-stability reports, and weaponized in the bubble debate by both sides, because it converts a rhetorical question—is AI demand real?—into a number that clears against actual positions. It bears remembering how much money now rides on the answer: nearly $700 billion of hyperscaler capital expenditure in 2026 against free cash flows already under visible pressure, an investment wave large enough that the International Monetary Fund credits it with offsetting a global oil shock while simultaneously flagging its concentration as a financial-stability concern [19][22][37]. Kristalina Georgieva, the Fund’s Managing Director, captured the macroeconomic promotion of this spending in August 2026 [37]:

“Is now becoming a growth engine for the global economy.”

— Kristalina Georgieva, Managing Director, IMF, on AI investment, August 2026 [37]

A forward curve for compute will be the closest thing the world gets to a daily referendum on whether that engine is running on demand or on momentum.


4.2 The Arrival of Compute Arbitrage

Price disparities of the magnitude documented in Section 2—fivefold across providers on a routine day, twenty-one-fold at the extremes—create arbitrage incentives that a financialized market will organize capital to pursue [8][10]. But compute arbitrage cannot work the way commodity arbitrage works, because the commodity cannot be transported; the arbitrage must instead move everything around the commodity. Workload migration moves the demand to the cheap supply, within the hard limits that latency, data gravity, and software friction impose. Cloud switching moves the customer relationship, paying the CUDA-ecosystem conversion tax along the way. Capacity brokerage—the business model of the marketplaces whose listings define the wide end of the price spread—moves information, matching stranded supply with unserved demand. Financial hedging moves only the risk, which is frequently all a participant actually needs moved. Regional datacenter development moves future supply toward cheap power and favorable regulation, arbitraging the geographic basis on a five-year horizon. Software adaptation—porting stacks to alternative accelerators—arbitrages the hardware basis on a similar horizon. Each channel is real; each is slow, costly, or capped; and the aggregate limit they impose is precisely why the price dispersion persists at all. Compute Friction, in the framework of this paper, is the quantitative answer to the question of how much of the apparent arbitrage can actually be captured—and the portion that cannot be captured is not waste. It is the equilibrium price of physics, law, and switching costs, and it will be traded as basis rather than harvested as profit.


4.3 GPU Collateral and Infrastructure Lending

Nowhere will the forward curve matter more concretely than in credit. The sector has already constructed a multi-hundred-billion-dollar lending complex on top of accelerator collateral and compute contracts—CoreWeave’s $30 billion-plus of 2026 raises and its $8.5 billion Meta-backed facility are only the most visible instances—and it has done so without any market-based mechanism for valuing the collateral’s future earning power [36]. Lenders currently underwrite against manufacturer roadmaps, sponsor projections, and the credit of a small number of anchor tenants; the concentration is severe enough that a single contract renewal can be the difference between an asset-backed loan and an unsecured one [36]. A credible forward rental curve changes the entire practice: loan-to-value ratios can key to market-implied revenue rather than appraisal fictions; collateral haircuts can widen mechanically as the curve prices generational decay; covenant packages can require hedging of projected GPU revenue exactly as power-project lenders require merchant-price hedges; and securitization—the packaging of long-dated GPU revenue streams into rated paper—becomes feasible at scale once the underlying cash flows can be marked against a public settlement. The stabilizing potential is genuine, and so is the destabilizing one, because the same curve that enables prudent lending enables leveraged speculation on it, and the collateral in question sits at the center of the fiercest valuation dispute in the market. Hyperscalers depreciate AI hardware over five to six years; Nvidia now ships new architectures annually; Michael Burry has publicly estimated that the gap overstates industry earnings by roughly $176 billion across 2026 through 2028, and the resale value of the collateral class he is describing can fall by most of its purchase price within three to four years [34][35][36]. Microsoft’s own chief executive, Satya Nadella, conceded the underlying anxiety in a single unguarded sentence about pacing chip purchases [34]:

“Stuck with four or five years of depreciation on one generation.”

— Satya Nadella, CEO, Microsoft, on what he sought to avoid [34]

A traded forward curve will not settle the depreciation war, but it will do something arguably more consequential: it will price it, continuously and in public, because the spread between near-dated and far-dated rental futures on an aging chip generation is the market’s real-time estimate of exactly the economic decay the accountants are arguing about. Financialization, here as everywhere, lowers some risks while manufacturing new ones—less appraisal fiction, more mark-to-market volatility; fewer unpriced assumptions, more correlated deleveraging when the curve gaps down. The net is genuinely uncertain, which is itself a reason regulators should be paying attention now rather than after the first credit event.


4.4 Hyperscaler Versus New Cloud Economics: What Exactly Is the Futures Market Pricing?

The enormous, persistent, and only partially converging spread between hyperscaler and neocloud GPU pricing forces a question that every user of these contracts must answer before trading them: what is the benchmark a price of? Raw silicon-hours? Usable accelerator time within a managed environment? A particular provider class? Enterprise-grade firm capacity, or marginal interruptible surplus? The honest answer is that the benchmark prices whatever its methodology says it prices—a normalized composite in Silicon Data’s case, a transaction-weighted clearing level in Ornn’s—and that both composites sit somewhere in the wide space between a $2.53 neocloud index print and hyperscaler list rates several times higher [5][10][11]. The consequence is structural and permanent: a benchmark based heavily on one market segment will represent the other segments poorly, and every hedger whose physical exposure lives in a poorly represented segment will carry basis risk in proportion to the distance. An enterprise paying hyperscaler premiums for compliance and integration hedges its bill only approximately with a new cloud-weighted future; a marketplace seller of interruptible capacity hedges only approximately with a composite that embeds reserved-term pricing. None of this is a flaw to be fixed. It is the ordinary condition of every benchmark-anchored physical market on earth, and the mature response—already visible in the segment-level index publication that providers have adopted—is a family of published differentials that let each participant locate its own exposure relative to the traded center [11].


4.5 Nvidia’s Benchmark Advantage

If markets increasingly quote artificial-intelligence compute economics in Nvidia hardware units, Nvidia benefits even from transactions it never touches, through a set of reinforcing loops familiar from every dominant standard in economic history. Developers optimize for the benchmark hardware because that is where tooling, talent, and documentation concentrate. Lenders prefer collateral with transparent reference pricing, so Nvidia-based fleets borrow more cheaply than fleets of any rival accelerator, which tilts every procurement decision at the margin. Datacenter operators install the hardware with the deepest secondary markets, because exit value is part of the underwriting. Customers prefer infrastructure whose costs can be financially hedged, and the only listed hedges reference Nvidia silicon. Traders concentrate liquidity in the contracts already trading, which raises the cost of launching any competing benchmark. Each loop strengthens the others, and together they convert an engineering lead into a piece of financial-market infrastructure—a moat made not of transistors but of open interest. The benchmark, in short, strengthens the ecosystem that made the benchmark possible. This is the strongest form of the Central Bank of AI thesis, and it carries the corresponding governance question: when a single company’s product roadmap can move the settlement value of a regulated futures complex—when a surprise architecture announcement is, functionally, a rate decision—the line between corporate communication and market-moving policy becomes something exchanges, the CFTC, and Nvidia itself will have to think about with a seriousness the sector has not yet displayed [1][5].


4.6 The Custom-Silicon Challenge

The countervailing force is already in the field. Google’s TPUs, Amazon’s Trainium, Microsoft’s Maia, Meta’s internal accelerators, OpenAI’s custom-silicon program, and the Chinese accelerator families maturing behind the export-control wall collectively represent the largest coordinated attempt in semiconductor history to escape a single vendor’s economics—and the hyperscalers pursuing it control both the largest compute fleets and the datasets on which any rival benchmark would be built. If custom architectures capture a substantial share of deployed capacity, the compute market fragments into multiple benchmark families, with cross-hardware spreads becoming the market’s continuous referendum on ecosystem competition; a TPU-hour trading at a widening discount to the H100 strip would say more about CUDA’s moat than any analyst report. ICE’s contract design, which contemplates adding GPU types as the market develops, leaves the institutional door open to exactly this evolution [5]. The plausible end state is therefore neither a single global benchmark on the Brent model nor a chaotic scatter, but something closer to the world’s interconnected energy markets: numerous fuels, grades, locations, and hubs, tied together by a dense web of traded basis relationships, with one or two contracts serving as the liquidity anchors against which everything else is quoted. In that world the analytical unit of this paper—the basis relationship—is not a complication of the market. It is the market.


Section 5: Policy, Regulation, Geopolitics, and the 2027–2030 Compute Market

Every previous financialization of a strategic input eventually summoned a regulatory architecture—position limits and benchmark oversight in energy, macroprudential machinery in housing finance, an entire supervisory college for reference rates after LIBOR. Compute will be no different, and the distinctive feature of this cycle is that the policy questions are arriving before the first contract has even settled. This section maps the terrain across five fronts—derivatives regulation, export controls, state-level price formation, grid congestion, and benchmark governance—and closes with a preparatory agenda for the years before 2030.


5.1 Who Regulates the Price of Intelligence?

As compute derivatives expand, regulators will confront the full standard inventory of derivatives-market questions—benchmark methodology and auditability, manipulation surveillance, market concentration, conflicts of interest between index providers and affiliated trading firms, clearing risk, position limits, disclosure regimes, settlement integrity, and, eventually, systemic importance. Some of these questions carry unusual weight here. The index providers at the center of the market are young companies with commercial relationships throughout the industry they measure; Silicon Data is backed by DRW, one of the world’s largest proprietary trading firms, a fact disclosed prominently in the launch materials and one that benchmark-governance frameworks exist precisely to manage [1]. The underlying market is opaque, bilateral, and dominated by a handful of sellers whose own reporting choices feed the indexes. And the notional exposures that could eventually reference these settlements—datacenter debt, GPU-backed credit, securitized compute revenue—are growing at a pace that suggests systemic relevance within years, not decades [22][36]. The deeper point is architectural: artificial-intelligence policy and derivatives regulation are converging on the same object, and the two policy communities barely speak. Model-governance debates absorb enormous public attention; the question of who audits the price of intelligence has received almost none. The IMF’s institutional warnings about AI amplifying financial-stability risk—articulated by Gita Gopinath well before the futures were announced—apply with special force to a channel in which AI exposure becomes directly tradable leverage [38]:

“A real need to have parallel effort to make sure that we’re also AI-proofing the global economy.”

— Gita Gopinath, First Deputy Managing Director, IMF [38]


5.2 Export Controls Create International Compute Basis

United States controls on advanced accelerators partition the world into jurisdictions of differential chip access, and the futures market will convert that partition from a geopolitical abstraction into a printed price. Restricted regions experience scarcity premiums; freely supplied regions enjoy lower effective compute costs; and the Chinese accelerator ecosystem that controls helped call into existence operates a parallel market whose implied prices will be compared, continuously, against the dollar benchmark. The result is not merely geopolitical but financial: export policy becomes a determinant of global compute basis, a lever whose effects can be read off screens in a way sanctions architects have never before enjoyed—or endured, since visible price gaps also measure evasion incentives with uncomfortable precision. The policy machinery is live and volatile: the past two years have seen license regimes revised repeatedly, revenue-sharing arrangements floated for China-bound accelerators, and Chinese domestic alternatives improving quickly enough that commentary now speaks of genuine competitive pressure on the reference-grade vendor’s position in that market. A compute forward curve segmented by jurisdiction—when it eventually exists—will be one of the more consequential geopolitical dashboards of the 2030s, and its architects should expect the same scrutiny that attends any instrument capable of pricing statecraft.


5.3 Governors Become Participants in Compute Price Formation

State policy shapes the effective price of compute through channels that are individually mundane and collectively decisive: electricity regulation and utility rate design, tax incentives and their increasingly contested repeal, datacenter permitting, transmission siting, nuclear and natural-gas development, water regulation, property taxation, zoning, and the community consent that polling suggests is eroding rapidly [28][30]. The fiscal stakes are documented state by state—Virginia has forgone an estimated $1.6 billion in exemptions, Georgia expects $2.5 billion, and legislators in several states are moving to claw such preferences back as household electricity anger sharpens into a voting issue [31]. A governor may never trade a GPU futures contract, yet decisions made in Austin, Richmond, Harrisburg, Indianapolis, Lansing, Phoenix, Columbus, or Sacramento alter the regional cost structures underlying those contracts, and once regional compute differentials are published and traded, the alteration will be visible—attributable, quotable, and politically consequential—within days of the policy change rather than years. States, in other words, are about to acquire compute-price report cards whether they want them or not, and the sophisticated ones will begin treating compute basis the way they already treat bond spreads: as a market verdict on governance.


5.4 Grid Congestion Becomes Compute Congestion

Regions able to deliver large increments of power, transmission, and fast interconnection will attract computational load; regions that cannot will develop scarcity premiums; and the compute market will thereby begin transmitting information about infrastructure bottlenecks that electricity markets themselves capture only partially. Locational power prices reveal congestion for load that already exists; compute basis will reveal congestion for load that wants to exist—the datacenter that was not built, the campus that migrated to a different interconnection queue—which is exactly the forward-looking signal that grid planners chronically lack. The feedback also runs in reverse, and it runs through household bills: the $6.3 billion of consumer costs PJM attributes to datacenter demand, the record rate-increase requests, and the 78 percent of Americans worried about datacenters raising their bills together define the political constraint surface on which the entire buildout now operates [28][29][31]. A compute market that makes regional scarcity visible will, in the best case, direct investment toward genuinely spare capacity and away from saturated corridors—performing the allocative service price signals exist to perform—and, in the worst case, industrialize the race for whatever cheap power remains ahead of the communities that live beside it. Which case obtains is substantially a policy choice, which is why it belongs in this section rather than in the market-design one.


5.5 Benchmark Governance Becomes Industrial Policy

If H100 and B200 settlement prices become major reference points for global compute, then the methodological minutiae of index construction—which providers enter the sample, which geographies, which contract durations, which network configurations, which service tiers, how outliers are filtered, how the new cloud and hyperscaler segments are weighted—acquire strategic importance of a kind normally associated with trade policy. These choices can shift billions of dollars of financial exposure; they can flatter or punish entire provider classes; they can make a nation’s compute look cheap or scarce to the global capital that reads the number. The post-LIBOR benchmark-governance regime—IOSCO principles, administrator accountability, methodology transparency, conflict management—provides the template, and the case for applying it here early is stronger than it was for rates, because the compute benchmarks are being born directly into futures settlement rather than growing informally for decades first [1][5]. Benchmark governance, in short, deserves scrutiny comparable to other economically significant reference rates, and it deserves it now, while the methodologies are young enough to be shaped rather than merely grandfathered.


5.6 What Policymakers Should Prepare for Before 2030

Federal and state policymakers should begin treating compute as a financialized infrastructure market rather than merely a technology industry, and the preparatory agenda follows directly from the frictions this paper has catalogued. First, monitor regional compute-price disparities as standing economic indicators, with the same institutional seriousness applied to regional power prices. Second, build the analytical capacity to trace how power shortages transmit into artificial-intelligence prices, because the transmission is now mechanical and fast. Third, evaluate concentration in benchmark hardware, recognizing that a single vendor’s roadmap functions as quasi-monetary policy for the sector. Fourth, examine the interaction between export controls and compute markets, including the price signatures of circumvention. Fifth, protect benchmark integrity under IOSCO-grade governance before, not after, the first manipulation scandal. Sixth, assess the systemic consequences of GPU-backed lending while the collateral class is still young enough to stress-test honestly, with the depreciation dispute treated as the material accounting question it is [34][35][36]. Seventh, prepare the regulatory perimeter for derivatives on additional accelerator architectures and, eventually, on inference-token and output-based indexes [42]. Eighth, develop tools for distinguishing speculative compute demand from productive demand, since forward curves will aggregate both. Ninth, study whether financial hedging genuinely accelerates datacenter investment—the industry’s central promise for these products—or primarily redistributes its rents. And tenth, incorporate compute-price indicators into national artificial-intelligence capacity planning, because a government that cannot read the compute curve in 2030 will be planning industrial strategy with one eye closed.


Section 6: What Have We Learned? Seven Pillars

Long arguments deserve compression at the end, but compression of the right kind: not a shorter restatement, but a set of load-bearing conclusions that can stand on their own when the supporting scaffolding is removed. What follows are seven such pillars—the original five of this framework, joined by two that the evidence assembled above now makes unavoidable.


Pillar 1: Compute Can Become Financially Standardized Without Becoming Physically Fungible

The first lesson is the most important, and everything else in this paper is commentary upon it. The creation of a futures contract does not transform GPU-hours into identical commodities, and no refinement of contract design ever will, because the sources of heterogeneity—topology, electricity, geography, latency, software, service quality, law—are not specification errors but properties of the physical world. Financial markets require standardized references; artificial-intelligence infrastructure remains heterogeneous; both statements are true simultaneously and will remain so. CME and ICE can build increasingly sophisticated price-discovery machinery, with contract families spanning hardware generations and eventually geographies, while enormous differences persist across accelerator generations, providers, datacenters, power markets, networks, jurisdictions, workloads, and service agreements [1][5][9][12]. The persistent gap between the financial abstraction and the delivered reality is Compute Friction, and recognizing it as permanent rather than transitional is the precondition for using these markets intelligently instead of being used by them.


Pillar 2: GPU Basis May Become as Important as the Benchmark Itself

A reference price is valuable precisely because real prices differ from it; a benchmark nobody deviates from would be measuring nothing. The durable analytical opportunity therefore lies not in asking what the H100 contract trades for but in measuring the structured distance between that settlement and effective compute delivered somewhere specific. Texas will carry one basis and Northern Virginia another; Europe, the Gulf, and East Asia others still; a hyperscaler will carry one premium, a new cloud another, a tightly interconnected flagship cluster a third—and the early data already trace these surfaces, from the 120-to-250-percent hyperscaler premium band to the swinging H200-over-H100 generation spread to the segment-split index publications that acknowledge the composite is not one price [9][11][14]. The benchmark gives the market a center. GPU Basis describes the distance from that center, and in mature commodity markets the basis book is where the real information—and frequently the real money—lives.


Pillar 3: Location and Electricity Are Becoming Properties of Compute

The artificial-intelligence industry has long spoken of compute as abstract processing power, an infinitely divisible cloud resource summoned by API call. The economics of 2025 and 2026 have ended that manner of speaking. Compute has geography; it consumes electricity at a scale that now moves national demand statistics; it occupies buildings that communities increasingly contest; it travels across fiber with a latency that prices distance; it encounters regulators, tax codes, water constraints, and interconnection queues [24][25][27][28]. Every GPU-hour is embedded inside Layer 1 energy systems, Layer 2 silicon supply chains, and Layer 3 physical facilities before it ever reaches a Layer 4 model or a Layer 5 agent, and the futures market—far from abstracting these dependencies away—will publish them daily as basis. The Five-Layer AI Economy is becoming financially interconnected from bottom to top, and the direction of causation runs both ways: power reprices compute, and compute demand now reprices power, all the way down to the household bill [29][31].


Pillar 4: Compute Derivatives Could Reorganize AI Infrastructure Finance

The significance of compute futures will ultimately be measured far from the trading floor, in the underwriting standards of the capital that builds the machine. Forward prices can inform datacenter valuations, GPU-backed loan covenants, cloud contract structures, accelerator purchasing decisions, laboratory budgets, and enterprise deployment economics; a curve converts an uncertain revenue stream into something investors believe they can model and partially hedge, and that belief—warranted or not—unlocks capital. The sector’s financing arc makes the stakes concrete: hyperscaler debt rising from 9 to 32 percent of capital expenditure in eighteen months, single GPU-collateralized facilities in the billions secured by individual tenant contracts, and an unresolved depreciation dispute measured in the hundreds of billions sitting underneath the collateral values [22][35][36]. Financial machinery of this sophistication can stabilize the buildout by distributing its risks to willing holders, or it can synchronize the buildout’s failure modes by marking everyone to the same curve on the same bad day. Both outcomes have precedents. Which one compute gets will depend substantially on choices—covenant design, haircut discipline, benchmark governance—being made right now, mostly out of public view.


Pillar 5: Compute Friction Is Not Merely an Inefficiency; It Is Strategic Information

It is tempting to assume markets will grind Compute Friction away, and some components will indeed decline: software will grow more portable, networks will improve, price transparency will spread, and segments of cloud pricing have already converged dramatically from their 2024 extremes [9]. But the remainder—geography, electricity, latency, sovereignty, regulation, hardware architecture, service quality—cannot be homogenized, and the correct response is to stop reading the residual friction as failure and start reading it as signal. A widening regional basis reveals energy scarcity before the utility commission convenes. A hardware premium reveals technological preference before the analyst decks are printed. A cloud premium reveals what enterprises actually pay for trust. A latency premium reveals where the application economy is physically anchoring itself. A sovereign premium prices geopolitical risk in dollars per GPU-hour. Compute Friction, systematically observed, becomes one of the primary instruments through which the artificial-intelligence economy reports the location of its real bottlenecks—and an economy that can see its bottlenecks is an economy that can, at least in principle, fix them.


Pillar 6: The Forward Curve Will Discipline the Boom—or Expose It

A sixth conclusion follows from joining the financial evidence of Section 4 to the macroeconomic setting. The buildout is now large enough to register in global growth accounting—the IMF credits AI investment with offsetting an oil shock—while its skeptics, from a Nobel-laureate economist projecting roughly a one-percent decadal GDP contribution and warning that the consumer-application layer needed to justify the capacity remains missing [39] to a short-seller alleging $176 billion of understated depreciation, argue that the capacity is running far ahead of the demand that must eventually pay for it [35][37][40]. Daron Acemoglu’s framing of the underlying uncertainty is the appropriately measured one [40]:

“It’s not that you cannot get big productivity gains from automation.”

— Daron Acemoglu, Institute Professor, MIT, Nobel Laureate in Economics [40]

—it is that such gains are harder and slower than the spending assumes. Until now this dispute has been conducted through equity prices, which bundle compute expectations with everything else a company is. A liquid compute curve unbundles the question. If demand is real, the curve will stay firm through the largest capacity additions in computing history, and the bulls will have a daily settlement to point to. If demand disappoints, the front of the curve will crack before the earnings do, and the curve will have served as the early-warning instrument this cycle otherwise lacks. Either way, the era in which the central quantitative claim of the AI boom could remain untested by a traded price is ending—and that, more than any hedging convenience, may be the futures market’s deepest contribution.


Pillar 7: The Unit of Account Will Outlive the Hardware That Defined It

Finally, the framework must survive its own examples. The H100 that anchors the first contract is already midway through its commercial life, trading at a widening discount to its successors and destined for the legacy-pricing decay curve its predecessors have traced [14]. The B200 will follow; Rubin-generation accelerators, alternative architectures, and measures that do not yet exist—inference-token indexes, agentic-compute indexes, utilization-adjusted output benchmarks already prefigured in the academic literature—will succeed them [42]. What persists is the institutional form: a benchmark, a curve, a basis surface, and a market that prices the distance between standardized intelligence-capacity and its delivered reality. Compute Friction names that permanent structure. The hardware in the examples will date within thirty-six months. The framework is built so that nothing else in this paper will.


Conclusion: Why “Compute Friction” Fits the Emerging AI Economy

The development scheduled for October 5, 2026 may eventually be remembered as far more than the introduction of another derivatives product. If compute futures become liquid and widely referenced, artificial-intelligence infrastructure will have crossed an institutional boundary of the kind that only a handful of resources ever cross. Computing power will no longer exist solely as servers purchased by technology companies or capacity rented through cloud dashboards; it will exist simultaneously as an exposure that financial institutions can price forward, hedge, trade, finance, securitize, and incorporate into investment decisions across the economy. CME’s planned H100 and B200 contracts, ICE’s competing Ornn-based suite, the perpetual-futures initiatives behind them, the ETF filings that appeared before a single contract traded, and a White House policy document that explicitly calls for a financial market in compute all indicate that this transformation is not hypothetical [1][5][10][41]. It has already begun.

But the transformation contains a contradiction, and this paper has argued that the contradiction is the story. The financial system wants a GPU-hour to behave like a barrel of oil. The physical artificial-intelligence economy refuses. An H100 offered through one cloud is not economically interchangeable with an H100 offered through another when the observed spread between them is fivefold on an ordinary day and twenty-one-fold at the extremes [8][10]. An accelerator in Texas is not automatically equivalent to one in Northern Virginia when the grids beneath them are diverging in cost, congestion, and political tolerance by double-digit percentages [29][30]. A processor woven into a flagship cluster by high-bandwidth optical fabric is not equivalent to an isolated processor of identical manufacture. A machine running on abundant contracted power does not share the economics of one operating behind a constrained grid whose regulator faces voters holding record utility bills [28][31]. A GPU thousands of miles from an inference workload cannot substitute for one adjacent to it, because light declines to hurry. An accelerator without the appropriate software environment is worth less to a particular customer than a slower chip inside the right ecosystem. These differences do not disappear when a futures contract lists. They become measurable against it—and measurement, not elimination, is what the contract is actually for.

That is precisely why the title Compute Friction fits. The term names the resistance encountered where the financial abstraction of standardized compute meets the physical reality of artificial-intelligence infrastructure, and it insists that the resistance is structural, informative, and permanent rather than a defect awaiting liquidity. Every claim in this paper reduces to a way of reading that resistance: as basis to be measured, as risk to be allocated, as collateral to be haircut, as policy to be scrutinized, as signal to be mined.

The resistance will matter more, not less, as the economy grows. Today the reference instrument is an H100 GPU-hour; tomorrow it may be a B200-hour, a Rubin-generation accelerator-hour, a TPU-hour, a Trainium-hour, an inference-token index, an agentic-compute index, or some future measure of machine intelligence that does not yet have a name. The hardware will change, the benchmarks will change, the exchanges will change, the models will change. The underlying interrogation will not: Where is the compute? What chip produces it? What electricity powers it? What network connects it? What software can use it? How reliably can it be delivered? How much latency does it carry? What jurisdiction controls it? And how much is that particular unit of compute actually worth relative to the financial benchmark? Those questions make GPU Basis a permanently useful measurement concept. Together, they constitute Compute Friction.

The emergence of compute futures therefore does not mean that artificial-intelligence compute has finally become a commodity. It means something more interesting: compute is becoming a financial market while remaining a fragmented physical infrastructure system, and the tension between the clean price on the trading screen and the messy reality of chips, clouds, datacenters, grids, networks, geography, latency, regulation, and software is where the next generation of artificial-intelligence infrastructure economics will be written. The exchanges have built the screen. The physical world retains the veto. The distance between them now has a name, will soon have a price, and deserves—this paper has argued—a literature.

And that is why Compute Friction fits this paper so well.


Footnotes / Endnotes:

[1] CME Group & Silicon Data, “CME Group and Silicon Data to Launch Compute Futures on October 5 to Unlock New Way to Hedge AI Risks,” Press Release, August 11, 2026. https://www.cmegroup.com/media-room/press-releases/2026/8/11/cme_group_and_silicondatatolaunchcomputefuturesonoctober5tounloc.html

[2] Terry Duffy (Chairman & CEO, CME Group), in CME Group & Silicon Data, “CME Group and Silicon Data Partner to Launch First Compute Futures,” Press Release, May 12, 2026. https://www.cmegroup.com/media-room/press-releases/2026/5/12/cme_group_and_silicondatapartnertolaunchfirstcomputefutures.html

[3] Carmen Li (CEO, Silicon Data), quoted in CNBC, “AI Computing Power Is Becoming a Tradable Asset Class as CME Launches Futures Contracts,” August 11, 2026. https://www.cnbc.com/2026/08/11/ai-computing-power-becomes-a-tradable-asset-class-as-cme-starts-futures.html

[4] Trabue Bland (SVP of Futures Markets, ICE), in Intercontinental Exchange & Ornn, “ICE and Ornn to Launch GPU Compute Futures Contracts,” Press Release, May 19, 2026. https://ir.theice.com/press/news-details/2026/ICE-and-Ornn-to-Launch-GPU-Compute-Futures-Contracts/default.aspx

[5] Yahoo Finance, “ICE and Ornn to Launch GPU Compute Futures Contracts,” May 19, 2026 (OCPI coverage of H100, H200, B200, RTX 5090; cash-settled, USD-denominated, transaction-based). https://finance.yahoo.com/markets/options/articles/ice-ornn-launch-gpu-compute-142609556.html

[6] Siôn Geschwindt, The Next Web, “NYSE Owner ICE Plans GPU Compute Futures with Index Partner Ornn,” May 19, 2026. https://thenextweb.com/news/ice-nyse-compute-futures-market-gpu-ai

[7] CME Group, “Initial Listing of Two (2) Compute Futures Contracts — Silicon Data H100 Rental Index Futures and Silicon Data B200 Rental Index Futures,” Special Executive Report SER-9785, August 2026. https://www.cmegroup.com/notices/ser/2026/08/ser-9785.html

[8] Spheron Network, “Compute Futures: What CME’s GPU Contracts Mean for AI Buyers,” August 2026 (SemiAnalysis one-year H100 contract data; August 2026 on-demand spread of $2.19–$11.06/hr). https://www.spheron.network/blog/compute-futures-cme-gpu-contracts-ai-buyers/

[9] Silicon Data, “GPU Futures: How Compute Is Becoming a Tradable Commodity,” July 22, 2026 (December 2025–January 2026 H100 spike; hyperscaler premium compression from ~250% to ~120%; H200/H100 spread dynamics; basis-risk analysis). https://www.silicondata.com/blog/gpu-futures

[10] FinStrat Management, “GPU Compute Futures: Price Discovery or a New Layer of Risk,” September 2026 ($0.72–$15.14 intraday spread across 24 marketplaces; index methodology comparison; Architect Financial Technologies perpetual futures). https://finstratmgmt.com/the-frontier/gpu-compute-futures-price-discovery-or-a-new-layer-of-risk/

[11] Silicon Data, “H100 Rental Price Index” (SDH100RT), accessed September 2026 (standardization for term, cluster scale, interconnect; separate neo-cloud and hyperscaler readings). https://www.silicondata.com/products/silicon-index/h100

[12] IntuitionLabs, “Data Center GPU Pricing 2026: The Full AI Pricing Index,” July 20, 2026 ($1.38–$12.29/GPU-hour observed band). https://intuitionlabs.ai/articles/data-center-gpu-pricing-2026

[13] GPUSmith, “Nvidia H100 Rental Price History and H200 Cost Trends 2026,” July 2026 (H100 purchase price $25,000–$40,000; ~8x on-demand range; hyperscaler customer-base analysis). https://gpusmith.com/articles/en/h100-rental-price-history-trends

[14] AIMultiple, “Cloud GPU Rental Price Index,” September 2026 (newest-generation pricing doubling; legacy decay $1.77→$1.14; hyperscalers at 3–5x neocloud floors in first year; 1-year committed vs. on-demand divergence). https://aimultiple.com/gpu-index

[15] NVIDIA Corporation, “NVIDIA Announces Financial Results for Third Quarter Fiscal 2026,” Press Release, November 2025 (Jensen Huang statement). https://nvidianews.nvidia.com/news/nvidia-announces-financial-results-for-third-quarter-fiscal-2026

[16] Yahoo Finance, “3 Semiconductor Stocks to Buy Heading Into September,” August 2026 (NVIDIA Q2 FY2027: $96B revenue, $89B data center, Q3 guidance $108B ±2%). https://finance.yahoo.com/technology/articles/3-semiconductor-stocks-buy-heading-110101367.html

[17] Seeking Alpha, “Nvidia CEO Huang Doubles Down, Eyes $1T in Product Demand Through 2027: GTC,” March 16, 2026. https://seekingalpha.com/news/4564887-nvidia-ceo-huang-doubles-down-eyes-1t-in-product-demand-through-2027-gtc

[18] Futurum Group, “AI Capex 2026: The $690B Infrastructure Sprint,” February 12, 2026. https://futurumgroup.com/insights/ai-capex-2026-the-690b-infrastructure-sprint/

[19] Statista, “Big Tech’s AI Spending to Reach $760 Billion in 2026,” July 31, 2026. https://www.statista.com/chart/35046/capital-expenditure-of-meta-alphabet-amazon-and-microsoft/

[20] I/O Fund, “AI Capex to Hit $1 Trillion – And Estimates Are Still Too Low,” August 5, 2026 (H1 2026 capex $301B; $732.5B 2026 guidance; Goldman Sachs $7.6T 2026–2031 baseline). https://io-fund.com/ai-stocks/ai-capex-1-trillion-estimates-too-low

[21] Brian Sozzi, Yahoo Finance, “Meta, Microsoft, Amazon, and Alphabet Are About to Spend a Shocking Amount of Money to Dominate the AI Era,” June 3, 2026 (Goldman Sachs $5.3T FY2025–FY2030; $725B 2026, +77%). https://finance.yahoo.com/sectors/technology/article/meta-microsoft-amazon-and-alphabet-are-about-to-spend-a-shocking-amount-of-money-to-dominate-the-ai-era-115359575.html

[22] FactSet Insight, “Hyperscalers Tap External Financing as AI Capex Outruns Cash Flow,” July 23, 2026 (incremental debt 9%→32% of capex; Meta external cloud “on the table”). https://insight.factset.com/hyperscalers-tap-external-financing-as-ai-capex-outruns-cash-flow

[23] TMT Finance, “2026 Hyperscaler Capex Tops US$700bn – Analysis,” August 2026 (Amazon capacity constraints through 2027; Oracle ~US$70bn FY2027 capex). https://www.tmtfinance.com/intel/2026-hyperscaler-capex-tops-us700bn-analysis

[24] International Energy Agency, “Energy and AI — Energy Demand from AI,” 2025 (base-case ~945 TWh data-centre demand by 2030; Lift-Off Case >1,700 TWh by 2035). https://www.iea.org/reports/energy-and-ai/energy-demand-from-ai

[25] International Energy Agency, “Key Questions on Energy and AI — Executive Summary,” 2026 (data-centre electricity demand +17% in 2025; AI-focused data centres +50%). https://www.iea.org/reports/key-questions-on-energy-and-ai/executive-summary

[26] Enlit World, “AI and Data Centre Electricity Use Continues to Surge – IEA,” April 2026 (top-five capex >$400B in 2025, +75% in 2026; per-task power consumption falling by an order of magnitude annually). https://www.enlit.world/library/ai-and-data-centre-electricity-use-continue-to-surge-iea-finds

[27] Harvard Belfer Center, Power and AI Initiative, “AI, Data Centers, and the U.S. Electric Grid: A Watershed Moment,” February 10, 2026 (LBNL: 176 TWh in 2023 → 325–580 TWh, 6.7–12.0% of U.S. consumption, by 2028). https://www.belfercenter.org/research-analysis/ai-data-centers-us-electric-grid

[28] Consumer Reports, “AI Data Centers: Big Tech’s Impact on Electric Bills, Water, and More,” March 20, 2026 (78% of Americans concerned; Virginia polling; Ari Peskoe on transmission cost pass-through). https://www.consumerreports.org/data-centers/ai-data-centers-impact-on-electric-bills-water-and-more-a1040338678/

[29] Ari Peskoe (Harvard Law School), interviewed in Harvard Salata Institute, “The Data Center Boom Is Colliding with the Grid’s Hardest Problems,” May 2026 (PJM monitor: >$20B capacity-market costs; >$20B regional transmission). https://salatainstitute.harvard.edu/data-centers-ai-artificial-intelligence-grid-permitting-transmission-electricity-energy/

[30] Fortune, “Data Centers Could Hike Power Costs in Some States Over 50% by 2030,” May 19, 2026 (Environmental Research Letters modeling: wholesale +6–29% nationally by 2030; Virginia up to 57%; Gallup opposition polling). https://fortune.com/2026/05/19/data-centers-electricity-costs-us-public-opinion/

[31] Saving to Invest, “The AI Data Center Boom Is Rapidly Raising Your Electric Bill,” September 2026 ($31B 2025 rate-increase requests; $18.6B in H1 2026; state tax-exemption figures); and Fortune, “Data Centers Were Actually Making Electricity Costs Cheaper…,” July 26, 2026 (PJM $6.3B consumer cost attribution). https://fortune.com/2026/07/26/data-centers-electricity-costs-cheaper-7billion-buildout-ai-demand/

[32] Steve Clemmer (Union of Concerned Scientists), quoted in E&E News by POLITICO, “What AI Data Centers Do — and Don’t Do — to Electricity Prices,” April 20, 2026 (UCS estimate: $886–978B data-center-linked electricity costs by 2050). https://www.eenews.net/articles/what-ai-data-centers-do-and-dont-do-to-electricity-prices/

[33] Brookings Institution, “Global Energy Demands Within the AI Regulatory Landscape,” June 10, 2026 (data-center consumption trajectory; per-query energy estimates). https://www.brookings.edu/articles/global-energy-demands-within-the-ai-regulatory-landscape/

[34] CNBC, “The Question Everyone in AI Is Asking: How Long Before a GPU Depreciates?,” November 14, 2025 (Satya Nadella on depreciation pacing; industry useful-life practices). https://www.cnbc.com/2025/11/14/ai-gpu-depreciation-coreweave-nvidia-michael-burry.html

[35] Deep Quarry (Substack), “Depreciation of GPUs: Between Useful Lives and Useful Myths,” December 2025 (Michael Burry’s ~$176B understated-depreciation estimate, 2026–2028; hyperscaler useful-life history). https://deepquarry.substack.com/p/depreciation-of-gpus-between-useful

[36] Angel Investors Network, “Private Credit’s AI Datacenter Boom: Hidden Concentration Risk,” August 2026 (CoreWeave >$30B 2026 raises; $8.5B facility backed by Meta contract worth up to $14.2B; GPU collateral resale dynamics). https://angelinvestorsnetwork.com/alternative-investments/private-credit-ai-datacenter-concentration-risk

[37] Kristalina Georgieva (Managing Director, IMF), quoted in Business Standard, “Global Economy in Tug-of-War Between Oil Shock and AI Boom: IMF Chief,” August 26, 2026. https://www.business-standard.com/world-news/global-economy-in-tug-of-war-between-oil-shock-ai-boom-imf-chief-126082600327_1.html

[38] Gita Gopinath (First Deputy Managing Director, IMF), quoted in Fortune, “IMF Official Delivers Stark Warning on AI’s Potential to Turn an Ordinary Downturn into a Severe Economic Crisis,” June 9, 2024. https://fortune.com/2024/06/09/ai-risks-recession-economic-crisis-job-losses-financial-markets-supply-chains-imf/

[39] MIT Technology Review, “Three Things in AI to Watch, According to a Nobel-Winning Economist,” May 11, 2026 (Daron Acemoglu on agentic AI, lab economists, and missing consumer applications). https://www.technologyreview.com/2026/05/11/1137090/three-things-in-ai-to-watch-according-to-a-nobel-winning-economist/

[40] Daron Acemoglu (Institute Professor, MIT; 2024 Nobel Laureate in Economics), interviewed in Fortune, “Nobel Laureate Daron Acemoglu on the ‘Brainless’ AI Discourse…,” June 21, 2026 (~0.55% decadal TFP estimate; ~1–1.5% GDP). https://fortune.com/2026/06/21/nobel-laureate-daron-acemoglu-ai-productivity-capitalism-democracy/

[41] The White House, “America’s AI Action Plan,” July 2025 (recommendation to improve the financial market for compute, including spot and forward markets). https://www.whitehouse.gov/wp-content/uploads/2025/07/Americas-AI-Action-Plan.pdf

[42] “AI Token Futures Market: Commoditization of Compute and Derivatives Contract Design,” arXiv working paper 2603.21690 (2026), drawing on Cheng & Xiong (2014), Tang & Xiong (2012), Basak & Pavlova (2016), and Singleton (2014) on commodity financialization. https://arxiv.org/pdf/2603.21690