Introduction: The Chip Nvidia Did Not Have to Make

On September 10, 2026, a relatively small announcement from Santa Clara offered an unusually revealing glimpse into the next phase of the artificial-intelligence infrastructure race. d-Matrix, an inference-chip startup that has spent the better part of a decade developing an alternative to conventional GPUs, announced a multi-year product-roadmap collaboration with Nvidia under which its next-generation Raptor XPU will be integrated directly into Nvidia’s AI infrastructure through NVLink Fusion, with the first Raptor-equipped MGX racks expected to reach initial availability in the fourth quarter of 2027 and the Raptor silicon itself scheduled to complete its final design stage—its tape-out—before the end of 2026.[1] At first glance, the announcement looked like just another interoperability agreement between a dominant technology company and an ambitious semiconductor startup, the kind of press release that circulates for a day and disappears beneath the next wave of trillion-dollar datacenter commitments. But its strategic significance is much larger than its immediate commercial terms, and it deserves to be read slowly, because inside its unglamorous technical language sits one of the clearest available previews of how power in the artificial-intelligence economy may be reorganized between now and 2030.

Consider what d-Matrix actually is. The company exists for one reason: its founders believe that inference—the continuous, high-volume, latency-sensitive work of running trained models for billions of users—can be performed differently, and for many workloads more efficiently, than on the general-purpose GPU architectures that dominated the first phase of the generative-AI boom. Its Raptor accelerator, the successor to its Corsair platform, introduces a 3D in-memory compute architecture, known as 3DIMC, that stacks computation and memory closely together precisely in order to reduce the latency and energy cost of moving data, which d-Matrix regards as the true bottleneck of large-scale generative inference.[2] In other words, d-Matrix is not an Nvidia clone and has never wanted to be one. It is an architectural dissident, a company whose entire reason for existing is the conviction that the dominant accelerator paradigm is not optimal for the workload that will define the next decade of computing.

And yet, instead of forcing customers to choose between a d-Matrix system and an Nvidia system, Raptor is being designed to enter the Nvidia rack itself. Under the announced roadmap, d-Matrix will work with Nvidia to incorporate its inference XPUs directly into Nvidia’s latest rack reference architecture, an environment populated by Nvidia Vera CPUs, NVLink switches, BlueField-4 DPUs, ConnectX-9 SuperNICs, and Spectrum-X Ethernet networking, with connectivity specialist Astera Labs—itself a member of the NVLink Fusion ecosystem—supplying custom high-speed data-flow solutions, and with the entire rack built from modular, cable-free trays drawn from Nvidia’s mature MGX ecosystem and global supply chain.[1] Nvidia, for its part, described the arrangement in the language of ecosystem expansion, presenting NVLink Fusion as an accelerated, lower-risk path from custom silicon to large-scale deployment, and welcoming d-Matrix into a partner roster that already includes AWS, Arm, Intel, Fujitsu, SiFive, Alchip, Astera Labs, GUC, Marvell, MediaTek, Samsung, Cadence, Synopsys, Ayar Labs, and Lightmatter.[3] d-Matrix cofounder and chief executive Sid Sheth framed the economics of the decision with unusual candor during the press briefing:

“Demand for inference is soaring, but capital, time and energy remain finite.” — Sid Sheth, cofounder and CEO, d-Matrix [3]

The architecture described by the two companies is therefore revealing in a way that transcends this single partnership. The XPU may belong to d-Matrix; the rack still speaks Nvidia. The processor may be specialized; the scale-up fabric, the scale-out network, the data-processing infrastructure, the reference architecture, the deployment supply chain, and the management software can all remain organized around Nvidia technology. Nvidia does not need to manufacture the inference accelerator at the center of every workload in order to remain deeply, perhaps decisively, embedded in the system surrounding it.

This distinction matters because the AI industry is entering a period in which the accelerator itself is becoming structurally more heterogeneous. Training, post-training, reinforcement learning, long-context reasoning, retrieval, video generation, robotics, and high-volume inference do not necessarily favor an identical processor architecture, and the largest buyers of compute have concluded that they cannot afford to pretend otherwise. Amazon continues to advance Trainium, whose third generation is delivering up to 40 percent better price-performance than its predecessor and whose future revenue commitments from customers have been reported at more than $225 billion.[4] Google continues to iterate its TPU line. Meta plans to begin manufacturing Iris, the newest chip in its MTIA accelerator program, in September 2026, as part of an infrastructure plan that contemplates as much as $145 billion of AI spending this year and a doubling of computing capacity from seven gigawatts to fourteen gigawatts by 2027.[5] OpenAI has committed, with Broadcom, to deploying ten gigawatts of its own custom-designed accelerators between late 2026 and 2029.[6] Qualcomm is moving into datacenter AI silicon; Broadcom and Marvell are building custom-chip franchises measured in the tens of billions of dollars; and startups such as d-Matrix and Groq have designed processors around inference rather than general-purpose GPU computation. The semiconductor market could therefore become considerably more fragmented at the level of the processor even while the infrastructure surrounding those processors becomes more standardized—and it is precisely in that gap, between fragmenting silicon and consolidating architecture, that this paper locates the next great contest of the AI economy.

That possibility changes the strategic question facing Nvidia. For much of the first phase of generative AI, the central question was straightforward: who owns the best accelerator? Nvidia’s answer was overwhelmingly persuasive, and the financial results of that persuasion remain staggering. In its fiscal second quarter of 2027, ended July 26, 2026, Nvidia reported revenue of $96.2 billion, up 106 percent from a year earlier, of which $89.0 billion came from the data center segment alone, with gross margins of 75 percent and guidance pointing toward $108 billion in the following quarter—while assuming zero data-center compute revenue from China.[7] Jensen Huang, announcing the results, compressed the company’s worldview into a single formulation:

“AI has reached its inflection point. It’s doing useful work.” — Jensen Huang, founder and CEO, Nvidia [7]

But the next stage of the industry may revolve around a different question entirely: who defines how hundreds or thousands of heterogeneous accelerators communicate, share memory, move data, connect to networks, fit into racks, obtain software support, and operate together as one AI factory? Nvidia’s Vera Rubin architecture, unveiled in detail at CES 2026 and entering production shipments in the fall of 2026, already illustrates how far the unit of competition has moved beyond a discrete GPU. Vera Rubin NVL72 combines 72 Rubin GPUs, 36 Vera CPUs, NVLink 6 switching, ConnectX-9 SuperNICs, and BlueField-4 DPUs into a single liquid-cooled rack that Nvidia explicitly describes as one AI supercomputer, while Spectrum-X Ethernet and Quantum-X800 InfiniBand extend the system outward across pods and campuses.[8] Nvidia increasingly describes the entire datacenter as the computing unit rather than treating each accelerator as an isolated product, and its published Vera Rubin materials are unambiguous on this point: the rack is designed to make multiple specialized components behave as one coherent machine for producing intelligence.[9]

NVLink Fusion takes the strategy one step further. Instead of requiring every accelerator inside that infrastructure to carry the Nvidia name, Nvidia can invite custom silicon into the architecture. That is why the d-Matrix announcement deserves more attention than its immediate revenue implications suggest: it demonstrates, in working commercial form, that Nvidia may be developing a strategy capable of surviving the partial commoditization—or, more precisely, the specialization—of the accelerator itself.

This paper calls that possibility Fabric Hegemony.

Fabric Hegemony describes a condition in which competitive advantage migrates upward from control of an individual processor toward control of the high-bandwidth interconnects, switching systems, memory interfaces, networking, rack architectures, system software, validation processes, and deployment ecosystems that allow thousands of processors to operate collectively. Under Fabric Hegemony, a company does not need to manufacture every accelerator. It needs to establish the architecture through which those accelerators become economically useful at scale.

The argument is particularly important within the Five-Layer AI Economy framework that organizes this paper. Layer 2—the chip layer—cannot be understood independently from Layer 3, the datacenter layer, because increasingly the economically meaningful product is neither a chip nor a building but a tightly integrated computational system occupying tens or hundreds of megawatts. Layer 4 models and Layer 5 applications then depend on how efficiently that system can produce intelligence: tokens per second, tokens per watt, latency per request, and ultimately dollars per unit of useful reasoning. Layer 1, energy, sits beneath everything, because every unnecessary movement of data is paid for twice, once in wasted time and once in wasted watts. The fabric connecting accelerators consequently becomes an economic bridge running through the entire stack, which is why the emerging contest is larger than Nvidia versus AMD, or GPUs versus ASICs. It is a contest over the constitutional architecture of the AI factory.

And there is already a competing vision. The UALink Consortium—whose board includes Alibaba, AMD, Apple, Astera Labs, AWS, Cisco, Google, HPE, Intel, Meta, Microsoft, and Synopsys—has published an open scale-up interconnect specification designed to connect heterogeneous accelerators from any vendor. Its UALink 200G 1.0 specification envisions high-bandwidth, low-latency, load/store-semantics communication across as many as 1,024 accelerators in a single pod, and its subsequent specifications add in-network compute and management capabilities as the consortium’s roadmap extends toward still larger systems.[10] UALink therefore represents a fundamentally different theory of infrastructure: rather than allowing one vendor’s architecture to become the connective standard of the AI era, an industry consortium is attempting to establish an interoperable alternative in which the choice of accelerator and the choice of fabric are deliberately decoupled.

The outcome of this contest could shape the economics of artificial intelligence through 2030 and beyond. If specialized accelerators proliferate but most of them enter Nvidia-designed racks, Nvidia may lose accelerator share without losing architectural power. If open fabrics become sufficiently performant and mature, hyperscalers could combine silicon from multiple vendors without passing through Nvidia’s infrastructure at all. And if the market fragments into several incompatible fabrics, the AI industry could develop something resembling competing technological jurisdictions, each with its own hardware, software, suppliers, and economics—a Balkanization of compute whose costs would ultimately be paid in the price of every token.

The September 10 d-Matrix announcement should therefore be read less as a partnership between two semiconductor companies than as an early constitutional experiment: an experiment in who gets to define the machine around the chip.


Why I Chose the Title “Fabric Hegemony”

I chose Fabric Hegemony because the phrase captures a possible transformation in the nature of Nvidia’s power, and because each of its two words carries analytical weight that more conventional vocabulary—“market share,” “moat,” “platform”—fails to convey. “Fabric” refers not merely to networking cables but to the entire connective architecture through which CPUs, GPUs, XPUs, memory, switches, DPUs, and complete racks operate as one computational system: the NVLink scale-up domain, the Spectrum-X scale-out network, the chip-to-chip links, the custom memory interfaces, the reference rack designs, and the software that orchestrates them. “Hegemony” describes something broader and more durable than conventional market dominance: the ability to influence the rules, interfaces, standards, and economic environment within which other companies compete, such that participation in the hegemon’s system becomes more attractive than resistance to it. Hegemonic power, in international relations as in technology, is distinguished precisely by the fact that it does not require coercion at every point of contact; it operates through the voluntary participation of actors who calculate, correctly, that joining the system serves their interests better than opposing it. Nvidia could consequently face increasingly capable accelerator competitors while simultaneously making those competitors more valuable to themselves—and to Nvidia—when they participate inside Nvidia’s architectural ecosystem.

The subtitle—“When Nvidia Can Lose the Accelerator—and Still Control the Architecture That Connects the AI Factory”—makes the paradox explicit. The important question for 2027 through 2030 may not be whether Nvidia maintains nearly every point of accelerator share, a proposition that the custom-silicon programs of Amazon, Google, Meta, Microsoft, and OpenAI have already rendered doubtful. It may be whether Nvidia can transform itself from the dominant manufacturer of AI processors into the dominant architect of heterogeneous AI infrastructure. If that transformation succeeds, losing some chips would not mean losing the system; it might, perversely, strengthen it, because every specialized accelerator that plugs into an Nvidia rack validates the rack as the industry’s default constitutional form. That is why Fabric Hegemony fits the central argument of this paper, and why the framework developed below treats the interconnect not as an engineering detail but as the decisive political-economic terrain of the coming AI decade.


Section 1: From GPU Dominance to Rack-Scale Power


1.1 The Accelerator Was the First Battlefield

The first generative-AI infrastructure cycle, running roughly from the release of ChatGPT in late 2022 through the Blackwell deployments of 2025, centered overwhelmingly on a single scarce object: the Nvidia data-center GPU. The reasons were structural rather than accidental. Frontier model training required enormous quantities of dense floating-point arithmetic; Nvidia’s GPUs delivered that arithmetic at scale; and, critically, fifteen years of accumulated investment in the CUDA software ecosystem meant that virtually every research framework, every optimization library, every kernel, and every trained engineer in the field assumed Nvidia hardware as the substrate of serious work. Access to H100s, and later to Blackwell systems, became the defining constraint on which laboratories could train frontier models at all, and the GPU allocation queue became a kind of shadow capital market in which compute functioned as currency. Nvidia’s control of an estimated 90 percent or more of the merchant market for high-end AI training silicon during this period was not merely a commercial statistic; it was the organizing fact of the entire industry, the gravitational center around which capital expenditure plans, national industrial strategies, venture financing, and even foreign policy—in the form of U.S. export controls—were arranged.[11] The physical GPU became the defining asset of the early AI boom in the way that spectrum licenses defined early mobile telephony or drilling rights defined early petroleum: the thing whose possession separated participants from spectators.

Yet even at the peak of this first cycle, the seeds of its transformation were visible. A single GPU, however powerful, trains nothing of consequence. Frontier models were trained on clusters of tens of thousands of accelerators, and the performance of those clusters depended as much on how the chips communicated as on how fast each chip computed. The battlefield was already migrating, quietly, from the processor to the system—and Nvidia, more clearly than any of its rivals, understood that migration and began building for it years before its competitors recognized that the war had changed.


1.2 The Unit of Compute Is Getting Larger

The history of AI infrastructure over the past four years can be told as the steady enlargement of the unit of compute: from the individual accelerator, to the eight-GPU server, to the NVLink-connected rack, to the multi-rack pod, and finally to what Nvidia now calls, without embarrassment, the AI factory—a building, or campus of buildings, engineered end to end as a single machine for converting electricity into tokens. Each enlargement followed from the same underlying pressure. Model parameter counts and context lengths grew faster than the memory and bandwidth of any single device, forcing computation to be distributed; distribution created communication; and communication, at sufficient scale, became the dominant cost. Mixture-of-experts architectures route tokens dynamically across dozens of expert subnetworks residing on different chips. Long-context reasoning models maintain key-value caches measured in terabytes that must be held, moved, and shared across devices. Agentic workloads chain together model calls, tool invocations, retrievals, and verifications in patterns that touch many processors in rapid succession. In every case, the economically meaningful question stopped being “how fast is the chip?” and became “how fast is the system of chips, taken as a whole, at the workload the customer is actually paying for?”

This is why the announcements of 2026 concern racks and gigawatts rather than devices. When Amazon and Nvidia agreed in August 2026 to deploy an additional two million Nvidia GPUs across AWS infrastructure in 2027 and 2028—after AWS demand had already consumed the one million GPUs pledged only months earlier—the agreement was described by both companies not as a chip order but as a co-engineered expansion spanning GPUs, CPUs, networking, memory, and software.[12] The unit of purchase, the unit of engineering, and the unit of competition have converged on the system, and the system’s defining property is not arithmetic but coordination.


1.3 Why Data Movement Becomes Compute

To understand why the fabric matters economically, one must begin from an accounting identity that the industry has internalized only gradually: idle compute created by communication bottlenecks is economically equivalent to unavailable compute. A $50,000 accelerator that spends forty percent of its cycles waiting for gradients to synchronize, for expert-routing traffic to arrive, for KV-cache pages to migrate, or for collective-communication operations to complete is, from the perspective of the income statement, a $30,000 accelerator that cost $50,000. Every improvement in interconnect bandwidth and latency therefore functions as a direct increase in the effective capital efficiency of the entire installed fleet, and every deficiency functions as a hidden depreciation charge.

The technical sources of this dependence are worth enumerating, because they explain why the problem intensifies rather than diminishes as models advance. Memory bandwidth constrains how quickly parameters and activations can be fed to compute units, which is why Rubin-generation GPUs carry 288 GB of HBM4 at 22 TB/s per device and why the industry’s supply commitments are increasingly dominated by memory rather than logic.[13] Collective-communication operations—all-reduce, all-gather, all-to-all—scale super-linearly in cost with cluster size and sit directly on the critical path of every training step. Mixture-of-experts routing converts model architecture into network traffic, so that the shape of the neural network literally becomes the shape of the packet flow. Disaggregated inference, in which prefill and decode phases run on different hardware pools, turns the KV cache into a constantly migrating asset whose movement cost is priced into every served token. In each case, arithmetic that cannot be fed is arithmetic that does not exist, and the fabric that does the feeding becomes, functionally, part of the compute itself. The industry’s engineers have a shorthand for this recognition—“data movement is the new compute”—and the entire architecture of the modern AI rack is an attempt to act on it.


1.4 Vera Rubin and the Rack as a Computer

Nvidia’s Vera Rubin platform, announced in full at CES 2026 and entering production shipments in the fall of 2026, is the clearest available expression of this architectural transition, and it rewards close reading. The Vera Rubin NVL72 rack unifies 72 Rubin GPUs and 36 Vera CPUs—the latter built on 88 custom Olympus Arm cores designed specifically for the data-movement and orchestration demands of agentic AI—through sixth-generation NVLink into a single shared-memory fabric, supported by ConnectX-9 SuperNICs, BlueField-4 DPUs, and Spectrum-X Ethernet or Quantum-X800 InfiniBand for scale-out connectivity.[8] Nvidia describes the design philosophy as “extreme co-design”: six distinct chips developed together so that the rack functions as one distributed accelerator rather than as a cabinet of servers.[14] The published specifications convey the scale of the integration: NVLink 6 provides 3.6 TB/s of all-to-all scale-up bandwidth per GPU and 260 TB/s across the rack fabric—more than twice the cross-sectional bandwidth of the global internet—while the rack as a whole delivers on the order of 3,600 petaflops of NVFP4 inference, roughly 21 TB of HBM4 GPU memory, and, by Nvidia’s claims, up to ten times lower cost per token than Blackwell-generation systems.[8]


Table 1. Nvidia Vera Rubin NVL72: The Rack as a Single Computer

ComponentRole in the SystemKey Published Figures
Rubin GPU (×72)Training and inference arithmetic288 GB HBM4 per GPU at 22 TB/s
Vera CPU (×36)Orchestration, data movement, agentic processing88 custom Olympus Arm cores; NVLink-C2C at 1.8 TB/s
NVLink 6 SwitchScale-up fabric3.6 TB/s per GPU; 260 TB/s per rack
ConnectX-9 SuperNICScale-out connectivity1.6 Tb/s per GPU
BlueField-4 DPUNetworking, storage, security offload; KV-cache extension~6× compute of BlueField-3
Spectrum-X / Quantum-X800Rack-to-rack and pod-scale networking102 TB/s switching with co-packaged optics (Spectrum-6)

Sources: Nvidia, StorageReview, Supermicro, SemiAnalysis.[8][9][14][15]


The analytical point matters more than the specifications, impressive as they are: Nvidia increasingly sells the coordinated machine, not merely the chip. A customer who buys Vera Rubin NVL72 is not buying seventy-two GPUs; the customer is buying a rack-scale locality domain, a pre-validated thermal and power envelope, a supply chain of more than eighty MGX ecosystem partners, and a software stack that treats the rack as the atomic unit of scheduling.[9] Once the atomic unit of the industry is the rack, whoever defines the rack defines the industry—and that recognition is the hinge on which this entire paper turns.


1.5 Training Architecture Versus Inference Architecture

The second force reshaping the competitive landscape is the divergence between training and inference as economic activities. Training is episodic, throughput-oriented, tolerant of batch latency, and concentrated in a small number of frontier laboratories; inference is continuous, latency-sensitive, exploding in volume, and distributed across every application that touches an end user. The assumption that a single general-purpose accelerator architecture should dominate both workloads was always a convenience of the early boom rather than a law of computer architecture, and as inference grows into the larger share of total compute demand, that convenience is dissolving. Low-latency interactive serving rewards memory-centric designs that minimize data movement per token; enormous token volumes reward architectures optimized for cost and power per token rather than peak flops; power-constrained datacenters reward performance per watt above performance in the absolute; and specific model structures—mixture-of-experts, speculative decoding, retrieval-augmented pipelines—create niches in which specialized silicon can beat general-purpose silicon by meaningful margins. This is the opening through which d-Matrix, Groq, the hyperscaler ASIC programs, and Qualcomm’s datacenter ambitions have all moved, and it is why Nvidia itself now segments its own product line, pairing Vera Rubin NVL72 with specialized companions—including the Groq-derived LPX racks for ultra-low-latency serving—rather than insisting that one architecture serve every purpose.[9] The era of the universal accelerator is ending; the question is what organizes the plurality that replaces it.


1.6 The d-Matrix Raptor Experiment

Return, then, to September 10, 2026, with the preceding analysis in hand. d-Matrix’s Raptor embodies precisely the specialization thesis: a memory-centric, 3DIMC-based architecture aimed at the premium tier of the token economy, the latency-critical serving of coding assistants, real-time chatbots, and voice agents where interactivity commands a price premium.[1][2] If accelerator differentiation automatically produced infrastructure independence, d-Matrix would be building its own racks, its own scale-up fabric, its own networking stack, and its own supply chain, and it would be asking customers to bet on all of them simultaneously. Instead, d-Matrix chose to place its differentiated silicon inside Nvidia’s environment—Vera CPUs, NVLink switches, BlueField-4, ConnectX-9, Spectrum-X, MGX trays, Nvidia’s supplier network—because doing so converts a multi-year, capital-devouring systems-engineering problem into a socket-level integration problem, and because customers who already operate Nvidia AI factories can adopt Raptor without abandoning anything they have built.[1] The partnership thereby becomes the first clean empirical illustration of this paper’s thesis: accelerator differentiation does not automatically produce infrastructure independence. A company can win the argument about how inference should be computed and still conduct that argument entirely inside another company’s architecture—which means the architecture, not the argument, may be where the durable power resides.


1.7 Fabric Hegemony Within the Five-Layer AI Economy

The concept of Fabric Hegemony acquires its full significance only when positioned across the Five-Layer AI Economy that frames this series of papers, because the fabric is the rare piece of infrastructure whose economics propagate through four of the five layers simultaneously.


Layer 1 — Energy. Every byte moved consumes joules, and at the scale of hundred-megawatt and gigawatt campuses, fabric efficiency is indistinguishable from power economics. A superior interconnect that raises useful computation per megawatt functions, from the grid’s perspective, like additional generating capacity that no utility had to build.


Layer 2 — Chips. As accelerators become heterogeneous—training silicon, inference silicon, reasoning silicon, video silicon, hyperscaler ASICs—the value of any individual chip increasingly depends on the fabric’s willingness and ability to connect it. The fabric becomes the market-access layer for silicon itself.


Layer 3 — Datacenters. Racks, switching, cabling, cooling, and reference architectures determine how chips become clusters and how clusters become factories. This is the layer where NVLink, MGX, and Spectrum-X live, and where UALink proposes its alternative constitution.


Layer 4 — Models. Fabric bandwidth and latency determine which model architectures are economically servable at all: trillion-parameter mixture-of-experts models with million-token contexts exist as commercial products only because rack-scale fabrics make their communication patterns affordable.


Layer 5 — Applications and Agents. Latency and cost per token, both fabric-dependent, decide whether agentic systems—which multiply model calls by orders of magnitude—are economically viable as businesses rather than demonstrations.


The fabric, in short, is not a hidden technical detail buried in Layer 3. It is connective economic infrastructure running vertically through the stack, and control of it confers influence over every layer it touches. That is the structural reason the interconnect has become worth fighting over, and the remainder of this paper examines the fight.


Section 2: Nvidia Can Lose the Chip and Win the System


2.1 NVLink Fusion as Strategic Architecture

For most of its history, NVLink was the most jealously guarded perimeter in Nvidia’s empire: a proprietary scale-up interconnect available exclusively to Nvidia GPUs and, later, to Nvidia’s own Grace and Vera CPUs, functioning simultaneously as a performance advantage and as a wall. What changed in May 2025, when Nvidia first unveiled NVLink Fusion, and what has accelerated dramatically through 2026, is that Nvidia deliberately opened a gate in that wall—on its own terms. NVLink Fusion allows semiconductor partners and hyperscalers to build semi-custom AI infrastructure in which non-Nvidia accelerators and non-Nvidia CPUs join the NVLink scale-up domain directly, gaining access to the same high-bandwidth, low-latency fabric, the same MGX rack architecture, the same networking, power, cooling, and software environment, and the same globally validated supply chain that Nvidia’s own GPUs enjoy.[16] Nvidia’s own description of the resulting model is candid: a “semi-custom AI factory,” in which Nvidia provides the interconnect, rack architecture, software, and supply chain while the partner concentrates on differentiated silicon.[3] The strategic inversion is profound. Rather than insisting that every processing element be an Nvidia GPU—a position that the hyperscaler ASIC programs were eroding regardless of Nvidia’s wishes—Nvidia created a sanctioned pathway through which external silicon enters Nvidia’s world, on Nvidia’s fabric, inside Nvidia’s racks. The wall did not come down; it acquired a toll gate.


2.2 The Architecture Surrounding the XPU

To grasp how much of the system remains Nvidia’s even when the accelerator is not, it is worth walking deliberately through the strategic stack that surrounds a third-party XPU inside an NVLink Fusion deployment, because each element represents a point of engineering dependence, a revenue opportunity, or both. The Vera CPU hosts the accelerator, orchestrates agentic workloads, and terminates NVLink-C2C, the 1.8 TB/s chip-to-chip link that binds host and accelerator into a coherent memory domain. The NVLink switches implement the scale-up fabric itself, the terrain on which all accelerator-to-accelerator communication occurs. NVLink-C2C and the newly introduced NVHBM—Nvidia’s custom high-bandwidth memory technology, developed with memory suppliers and now being extended to partners including Amazon’s Annapurna Labs for future Trainium generations—govern how the partner’s silicon reaches memory at competitive bandwidth and power.[12][16] BlueField-4 DPUs offload networking, storage, security, and increasingly the management of KV-cache memory tiers. ConnectX-9 SuperNICs and Spectrum-X Ethernet carry the scale-out traffic that binds racks into pods and pods into factories. The MGX rack architecture, with its modular cable-free trays and its ecosystem of more than eighty manufacturing partners, defines the physical, thermal, and electrical world into which everything must fit.[9] And above the hardware sits Nvidia’s system software, telemetry, validation, and deployment tooling, the accumulated operational knowledge of the largest AI installed base on earth.

The paper must therefore pose its central counterfactual plainly: if another company’s accelerator occupies the socket, but Nvidia technologies occupy the CPU, the fabric, the memory interface, the DPU, the NIC, the switch, the rack, and the management plane—who actually owns the platform? The honest answer is that the socket’s occupant owns a component, while Nvidia owns the environment, and in platform economics the environment has historically been the more valuable possession.


2.3 d-Matrix as the Prototype

The d-Matrix arrangement illustrates the division of labor with almost diagrammatic clarity. d-Matrix supplies what it is genuinely best at: specialized, memory-centric inference compute embodied in Raptor, aimed at the ultra-low-latency premium tier of the token economy. Nvidia supplies essentially everything else that turns a chip into a deployable product: the Vera CPUs, the NVLink scale-up foundation, the BlueField and ConnectX networking layers, the Spectrum-X scale-out environment, the MGX rack design and its manufacturing ecosystem, and the accumulated validation and supply-chain maturity that allow a startup’s silicon to be manufactured into full racks and delivered to hyperscalers and neoclouds at production scale.[1][3] For d-Matrix, the calculus is compelling and Nvidia describes it precisely as such: building a custom XPU is hard, deploying one at AI-factory scale is harder, and NVLink Fusion lets the XPU maker skip years of rack-scale infrastructure development—sourcing, validating, cooling, powering, cabling, and supporting an entire systems business—and plug instead into a proven, globally deployed platform.[3] The startup trades a measure of architectural sovereignty for time-to-market, capital efficiency, and immediate compatibility with the world’s dominant AI infrastructure. It is, on its own terms, a rational and possibly brilliant trade. But the aggregate effect of many such rational trades, made independently by many silicon companies, is the subject of this paper: each trade individually strengthens the trader, and collectively they entrench the architecture into which all of them are trading.


2.4 Compatibility as Competitive Strategy

Traditional accounts of platform power emphasize exclusion: the dominant firm locks competitors out and customers in. But the deepest platform power in technological history has more often arisen from the opposite maneuver—making it advantageous for competitors to become compatible. Microsoft did not exclude application developers from Windows; it courted them, and every third-party application deepened the operating system’s indispensability. Visa and Mastercard did not issue every card or operate every bank; they defined the rails, and every additional issuer and merchant made the rails harder to abandon. The dominant cloud providers did not write every workload; they defined the APIs against which workloads were written, and each migration onto those APIs raised the cost of ever migrating off. In each case, the platform grew stronger precisely because rivals and complementors obtained genuine commercial benefits by joining it, which meant the platform’s expansion required no coercion and provoked, at least initially, no organized resistance.

NVLink Fusion applies this logic, for the first time at full scale, to AI silicon. Every accelerator company that adopts it receives something real: faster deployment, validated infrastructure, access to Nvidia’s customer base and supply chain. And every adoption returns something real to Nvidia: another constituency invested in NVLink’s continuation, another data point persuading customers that the Nvidia rack is where innovation arrives first, another reason for the next silicon startup to regard NVLink compatibility as the default rather than the exception. Compatibility, structured this way, is not a concession to competition. It is a competitive strategy—arguably the most sophisticated one available to a firm whose absolute share of sockets must mathematically decline.


2.5 Nvidia’s MediaTek and Marvell Strategy

The d-Matrix agreement, moreover, is not an isolated event but the latest installment in a deliberate, capital-backed pattern that unfolded across 2026 with remarkable speed. On March 31, 2026, Nvidia announced a $2 billion investment in Marvell Technology alongside a strategic partnership under which Marvell—one of the two dominant merchant designers of hyperscaler custom silicon—will provide custom XPUs and NVLink Fusion-compatible scale-up networking integrated with Nvidia’s Vera CPUs, ConnectX NICs, BlueField DPUs, NVLink interconnect, and Spectrum-X switches, while the two companies collaborate on silicon photonics and optical interconnect.[17][18] Jensen Huang framed the moment in the vocabulary of inference economics:

“The inference inflection has arrived. Token generation demand is surging.” — Jensen Huang, announcing the Marvell partnership [17]

Marvell’s chairman and chief executive Matt Murphy, for his part, emphasized the connective logic of the arrangement:

“Our expanded partnership with Nvidia reflects the growing importance of high-speed connectivity.” — Matt Murphy, chairman and CEO, Marvell [18]

Five months later, on August 31, 2026, Nvidia invested $3.5 billion in convertible bonds issued by MediaTek—one of the world’s largest chip-design houses and the other natural merchant partner for hyperscaler custom silicon—in an agreement under which MediaTek adopts NVLink Fusion as the platform for the custom AI accelerators it designs for cloud customers, incorporating NVLink Fusion chiplets, NVLink-C2C, NVHBM custom memory, advanced packaging, and MGX rack integration.[19] Industry observers noted the pattern immediately: Nvidia had now taken multi-billion-dollar positions in the very companies most capable of designing the chips meant to reduce dependence on Nvidia GPUs, and had bound their custom-silicon franchises to Nvidia’s interconnect in the process.[20] Alongside these stands the December 2025 arrangement with Groq—a reported $17–20 billion non-exclusive license to the inference startup’s technology, accompanied by the hiring of its founder Jonathan Ross and other executives—which brought low-latency LPU technology into Nvidia’s platform as the Groq LPX rack line and which is examined from the regulatory side in Section 4.[21][22]

The question this pattern forces is uncomfortable and important: is Nvidia’s balance sheet becoming a tool for expanding architectural adoption? With roughly $96 billion of quarterly revenue and commanding margins, Nvidia can deploy capital at a scale that reshapes the incentives of the entire custom-silicon industry, converting would-be architectural rivals into fabric tenants for sums that are, to Nvidia, rounding errors. Whatever one concludes about the competitive ethics of the strategy, its architectural logic is coherent: every dollar invested in Marvell, MediaTek, or Groq is a dollar invested in ensuring that the alternatives to Nvidia’s chips are built on Nvidia’s fabric.


2.6 The Architecture Toll

These developments justify introducing a central economic concept for the remainder of this paper: the architecture toll. Nvidia may increasingly capture value not only through accelerator margins—the dominant mechanism of the first boom—but through a diversified stream of revenues and advantages that accrue whenever anyone’s silicon, including a competitor’s, operates inside Nvidia’s architecture. The toll does not primarily take the form of literal royalties, and it need not. It arises through the sale of NVLink switches and NVLink Fusion chiplets; through Vera CPUs hosting third-party XPUs; through BlueField DPUs and ConnectX SuperNICs attached to every compute tray; through Spectrum-X networking carrying every rack’s scale-out traffic; through NVHBM memory interfaces licensed into partner silicon; through MGX reference designs steering the manufacturing ecosystem; through software, validation, and system-integration services; and, more subtly, through the increased demand for Nvidia’s own GPUs that follows whenever the Nvidia rack is confirmed as the venue where all serious silicon must appear. Analysts describing the Marvell arrangement captured the structure bluntly: deals of this kind ensure that custom chips designed for hyperscalers still generate Nvidia revenue through the platform components that surround them—an ecosystem tax collected in components rather than in cash.[23] The toll metaphor should be handled carefully, because tolls imply compulsion and NVLink Fusion’s participants join voluntarily; but the economics of a voluntary toll and a compulsory one converge once the tolled road is the only one that reaches the market at scale.


2.7 From Semiconductor Vendor to AI Systems Architect

The conclusion of this section is that Nvidia’s most consequential evolution may be organizational and conceptual rather than technological. The company has already lived three lives. The first Nvidia sold graphics processors to gamers and workstations. The second Nvidia sold AI accelerators to datacenters and, in doing so, became the most valuable semiconductor company in history. The third Nvidia—the Nvidia of Grace Blackwell and Vera Rubin—sold complete AI systems, rack-scale computers co-designed from silicon to software. The emerging fourth Nvidia attempts something categorically different: to define the architecture through which other companies’ silicon becomes an AI system. In this fourth incarnation, Nvidia’s product is not exactly a chip, a server, or even a rack; it is the constitutional order of the AI factory itself—the interfaces, standards, physical forms, and supply chains within which all participants, allies and rivals alike, conduct their business. Sellers of products can lose share; architects of constitutional orders are dislodged only by rival constitutions. Whether such a rival constitution exists, and whether it can win, is the subject of the next section.


Section 3: The Counter-Revolution: Open Fabrics and Custom Silicon


3.1 Hyperscalers Do Not Want Permanent Accelerator Dependence

No analysis of Fabric Hegemony is complete, or honest, without full weight given to the forces arrayed against it, and those forces begin with the balance sheets of Nvidia’s own largest customers. Amazon, Google, Meta, and Microsoft are collectively deploying capital expenditure at a scale without precedent in industrial history: roughly $301 billion in the first half of calendar 2026 alone, with full-year guidance across the four companies of approximately $725–733 billion—up more than 75 percent from 2025’s already record $410 billion—and Wall Street projections of $1 trillion or more in 2027, with Goldman Sachs modeling $5.3 trillion of hyperscaler capex between fiscal 2025 and fiscal 2030.[24][25][26] When spending reaches this magnitude, arithmetic that once seemed trivial becomes strategic: shaving even a few percentage points from cost per token, power per token, or hardware vendor margin translates into billions of dollars annually, and a supplier earning 75 percent gross margins on the largest line item in your budget is not merely a partner but a standing invitation to vertical integration.[7] Every hyperscaler has accepted the invitation. Their custom-silicon programs are no longer experimental side projects, hedges, or negotiating chips; they are strategic instruments of cost structure and supply security, funded at the level of national industrial programs and increasingly successful on their own technical terms.


3.2 The Custom-Silicon Explosion

The breadth of the custom-silicon movement as of September 2026 deserves systematic enumeration, because its scale is the single strongest argument that Nvidia’s accelerator share must decline even as the market grows.


Table 2. The Custom Silicon Landscape, September 2026

ProgramSponsor / PartnerStatus and Scale
TPU (v7 generation)GoogleMature; rack-scale before Nvidia; core of Gemini training and serving
Trainium 3 / InferentiaAmazon (Annapurna Labs)Trainium3 ~40% better price-performance than Trainium2; >$225B customer commitments; custom-chip business >$25B annual run rate
MTIA / IrisMeta, with Broadcom and TSMCIris entering production September 2026; ~$145B 2026 AI capex; 7 GW→14 GW capacity plan
MaiaMicrosoftSecond generation in development for Azure AI
OpenAI custom XPUOpenAI, with Broadcom10 GW of OpenAI-designed accelerators, deployments 2H 2026 through 2029
Merchant custom siliconBroadcom, Marvell, MediaTek, AlchipTens of billions in hyperscaler XPU design wins
Datacenter entrantsQualcomm (AI200/AI250 class)New rack-scale inference platforms
Inference specialistsd-Matrix, Groq, othersd-Matrix Raptor via NVLink Fusion; Groq licensed into Nvidia platform

Sources: company announcements and reporting compiled in this paper.[4][5][6][12][16][19][22]


Meta’s Iris illustrates the direction of travel with particular clarity. The chip—fourth in the MTIA line, designed with Broadcom, fabricated by TSMC—cleared its bug-testing phase in roughly six weeks without significant problems and enters production in September 2026, explicitly to reduce dependence on external vendors while supplementing, not replacing, Meta’s enormous Nvidia and AMD purchases.[5] Forrester analyst Mike Gualtieri distilled the strategic sentiment driving every one of these programs into a single sentence, observing that

“You can’t become an AI titan” while dependent on another company for your chips. — Mike Gualtieri, VP and Principal Analyst, Forrester [27]

OpenAI’s Sam Altman, announcing the Broadcom partnership, made the same point in the language of performance rather than sovereignty:

“We can get huge efficiency gains, and that will lead to much better performance.” — Sam Altman, CEO, OpenAI [28]

The custom-silicon explosion is therefore real, well funded, technically credible, and permanent. The question this paper poses is not whether it succeeds—at the level of the chip, it already has—but what it wins.


3.3 The Paradox of Custom Silicon

Here the analysis arrives at one of its most important paradoxes, and the events of August 2026 supplied its perfect empirical demonstration. A hyperscaler can design its own accelerator and still remain architecturally dependent on somebody else’s fabric. Amazon’s Annapurna Labs has built, in Trainium, the most commercially successful hyperscaler training ASIC outside Google—and yet, at re:Invent 2025 and again in the August 2026 expansion of the AWS–Nvidia partnership, Amazon announced that next-generation Trainium chips will support NVLink Fusion, and that Annapurna will adopt Nvidia’s custom NVHBM memory technology, precisely so that Trainium silicon and Nvidia GPUs can be combined within a common, shared, scale-up rack architecture.[12][16] Consider what this means with full seriousness: the flagship custom accelerator of the world’s largest cloud provider, the very chip built to reduce Nvidia dependence, will reach memory through Nvidia’s memory interface and reach its neighbors through Nvidia’s interconnect. Amazon retains, of course, its own formidable infrastructure assets—the Nitro system, the EFA scale-out fabric, and its position on the UALink board—and the arrangement is best read as deliberate optionality rather than surrender. But the paradox stands, and it generalizes: custom silicon does not automatically equal infrastructure sovereignty, because sovereignty resides not in the chip one designs but in the architecture through which the chip becomes a system. A hyperscaler that owns its accelerator but rents its fabric has relocated its dependence, not eliminated it.


3.4 UALink: The Open Challenge

The clearest institutional alternative to an Nvidia-controlled scale-up domain is the Ultra Accelerator Link Consortium, incorporated in October 2024 and now comprising more than 85 member companies, with a board that reads as a census of Nvidia’s largest customers and most motivated rivals: Alibaba, AMD, Apple, Astera Labs, AWS, Cisco, Google, HPE, Intel, Meta, Microsoft, and Synopsys.[29] The consortium’s UALink 200G 1.0 specification, published in April 2025, defines a low-latency, high-bandwidth, open-standard scale-up interconnect supporting up to 1,024 accelerators in a single pod—each accelerator addressed by a ten-bit routing identifier, with simple load/store memory semantics, atomic operations, and DMA across the entire pod—delivering 200 Gbps per lane while leveraging the ubiquitous Ethernet physical-layer ecosystem for cost and supply-chain leverage.[10][29] Subsequent specification work has added in-network compute, enabling computation within the fabric itself, along with management capabilities, and the consortium describes a continuing roadmap toward larger and faster systems.[10] Consortium president Peter Onufryk described the intent at the specification’s release, saying member companies are

“actively building an open ecosystem for scale-up accelerator connectivity.” — Peter Onufryk, President, UALink Consortium [30]

The consortium’s own January 2026 white paper articulates the economic theory of the challenge with precision: coupling accelerator selection with interconnect architecture creates compound lock-in that constrains silicon choices, limits negotiation leverage, and impedes topology evolution, whereas UALink deliberately decouples those decisions so that organizations can optimize each layer independently.[31] UALink is thus not merely a competing cable; it is a competing constitutional theory, holding that the fabric of the AI factory should be a commons governed by its users rather than an estate governed by its landlord.


3.5 Proprietary Maturity Versus Open Optionality

The competition between these constitutions can be framed as a contest between two different kinds of advantage. NVLink’s advantages are those of the incumbent estate: six generations of maturity; extreme co-design with CPUs, DPUs, NICs, and switches; a deployed base measured in millions of GPUs; a validated global supply chain of more than eighty rack-ecosystem partners; and the simple, commercially decisive fact that it ships in volume today, with 3.6 TB/s per GPU in the Rubin generation.[8][9] UALink’s advantages are those of the commons: openness, multi-vendor governance, architectural independence, freedom from single-supplier pricing power, and the backing of precisely the companies that control the majority of global AI capital expenditure. History cautions against assuming that the technically or philosophically superior protocol wins; it cautions equally against assuming the incumbent always does. The winner of a fabric war is the architecture that customers trust enough to deploy at hundred-megawatt scale with billions of dollars and years of roadmap at stake—and trust of that kind is built from silicon that exists, switches that ship, software that works, and failure modes that are understood. On that criterion NVLink holds a commanding lead as of late 2026, with UALink’s first full silicon ecosystem—switches, retimers, accelerator ports—still moving from specification toward volume deployment.[31] The strategic window matters enormously: every quarter in which open-fabric silicon remains immature is a quarter in which another wave of AI factories is poured, literally in concrete, around the proprietary alternative.


3.6 Time-to-Market Becomes a Moat

This temporal asymmetry deserves elevation into a principle, because it explains the otherwise puzzling willingness of Nvidia’s potential challengers to adopt its fabric. For a hardware startup—or even a hyperscaler—the accelerator is only the beginning of the bill. Building the accelerator, the memory subsystem, the scale-up switching, the rack architecture, the liquid-cooling integration, the power delivery, the management software, the validation regime, and the global supplier network is dramatically harder, slower, and more capital-intensive than designing the chip, and every quarter spent building infrastructure is a quarter of foregone revenue in a market growing at triple-digit rates. NVLink Fusion converts Nvidia’s fifteen years of accumulated systems experience into a product that competitors can buy, and in doing so it weaponizes time itself: the startup that adopts NVLink Fusion ships in 2027, while the startup that builds independently ships—perhaps—in 2029, into a market whose standards will by then have hardened further around the incumbent. Open specifications, however excellent, cannot by themselves reproduce supply chains, validation histories, or deployment ecosystems; those must be built in calendar time, and calendar time is precisely what the AI boom refuses to grant. Time-to-market has thus become a moat as real as any patent, and NVLink Fusion is the mechanism by which Nvidia rents that moat to its rivals—keeping them alive, keeping them close, and keeping them inside.


3.7 The Coming Dual-Fabric World

Prudence requires developing the possible outcomes for 2027 through 2030 as scenarios rather than predictions, and this paper’s analytical value lies in specifying what must occur for each to emerge.


Scenario A — Nvidia Fabric Dominance. Heterogeneous accelerators proliferate, but the overwhelming majority reach market as NVLink-compatible devices inside MGX-derived racks. Required conditions: NVLink Fusion’s partner roster continues compounding; UALink silicon arrives late or underperforms; hyperscalers conclude that dual-sourcing fabrics costs more than it saves; regulators decline to intervene structurally. In this world, Nvidia’s accelerator share erodes gracefully while its architectural share approaches totality, and the architecture toll becomes one of the great annuities in industrial history.


Scenario B — Dual Architecture. NVLink and UALink consolidate into two durable ecosystems: Nvidia’s semi-custom estate on one side, and an open federation—anchored by AMD, Google’s TPU pods, and portions of AWS, Meta, and Microsoft infrastructure—on the other. Required conditions: UALink switches and accelerator implementations ship at competitive bandwidth by 2027–2028; at least two hyperscalers commit flagship deployments to the open fabric; software abstraction layers mature enough to make dual-fabric operation tolerable. This is the world most analogous to historical precedent—two payment networks, two mobile platforms—and it preserves competition at the cost of fragmentation.


Scenario C — Open Fabric Commoditization. Open standards mature to the point that accelerator vendors move freely among systems, fabrics become interchangeable plumbing, and value migrates back to the silicon and up to the software. Required conditions: sustained multi-vendor governance discipline of the kind that carried Ethernet and PCIe to universality; a performance plateau that narrows proprietary advantages; and, plausibly, regulatory pressure toward interoperability. This is the world Nvidia’s strategy is explicitly designed to prevent, and its emergence before 2030 would require open-ecosystem execution of historic quality.


The paper does not assume Nvidia wins; it argues that the contest itself—not any single quarter’s benchmark—is the correct object of analysis, and Section 4 examines what happens when governments join it.


Section 4: The Political Economy of Fabric Hegemony


4.1 Hegemony Is More Than Market Share

Policy analysis of the semiconductor industry has traditionally been conducted in the vocabulary of concentration: market shares, HHI indices, merger thresholds, and the presence or absence of viable competitors in defined product markets. Fabric Hegemony requires a different and more demanding vocabulary, because the power it describes is categorically different from the power that concentration metrics measure. A processor monopoly, however large, can in principle be attacked by a faster processor; the history of computing is a graveyard of chip champions—DEC, Motorola, Intel in mobile—overthrown by exactly that mechanism. Architectural dominance is harder to dislodge, because it does not reside in any single product that a rival could out-engineer. It resides in the accumulated organization of an entire industry around a set of interfaces: customers whose facilities are physically built to the architecture’s power and cooling envelopes, suppliers whose factories are tooled to its mechanical forms, software whose assumptions encode its topology, engineers whose careers are trained to its idioms, and competitors whose own products are designed to plug into it. Each of these constituencies has independently invested in the architecture’s continuation, which means the architecture is defended not only by its owner but by everyone who has adapted to it. Overthrowing a chip requires a better chip; overthrowing an architecture requires coordinating an entire industry’s simultaneous migration—a collective-action problem of the highest order, and precisely the problem UALink exists to solve. This distinction, between share of sockets and control of interfaces, is the analytical foundation for everything that follows in this section, and it explains why a company could theoretically fall to fifty percent of accelerator sales while gaining power over the industry every single year.


4.2 When Interoperability Produces Dependence

The subtlest feature of NVLink Fusion, and the one most likely to confound conventional regulatory analysis, is that it simultaneously increases competition at one layer of the stack while consolidating control at another. At the accelerator layer, NVLink Fusion is unambiguously pro-choice: customers who once faced a binary decision between an all-Nvidia rack and a risky independent alternative can now select among Nvidia GPUs, Marvell- or MediaTek-designed XPUs, d-Matrix inference silicon, and their own custom chips, all within one validated environment. Customer choice expands; measured accelerator concentration falls; the market appears, by every traditional metric, to be getting healthier. Yet each exercise of that expanded choice deepens the industry’s commitment to the environment within which the choosing occurs. The more accelerator diversity flourishes inside NVLink Fusion, the more indispensable NVLink Fusion becomes, and the higher the cost of ever reconstituting that diversity on different rails. Interoperability, structured asymmetrically—many silicon vendors interoperating through one fabric vendor—produces dependence as its equilibrium outcome. This is not a novel pattern; it is the pattern of every successful platform from operating systems to app stores. But it has never before operated at the scale of hundred-billion-dollar annual infrastructure programs, and it presents policymakers with a genuinely hard question: how should competition law evaluate an arrangement that demonstrably increases choice in the product market while arguably foreclosing competition in the architecture market that traditional analysis does not even define?


4.3 The Antitrust Question

These implications are no longer theoretical, and the timing of their arrival could not have been more pointed. On September 9 and 10, 2026—the very days of the d-Matrix announcement—Reuters, Bloomberg, and the New York Times reported that the U.S. Department of Justice is investigating whether Nvidia structured its December 2025 arrangement with Groq, reported variously at $17 billion and $20 billion, specifically to avoid the antitrust review that an outright acquisition would have triggered: the deal took the form of a “non-exclusive license” to Groq’s inference-chip technology combined with the hiring of founder Jonathan Ross and other senior executives, a structure increasingly common across the AI industry precisely because it escapes automatic merger notification, and the Department has sent Nvidia a formal demand for information.[21][22] The investigation joins a broader current of scrutiny: Vanderbilt Law School antitrust professor Rebecca Haw Allensworth, commenting on Nvidia’s intertwined financial relationships across the AI ecosystem, identified the core structural concern that arises when the dominant supplier holds stakes in its own customers and rivals:

“They’re financially interested in each other’s success.” — Rebecca Haw Allensworth, Professor of Law, Vanderbilt University [32]

—a condition that, she warned, creates incentives to differentiate terms among competitors in ways ordinary arm’s-length supply relationships do not. Harvard senior fellow Paulo Carvão has similarly characterized the Nvidia-centered infrastructure alliances as the largest such project in history while flagging the governance questions that scale of that kind necessarily raises,[32] and legal scholarship has begun arguing explicitly that traditional antitrust tools are poorly matched to markets where lock-in operates through software ecosystems and interface control rather than through pricing.[33] This paper deliberately avoids assuming illegality anywhere in this pattern; the Groq inquiry may conclude without action, and every individual NVLink Fusion agreement is voluntary and plausibly efficiency-enhancing. The larger policy question stands regardless of any single case’s outcome, and it deserves to be stated in its most general form: at what point does beneficial integration become architectural foreclosure? Competition policy has spent a century learning to answer that question for railroads, operating systems, and payment networks. It has not yet answered it for fabrics, and the window in which the answer will matter is measured in years, not decades.


4.4 The Standard-Setting Question

Answering it will require regulators and legislators to adopt a lens that semiconductor policy has rarely used. The traditional apparatus examines manufacturing capacity, market share, mergers, and export controls—the visible mass of the industry. Fabric Hegemony demands attention to a different set of questions, quieter but arguably more consequential: Who controls the technical interfaces through which components must communicate? Are the specifications open, and on what licensing terms, with what governance over their evolution? Can competing accelerators interoperate with the dominant fabric without systematic disadvantage in bandwidth, latency, features, or time of access? Can customers migrate between fabrics at a cost that preserves the credibility of the threat to do so—since it is the credibility of exit, not its exercise, that disciplines an architecture’s owner? And do vertically integrated rack architectures impose switching costs that function, economically, as exclusivity even when no contract requires it? These are the questions standard-setting bodies, procurement offices, and competition authorities know how to ask about telecommunications and payments; the contribution of this paper’s framework is simply the insistence that AI infrastructure has crossed the threshold of systemic importance at which they must be asked here too.


4.5 China and the Fabric Chokepoint

The same lens transforms the technology dimension of U.S.–China competition. Export-control policy since 2022 has concentrated overwhelmingly on three visible chokepoints: advanced GPUs, semiconductor manufacturing equipment, and high-bandwidth memory. Those controls treat the accelerator as the strategic object. But if the argument of this paper is correct—if the economically meaningful system now includes scale-up switching, networking silicon, DPUs, optical interconnect and co-packaged optics, custom memory interfaces, and rack-level system software—then the strategic surface is far broader than the chip, and both the opportunities and the risks of control policy migrate accordingly. A future in which China fields competitive domestic accelerators but cannot reproduce rack-scale fabric performance is a future in which the effective compute gap persists even after the chip gap closes; conversely, a control regime blind to fabrics could watch its chip-level restrictions be substantially offset by system-level engineering. It is therefore reasonable to expect the export-control debate of 2027–2030 to migrate from chip restrictions toward architecture restrictions—interconnect switches, optical engines, fabric management software—with all the attendant complications that follow when the controlled object is a set of interfaces rather than a device. Notably, Nvidia’s own current guidance assumes zero data-center compute revenue from China, a fact that simultaneously demonstrates the reach of existing controls and frees the company to optimize its architecture strategy for the non-Chinese world.[7]


4.6 Sovereign AI Without Sovereign Fabric

Governments outside the two superpowers face their own version of the question. Dozens of states are now funding “sovereign AI” programs—national compute clusters, subsidized datacenters, domestic model initiatives—on the theory that intelligence infrastructure is too strategically important to rent entirely from foreign platforms. The framework of this paper suggests an uncomfortable audit question for every such program: what, exactly, is sovereign about it? A national AI factory whose accelerators, scale-up fabric, networking, rack architecture, management software, and supply chain are all proprietary foreign technology is sovereign in geography and in electricity, but in nothing else that matters; its capabilities, costs, upgrade paths, and continuity of operation remain functions of decisions made in Santa Clara, Washington, or elsewhere. This yields a distinction that governments would do well to internalize before, rather than after, the concrete is poured: compute sovereignty is not architecture sovereignty. Owning the machine is not owning the design of the machine, and only the latter confers the durable autonomy that sovereignty rhetoric promises. Open fabrics such as UALink are, among other things, an instrument by which smaller states and blocs could pursue architecture sovereignty collectively—which is why the standards fight examined in this paper is also, quietly, a fight about the future distribution of technological self-determination among nations.


4.7 The Governor and Federal Policy Agenda

Translated into the practical agenda of U.S. policymakers heading toward the November 2026 midterms and beyond, the framework generates concrete questions that move semiconductor policy decisively beyond the fab. Should federal procurement—which will fund substantial AI capacity for defense, energy, and civilian agencies, including the 100,000-GPU secure government AI factories now being co-developed by AWS and Nvidia[12]—favor interoperable infrastructure, or at minimum require documented migration paths, as a condition of award? Should CHIPS-style incentives, designed for a world where the fab was the chokepoint, be extended to networking silicon, optical interconnect, and advanced packaging, which are the chokepoints of the fabric era? Should the national laboratories deliberately maintain operational capability on multiple accelerator fabrics, treating fabric diversity as a strategic reserve in the way fuel diversity is treated in energy policy? Should critical national AI systems be subject to portability requirements that keep the exit threat credible? And should any national-security-critical AI capability be permitted to depend architecturally on a single vendor, however excellent? None of these questions presumes hostility to Nvidia, whose engineering achievements this paper has documented at length; all of them presume that a technology this central to national capability deserves the same structural attention that nations have always paid to their rails, their grids, and their networks.


Section 5: 2027–2030 — The Fabric Becomes the Market


5.1 Inference Changes the Economics

Training created the first GPU boom; inference will create the larger infrastructure economy, and the difference between the two regimes rewrites what “winning” means. Training demand, however vast, is ultimately bounded by the number of frontier laboratories and the cadence of frontier runs. Inference demand is bounded only by the world’s appetite for intelligence delivered as a service, and every indicator of 2026 suggests that appetite is effectively unbounded at current prices: billions of users, agents performing reasoning, coding, search, simulation, video generation, and enterprise workflows continuously rather than episodically, and token volumes compounding at rates that have repeatedly embarrassed forecasts. Nvidia’s own results narrate the transition—the company now speaks of tokens as productive and profitable output, of compute as revenue, and its CFO Colette Kress was compelled to address the systemic strain directly, telling investors:

“Memory scarcity today is being driven in large part by the AI buildout itself.” — Colette Kress, CFO, Nvidia [34]

The company’s own quarterly commentary records the composition of the surge—hyperscale revenue more than doubling year over year, and revenue from AI clouds, industrial users, and enterprises rising 138 percent—while its supply commitments, swollen to roughly $279 billion and dominated by memory for the Vera Rubin ramp, show the buildout already contracted years forward.[37][38] In an inference-dominated economy, the winning architecture is not the one that posts the highest benchmark but the one that produces the cheapest useful token at the required latency and within the available power—a full-system property in which fabric efficiency, memory movement, and orchestration weigh as heavily as raw arithmetic. This is why Nvidia claims one-tenth the cost per million tokens for Vera Rubin versus Blackwell, why d-Matrix aims Raptor at a premium latency tier priced above commodity tokens, and why every serious participant now quotes economics in dollars and joules per token rather than in flops.[1][9] The market’s unit of account has changed, and the fabric sits inside the new unit.


5.2 Specialized Accelerators Multiply

Extrapolating the specialization already visible in 2026, the AI factory of 2030 will plausibly contain not one class of processor but seven or eight, each optimized for a distinct region of the workload space: training accelerators tuned for sustained dense throughput; reasoning accelerators optimized for long-context, memory-heavy chain-of-thought serving; ultra-low-latency inference processors of the Groq LPX and d-Matrix Raptor type for interactive premium tiers; video and world-model accelerators with specialized codecs and bandwidth profiles; robotics and physical-AI processors operating under real-time constraints; memory-centric near-data processors; CPU fleets—of the Vera class—handling agentic orchestration, tool execution, and sandboxing; and network-compute devices performing collective operations inside the fabric itself. Every additional processor class multiplies the combinatorial complexity of making the factory coherent, and therefore raises the value of the one element all classes share: the fabric that binds them. Heterogeneity, in other words, is not a threat to fabric power; it is the source of it. A homogeneous world needs only a fast wire; a heterogeneous world needs a constitution—and constitutions are precisely what NVLink Fusion and UALink are competing to be.


5.3 The Fabric as an Operating System for Physical Compute

A conceptual analogy sharpens what is at stake. CUDA’s historical function was to abstract the GPU: it gave developers a stable programming target while Nvidia revolutionized the silicon beneath it, and that separation of interface from implementation is precisely what made the interface so valuable and so durable. The fabric is now assuming the same role one level up. A rack-scale architecture that presents heterogeneous processors, disaggregated memory, and tiered storage as one schedulable, addressable machine is performing the classical functions of an operating system—resource allocation, isolation, communication, abstraction—for physical compute at datacenter scale. Whoever defines that abstraction inherits the strategic position that operating systems have always conferred: applications (here, models and serving stacks) are written to it, hardware (here, accelerators and memory) is certified against it, and both sides of the market meet through it. NVLink Fusion plus MGX plus Nvidia’s system software is, on this reading, a candidate operating system for the AI factory, with the Fusion partner roster as its device-driver ecosystem; UALink plus open management standards is the rival candidate, the Unix to Nvidia’s proprietary incumbency. The analogy also carries the appropriate warning for the incumbent, since the history of operating systems includes both entrenchment of astonishing durability and, occasionally, the sudden victory of an open alternative that the incumbent dismissed for years.


5.4 Agentic AI Makes Communication More Important

The rise of agentic AI amplifies every argument above, because agents are, computationally, communication made flesh. A single agentic task decomposes into repeated model calls, tool invocations, code execution, memory retrievals, planning steps, and verification passes, each touching different resources: GPUs for generation, CPUs for tool execution and sandboxing, storage tiers for persistent memory, KV caches migrating between prefill and decode pools, and networks binding all of it into a latency budget the user experiences as a single response. Nvidia’s own positioning of the Vera CPU—explicitly marketed for the code execution, tool use, sandboxing, analytics, and orchestration behind agentic AI—demonstrates how thoroughly the company has internalized this shift, as does the appearance of dedicated context-memory storage platforms extending GPU memory into NVMe for KV-cache data.[12][15] As intelligence becomes distributed and persistent rather than monolithic and stateless, the ratio of data movement to arithmetic in the total cost of a delivered task rises structurally, and the economics of the tightly coordinated rack—where movement is cheap—strengthen against loosely coupled alternatives. Agentic AI is thus not merely another workload for the fabric; it is the workload that makes the fabric the primary determinant of end-user economics.


5.5 Energy Makes Fabric Efficiency Strategic

Return finally to Layer 1 of the Five-Layer AI Economy, because energy is where the fabric’s economics become national economics. Every unnecessary movement of data consumes energy twice over—once in the transfer itself and once in the idle compute that waits for it—and at the campus scales now under construction, those joules aggregate into utility-scale quantities. Meta alone plans fourteen gigawatts of computing capacity by 2027; hyperscaler capex guidance for 2026 approaches three-quarters of a trillion dollars, much of it ultimately constrained not by capital but by interconnection queues, turbine lead times, and transmission; and Microsoft has disclosed order backlogs it cannot serve for want of power.[5][25][26] In such a world, a fabric that raises useful computation per megawatt by even fifteen or twenty percent is functionally equivalent to commissioning gigawatts of new generation—without permits, turbines, or transmission lines—and fabric efficiency graduates from an engineering virtue into an instrument of energy policy. This is also the deepest economic defense of extreme co-design: Nvidia’s claim of up to ten times throughput per watt improvements at the rack level, if realized even in part, represents avoided power plants, and nations rationing grid capacity will notice.[15] The politics of AI energy consumption, already sharpening, will eventually discover the fabric; when they do, interconnect architecture will be debated in energy ministries as well as in standards bodies.


5.6 The AI Factory Becomes Modular

Synthesizing the trajectories above yields the plausible architecture of 2030: a modular AI factory in which Nvidia GPUs run frontier training and general inference beside specialized third-party XPUs serving premium latency tiers, hyperscaler custom accelerators absorbing internal baseline workloads, CPU fleets orchestrating agents, and network-compute devices operating inside the fabric—all housed in standardized rack and networking systems, provisioned by a common supply chain, and scheduled by common software. Modularity of this kind is unambiguously good for buyers at the component level: it disciplines every socket with the threat of substitution. But modularity requires a standard to be modular against, and this is the pivot on which the entire decade turns. If the standard is Nvidia’s—MGX racks, NVLink Fusion scale-up, Spectrum-X scale-out, NVHBM memory interfaces—then accelerator pluralism paradoxically strengthens Nvidia, because every substitution event occurs on Nvidia’s terms, inside Nvidia’s architecture, generating Nvidia’s toll. If the standard is open, the same pluralism disciplines Nvidia along with everyone else. The AWS arrangement of August 2026—Trainium and Nvidia GPUs sharing a common rack-scale architecture through NVLink Fusion and NVHBM, while AWS simultaneously holds a UALink board seat—shows the world’s most sophisticated infrastructure buyer refusing to prejudge that question, purchasing modularity now while preserving constitutional optionality for later.[12][29] Every major buyer will face the same choice, and the aggregate of their choices will decide which scenario of Section 3.7 becomes history.


5.7 Nvidia’s Ultimate Strategic Test

The analysis of this section resolves into a single formulation of Nvidia’s position. The question confronting the company is no longer whether it can defeat every accelerator competitor; the custom-silicon programs of its own largest customers have settled that it cannot and need not. The question is whether Nvidia can convince accelerator competitors—and the customers who fund them—that joining Nvidia’s architecture is durably more attractive than escaping it: that the time-to-market advantage, the supply-chain leverage, the co-design performance, and the customer proximity offered inside the estate exceed the pricing freedom, bargaining power, and sovereignty available outside it. Every NVLink Fusion signing is evidence that the persuasion is working; every UALink milestone is evidence that the counter-persuasion is organizing; and the balance between them, compounded over the deployment decisions of 2027 through 2030, is the strategic test of Fabric Hegemony. Empires of the socket are lost to faster chips. Empires of the architecture are lost only when the governed conclude, together, that the toll exceeds the service—and Nvidia’s entire fourth act is a wager that it can price the toll, forever, just below that threshold.


Section 6: What Have We Learned? Seven Pillars


Pillar 1 — The Accelerator Is No Longer the Entire AI Computer

Artificial intelligence has moved decisively from chip-level competition toward system-level competition. CPUs, accelerators, memory, switches, DPUs, optical links, networking, cooling, and software now operate as one economic machine, co-designed and co-priced, and the published architecture of Vera Rubin NVL72—six chips engineered as one rack-scale computer—is the incumbent’s formal acknowledgment of that reality.[8][14] The practical corollary is severe for challengers: a faster accelerator can underperform commercially, and often will, if it cannot integrate efficiently into a scalable infrastructure system, because the customer no longer buys arithmetic—the customer buys delivered tokens per dollar per watt, and that quantity is manufactured by the system, not the chip.


Pillar 2 — Nvidia Can Potentially Lose Share Without Losing Control

The d-Matrix agreement demonstrates the central paradox of this paper in commercial miniature: Raptor will compete with GPUs for inference workloads while entering the market inside a rack organized around Nvidia CPUs, Nvidia switches, Nvidia DPUs, Nvidia NICs, Nvidia networking, and Nvidia’s rack standard.[1][3] Accelerator competition and platform competition are therefore distinct games that must be analyzed separately, scored separately, and regulated separately. Losing the accelerator is not equivalent to losing the architecture—and a firm that understands the difference can deliberately trade the former for the latter, exchanging socket share it was destined to lose for architectural adoption it could not otherwise have purchased.


Pillar 3 — Open Standards Have Become Strategic Infrastructure

UALink and its companion interoperability initiatives are no longer obscure engineering projects conducted in conference rooms; they are attempts, backed by the balance sheets of Alibaba, AMD, Apple, AWS, Cisco, Google, HPE, Intel, Meta, and Microsoft, to prevent any single company’s fabric from becoming the default constitutional architecture of artificial intelligence.[29][30] The battle over AI infrastructure will consequently be fought in standards bodies as much as in fabs, and the outcome will be determined less by the elegance of specifications than by the speed at which open silicon ships, the discipline of consortium governance, and the willingness of at least two major buyers to anchor flagship deployments on the open side. Standards, in this decade, are industrial policy conducted by other means.


Pillar 4 — Fabric Control Connects All Five Layers of the AI Economy

Fabric architecture determines how efficiently chips become datacenters, how efficiently datacenters execute models, how cheaply models serve applications, and how much electricity the entire process consumes. Its effects propagate vertically—Energy → Chips → Datacenters → Models → Applications and Agents—which is precisely why control of it confers influence disproportionate to its visible revenue. The fabric is not a hidden technical detail; it is connective economic infrastructure, and the margin structure of the 2030 AI economy will be written substantially in its interfaces.


Pillar 5 — Capital Has Become an Architectural Weapon

The pattern of 2026—$2 billion into Marvell, $3.5 billion into MediaTek, a reported $17–20 billion licensing-and-talent arrangement with Groq, alongside $100 billion committed to OpenAI and multi-billion-dollar positions across the ecosystem—demonstrates that in a contest over architectural adoption, the balance sheet is itself a strategic instrument.[17][19][21][32] Nvidia is not merely selling its fabric; it is financing its fabric’s adoption by the very firms best positioned to build alternatives, converting potential architects of rival constitutions into invested tenants of its own. Any analysis of fabric competition that examines technology while ignoring capital flows will systematically underestimate the incumbent, and any competition-policy framework that reviews acquisitions while ignoring licensing-plus-investment structures will systematically miss the mechanism.


Pillar 6 — Custom Silicon Without Fabric Independence Relocates Dependence Rather Than Ending It

The hyperscaler ASIC boom is real, funded, and technically successful—and yet Trainium’s adoption of NVLink Fusion and NVHBM shows that even the most advanced custom-silicon program can end up reaching memory and reaching its neighbors through the incumbent’s interfaces.[12][16] Sovereignty over the die is not sovereignty over the system. For hyperscalers, the lesson is that fabric strategy deserves board-level attention equal to silicon strategy; for governments funding sovereign AI, the lesson is sharper still, because compute sovereignty without architecture sovereignty is a flag painted on rented infrastructure.


Pillar 7 — The Next AI Monopoly Debate May Be About Architecture, Not Chips

Public policy has concentrated on GPU market share, semiconductor fabs, and export controls—the visible mass of the industry. The next debate, foreshadowed by the Justice Department’s September 2026 inquiry into the Groq arrangement, will be more subtle and more consequential.[21][22] Who controls the interfaces? Who defines compatibility, and on what terms may rivals achieve it? Who owns the reference architecture around which supply chains tool themselves? Who determines which components can participate efficiently and which are relegated to degraded modes? And who collects economic value even when someone else’s logo appears on the accelerator? Those questions—questions of interface power rather than product power—define the political economy of Fabric Hegemony, and the institutions that learn to ask them first will shape the answers.


Conclusion: When Winning the Architecture Matters More Than Winning Every Chip

The September 10, 2026 announcement between d-Matrix and Nvidia may eventually look less consequential than a new GPU launch, a trillion-dollar datacenter commitment, or another spectacular advance in frontier models. There were no gigantic revenue projections attached to it; the companies did not even disclose financial terms. There was no announcement that Nvidia had acquired d-Matrix, and there was no suggestion that d-Matrix had abandoned its effort to build an accelerator capable of competing with conventional GPU-based inference. By the standards of a year in which Nvidia reported a $96 billion quarter, guided to $108 billion, and watched four customers commit three-quarters of a trillion dollars to infrastructure, a startup’s rack-integration roadmap barely registers as news.[7][25]

That is precisely why the announcement matters.

d-Matrix remains an accelerator company. Raptor remains an alternative architecture, built on the conviction that in-memory compute serves the token economy better than the GPU does. Its purpose is not to reproduce an Nvidia GPU but to approach inference differently—and yet Raptor is being designed to enter the Nvidia AI factory through NVLink Fusion and to operate within an environment containing Nvidia CPUs, Nvidia switches, Nvidia DPUs, Nvidia networking, and Nvidia rack infrastructure, manufactured through Nvidia’s supply chain and sold into Nvidia’s installed base.[1][2][3] The arrangement therefore demonstrates a strategic possibility that would have seemed counterintuitive, almost paradoxical, during the first years of the generative-AI boom, when every analysis of Nvidia’s future reduced to a single question about whether anyone could build a better chip.

Nvidia does not necessarily have to defeat every new AI accelerator.

It could connect them.

That distinction will grow more important, not less, as artificial intelligence moves from an era dominated by enormous homogeneous training clusters toward a heterogeneous economy of training, reasoning, inference, agents, robotics, simulation, and physical AI. Different workloads will create genuine and durable incentives for different processors. Hyperscalers will continue developing proprietary accelerators, as Meta’s Iris production ramp and Amazon’s Trainium franchise already prove.[4][5] Frontier laboratories will pursue their own silicon, as OpenAI’s ten-gigawatt Broadcom program demonstrates.[6] Startups will keep inventing architectures optimized around memory, latency, and power. Governments will seek sovereign alternatives; semiconductor companies from Qualcomm to Marvell will pursue opportunities previously ceded to GPUs. The heterogeneity is coming regardless of anyone’s strategy, because it is driven by physics and economics rather than by preference. The only open question is what organizes it.

It would therefore be dangerous—for investors, for competitors, and for policymakers—to assume that Nvidia’s future depends primarily on maintaining today’s accelerator market share. The stronger strategic position, and the one Nvidia’s conduct throughout 2026 reveals it to be pursuing deliberately, is to make accelerator diversity compatible with, and therefore constitutive of, Nvidia’s infrastructure: to become the ground on which the fragmentation happens.

That is the deeper meaning of Fabric Hegemony.

Hegemony does not require eliminating every rival; historically, hegemonic systems have been powerful precisely because other participants voluntarily operate within them, finding participation more profitable than resistance. A reserve currency does not require its issuing country to manufacture every product traded in that currency; the currency’s power grows with every transaction it denominates, including transactions between the issuer’s competitors. An operating system does not need to write every application; each third-party application deepens the platform’s indispensability. A payment network does not need to sell a single item purchased through it; it needs only to remain the rails on which purchasing occurs. Likewise, an AI infrastructure company may not need to design every processor if it controls enough of the architecture through which those processors communicate, share memory, obtain networking, enter racks, and become economically productive. In each historical case, the hegemon’s genius lay in converting rivals into participants and participation into reinforcement—and NVLink Fusion, with its roster of partners that now includes Nvidia’s own fiercest silicon competitors and largest customers, is a textbook instantiation of that genius.[3][12][16][19]

There are important limits to the thesis, and intellectual honesty requires stating them with the same force as the thesis itself. Nvidia faces serious technical and commercial competition at every layer, from AMD’s rack-scale platforms to Google’s mature TPU estate. Hyperscalers possess enormous bargaining power, sophisticated silicon teams, and structural incentives to reduce supplier dependence that no partnership announcement extinguishes; their simultaneous membership in UALink while adopting NVLink Fusion is not confusion but strategy—the deliberate preservation of exit. UALink offers an explicitly open alternative for accelerator-to-accelerator communication whose continued maturation could meaningfully weaken the economics of proprietary scale-up fabrics, particularly if open silicon ships at competitive bandwidth within the 2027–2028 window.[10][31] Custom-silicon programs could evolve into complete infrastructure stacks rather than merely alternative accelerators, as OpenAI’s full-rack ambitions with Broadcom already hint.[6] And regulatory authorities on multiple continents are demonstrably awake: the Justice Department’s inquiry into the Groq structure, whatever its outcome, signals that the era in which fabric strategies escaped scrutiny simply because they were novel is ending.[21][22] Academic economics, meanwhile, counsels humility about the entire edifice: if Daron Acemoglu’s estimate that AI adds only about 0.7 percent to total factor productivity over a decade proves closer to the truth than Erik Brynjolfsson’s far more expansive view—in which each dollar of tangible AI investment catalyzes nine or ten in complementary intangibles—then the token economy underwriting every fabric, every rack, and every gigawatt will be smaller than its architects assume, and the contest described in this paper will be fought over a lesser prize.[35][36]

Fabric Hegemony therefore should not be interpreted as a prediction that Nvidia inevitably controls the future. It is a framework for identifying where the struggle for control moves next—and for recognizing that struggle when it appears in forms that traditional metrics were never designed to register.

The first AI infrastructure contest concerned access to GPUs. The second concerned the ability to construct gigantic datacenters around them. The emerging contest concerns the architecture that turns many different processors into one functioning intelligence factory—and it will be decided in interconnect specifications, rack standards, memory interfaces, consortium governance, procurement rules, and investment structures, terrain far quieter than product launches but far more durable in its consequences.

This is why I chose the title: Fabric Hegemony: When Nvidia Can Lose the Accelerator — and Still Control the Architecture That Connects the AI Factory. The title fits because it names an apparent contradiction that may define the next stage of semiconductor competition. Nvidia could lose portions of accelerator market share while expanding the reach of NVLink, MGX, Spectrum-X, BlueField, ConnectX, NVHBM, and the surrounding ecosystem. Competitors could win individual sockets while Nvidia remains embedded in the racks. Custom silicon could proliferate while the architecture becomes more standardized. Greater chip competition could, paradoxically, produce greater fabric concentration—diversity at the layer everyone watches financing concentration at the layer almost no one does.

Within the Five-Layer AI Economy, that possibility is profound. The company controlling the accelerator controls an important component of Layer 2. But the company controlling the architecture through which energy, chips, datacenters, models, and applications are transformed into useful intelligence potentially influences the economics of the entire stack—the watts consumed, the tokens produced, the margins earned, and the applications that become viable at each layer above.

In the coming era of heterogeneous AI computing, the most powerful company may therefore not be the one that manufactures every accelerator.

It may be the company that convinces everyone else’s accelerator where to plug in.


Footnotes and Endnotes:

[1] d-Matrix, “d-Matrix Adopts NVIDIA NVLink Fusion Rackscale Infrastructure for Ultra-Low Latency AI Inference,” PR Newswire, September 10, 2026. https://www.prnewswire.com/news-releases/d-matrix-adopts-nvidia-nvlink-fusion-rackscale-infrastructure-for-ultra-low-latency-ai-inference-302875104.html

[2] Techtroduce, “d-Matrix Adopts NVIDIA NVLink Fusion for Next-Gen Raptor XPUs” (on Raptor’s 3DIMC in-memory compute architecture), September 10, 2026. https://www.techtroduce.com/d-matrix-nvidia-nvlink-fusion-raptor-xpu/

[3] NVIDIA Blog, “d-Matrix Adopts NVIDIA NVLink Fusion for Rack-Scale XPU Deployment” (including Sid Sheth press-briefing remarks and the NVLink Fusion partner roster), September 10, 2026. https://blogs.nvidia.com/blog/d-matrix-nvlink-fusion/

[4] Converge Digest, “AWS Links Trainium Strategy with NVIDIA NVLink, NVHBM and Vera” (Trainium3 price-performance, >$225B Trainium commitments, $25B custom-chip run rate), August 2026. https://convergedigest.com/aws-nvidia-2-million-gpus-ai-infrastructure/

[5] Reuters, via CNBC, “Meta to put AI chip into production in September as it looks to double computing capacity,” July 9, 2026. https://www.cnbc.com/2026/07/09/meta-to-put-ai-chip-into-production-in-september-report.html

[6] OpenAI, “OpenAI and Broadcom announce strategic collaboration to deploy 10 gigawatts of OpenAI-designed AI accelerators,” October 13, 2025. https://openai.com/index/openai-and-broadcom-announce-strategic-collaboration/

[7] NVIDIA Newsroom, “NVIDIA Announces Financial Results for Second Quarter Fiscal 2027” (revenue $96.2B; Data Center $89.0B; 75.0% gross margin; Q3 guidance $108.0B; no China Data Center compute assumed; Jensen Huang remarks), August 26, 2026. https://nvidianews.nvidia.com/news/nvidia-announces-financial-results-for-second-quarter-fiscal-2027

[8] StorageReview, “NVIDIA Launches Vera Rubin Architecture at CES 2026: The VR NVL72 Rack” (NVLink 6 at 3.6 TB/s per GPU and 260 TB/s per rack; Vera CPU with 88 Olympus cores; NVLink-C2C at 1.8 TB/s; Spectrum-6 with co-packaged optics), January 2026. https://www.storagereview.com/news/nvidia-launches-vera-rubin-architecture-at-ces-2026-the-vr-nvl72-rack

[9] NVIDIA, “Vera Rubin NVL72” product page (rack-scale architecture; MGX third generation; 80+ ecosystem partners; one-tenth cost per million tokens versus Blackwell; Groq 3 LPX pairing). https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72

[10] UALink Consortium, “Specifications” (UALink 200G 1.0: 200G per lane, up to 1,024 accelerators per pod; subsequent In-Network Compute and management specifications). https://ualinkconsortium.org/specification/

[11] U.S. Senator Elizabeth Warren, letter to DOJ Antitrust Division re Nvidia investigation (estimates of ~90% high-end AI chip share and ~98% data center GPU share), September 6, 2024. https://www.warren.senate.gov/newsroom/press-releases/warren-throws-support-behind-department-of-justice-probe-into-ai-chipmaker-nvidia-underscores-need-for-comprehensive-investigation

[12] NVIDIA Newsroom, “AWS and NVIDIA to Deliver 2 Million Additional GPUs and Next-Generation Infrastructure for Agentic and Physical AI” (Vera CPUs on AWS; NVLink Fusion and NVHBM for next-generation Trainium; 100,000-GPU U.S. government AI factories), August 2026. https://nvidianews.nvidia.com/news/aws-and-nvidia-to-deliver-2-million-additional-gpus-and-next-generation-infrastructure-for-agentic-and-physical-ai

[13] Spheron, “NVIDIA Vera Rubin NVL72: Specs, Price & Why There’s No H300” (Rubin GPU 288 GB HBM4 at 22 TB/s; rack totals; production timing), September 2026. https://www.spheron.network/blog/nvidia-vera-rubin-nvl72-guide/

[14] SemiAnalysis, “Vera Rubin — Extreme Co-Design: An Evolution from Grace Blackwell Oberon,” February 2026. https://newsletter.semianalysis.com/p/vera-rubin-extreme-co-design-an-evolution

[15] Supermicro, “Supermicro Solutions Featuring NVIDIA Vera Rubin” (3.6 exaflops inference, 75 TB fast memory, up to 10× throughput per watt and one-tenth token cost versus Blackwell; CMX context-memory storage). https://www.supermicro.com/en/accelerators/nvidia/vera-rubin

[16] Amazon, “AWS and NVIDIA to deploy 2 million more GPUs for AI in 2027–2028” (re:Invent 2025 NVLink Fusion support in next-generation Trainium; Annapurna Labs and NVHBM), August 2026. https://www.aboutamazon.com/news/aws/aws-nvidia-2-million-gpus-ai

[17] NVIDIA Newsroom, “NVIDIA AI Ecosystem Expands as Marvell Joins Forces Through NVLink Fusion” ($2B investment; Jensen Huang remarks), March 31, 2026. https://nvidianews.nvidia.com/news/nvidia-ai-ecosystem-expands-as-marvell-joins-forces-through-nvlink-fusion

[18] Data Center Dynamics, “Nvidia invests $2bn in Marvell as part of wider NVLink Fusion deal” (Matt Murphy remarks; custom XPUs and NVLink Fusion-compatible scale-up networking; silicon photonics collaboration), March–April 2026. https://www.datacenterdynamics.com/en/news/nvidia-invests-2bn-in-marvell-as-part-of-wider-nvlink-fusion-deal/

[19] Tom’s Hardware, “Nvidia pours $3.5 billion into MediaTek — company will adopt NVLink Fusion for its custom AI accelerators,” August 31, 2026. https://www.tomshardware.com/tech-industry/artificial-intelligence/nvidia-pours-usd3-5-billion-into-mediatek-company-will-adopt-nvlink-fusion-for-its-custom-ai-accelerators

[20] Tom’s Hardware, “Why Nvidia just poured $2 billion into AI ASIC competitor Marvell — NVLink Fusion turns into soft ecosystem lock-in,” April 2026. https://www.tomshardware.com/tech-industry/nvidia-invests-2-billion-in-marvell-whose-biggest-clients-are-trying-to-replace-nvidia-chips

[21] Bloomberg, “DOJ Probes Nvidia’s $20 Billion License Deal With Groq on Antitrust Concerns,” September 10, 2026. https://www.bloomberg.com/news/articles/2026-09-10/doj-probes-nvidia-s-license-deal-with-groq-on-antitrust-concerns

[22] Reuters, “DOJ probes Nvidia’s licensing deal with AI startup Groq, NYT reports” ($17B non-exclusive license; hiring of Jonathan Ross; formal information demand), September 9, 2026; see also Axios, “DOJ investigates Nvidia’s deal with Groq,” September 10, 2026. https://www.axios.com/2026/09/10/doj-nvidia-groq-antitrust

[23] The Next Web, “Nvidia’s $2 billion Marvell bet is not an investment. It is a toll booth,” April 2026. https://thenextweb.com/news/nvidia-marvell-nvlink-fusion-ecosystem-lock-in

[24] I/O Fund, “AI Capex to Hit $1 Trillion — And Estimates Are Still Too Low” (H1 2026 capex of $301B across Microsoft, Meta, Amazon, Google; $732.5B 2026 guidance; Goldman Sachs $7.6T 2026–2031 baseline), August 2026. https://io-fund.com/ai-stocks/ai-capex-1-trillion-estimates-too-low

[25] Yahoo Finance / Quartz, “Meta, Microsoft, Amazon, and Alphabet are about to spend a shocking amount of money to dominate the AI era” (Goldman Sachs: $725B 2026 combined capex, up 77% from $410B; $5.3T FY2025–FY2030), June 2026. https://finance.yahoo.com/sectors/technology/article/meta-microsoft-amazon-and-alphabet-are-about-to-spend-a-shocking-amount-of-money-to-dominate-the-ai-era-115359575.html

[26] CNBC, “Tech AI spending approaches $700 billion in 2026, cash taking big hit,” February 6, 2026. https://www.cnbc.com/2026/02/06/google-microsoft-meta-amazon-ai-cash.html

[27] Mike Gualtieri (Forrester), quoted in Lumien, “Meta’s ‘Iris’ AI Chip Enters Production in September,” July 2026. https://lumienai.com/news/meta-iris-ai-chip-production-september-2026-mtia

[28] Sam Altman, quoted in CNBC, “Broadcom stock pops 9% on OpenAI custom chip deal, adding to Nvidia and AMD agreements,” October 13, 2025. https://www.cnbc.com/2025/10/13/openai-partners-with-broadcom-custom-ai-chips-alongside-nvidia-amd.html

[29] StorageReview, “UALink Consortium Finalizes 1.0 Specification for AI Accelerator Interconnects” (85+ members; board composition; 1,024-accelerator pods; 200 GT/s per lane), April 2025. https://www.storagereview.com/news/ualink-consortium-finalizes-1-0-specification-for-ai-accelerator-interconnects

[30] Peter Onufryk (UALink Consortium President), quoted in SDxCentral, “UALink Consortium releases 200G 1.0 specification for AI accelerator interconnects,” April 2025. https://www.sdxcentral.com/news/ualink-consortium-releases-200g-10-specification-for-ai-accelerator-interconnects/

[31] UALink Consortium, “UALink: An Open, High-Efficiency Scale-Up Fabric” (white paper on compound lock-in and decoupling accelerator and interconnect decisions), January 2026. https://ualinkconsortium.org/wp-content/uploads/2026/01/UALink_White_Paper_Publication_Candidate_FINAL_VERSION.pdf

[32] Rebecca Haw Allensworth (Vanderbilt Law School) and Paulo Carvão (Harvard), quoted in Mogin Law LLP, “Nvidia’s $100B Investment in OpenAI Raises Antitrust Eyebrows,” September 2025. https://moginlawllp.com/nvidias-100b-investment-in-openai-raises-antitrust-concerns/

[33] Jacqueline Rodriguez, “Too Big To Fail? NVIDIA’s Dominance and Rethinking Antitrust Enforcement in the AI Era,” Duke Undergraduate Law Review, February 2026. https://www.dukeundergraduatelawreview.com/online-journal/too-big-to-fail-nvidias-dominance-and-rethinking-antitrust-enforcement-in-the-ai-era

[34] Colette Kress (Nvidia CFO), quoted in CNBC, “Nvidia earnings takeaways: Huang forecasts 70% fiscal 2028 revenue growth, far above estimates,” August 26, 2026. https://www.cnbc.com/2026/08/26/nvidia-nvda-earnings-report-q2-2027-live-updates.html

[35] Daron Acemoglu (MIT), “The Simple Macroeconomics of AI,” NBER Working Paper 32487 (TFP effects on the order of 0.7% over ten years), 2024. https://www.nber.org/system/files/working_papers/w32487/w32487.pdf

[36] Federal Reserve Bank of Richmond, “Will AI Investments Pay Off?” (Erik Brynjolfsson, Stanford: ~$9–10 of intangible investment per $1 of tangible AI investment; survey of Acemoglu, Wharton, and related estimates), Econ Focus, Q1–Q2 2026. https://www.richmondfed.org/publications/research/econ_focus/2026/q1-q2_feature2

[37] NVIDIA CFO Commentary, Q2 FY2027 (Form 8-K, Exhibit; Data Center revenue detail, hyperscale and ACIE growth, China Hopper shipments below 1% of Data Center revenue), August 26, 2026. https://www.sec.gov/Archives/edgar/data/0001045810/000104581026000073/q2fy27cfocommentary.htm

[38] 24/7 Wall St., “NVIDIA Q2 2027: $96 Billion Quarter Fueled by 117% Data Center Surge” ($279B supply commitments; AWS commitment context), August 26, 2026. https://247wallst.com/cards/nvidia-q2-2027-earnings-nvda-01m0zw3hstde6rb2mt0kwj6d2c