Introduction: The Border That Became a Login

Imagine a customs officer standing beside a sealed container at a major port. Inside are racks of advanced accelerators, high-bandwidth memory, networking equipment, and cooling components. The officer can inspect the manifest, identify the exporter and consignee, compare serial numbers, and determine whether the shipment requires a license. This is the familiar architecture of export control. It is physical, documentable, and territorial. A product leaves one jurisdiction and enters another. The law follows the movement of the item.

Now imagine a different transaction. The accelerators never leave an authorized data center in an allied country. The hardware remains inside a guarded facility. No restricted chip crosses a border. Instead, a newly created subsidiary leases remote capacity, pays through an intermediary, connects through encrypted credentials, and gives engineers thousands of miles away the ability to train, fine-tune, or operate a powerful system. The useful capability has traveled, even though the machine has not. The border is no longer the port. The border is the login.

In June 2026, the world watched a live demonstration of what enforcement at that new border looks like. The U.S. Department of Commerce ordered a leading American model developer, Anthropic, to obtain a license before allowing any foreign national to access its two most capable frontier models, citing concerns that the systems could be jailbroken and used to discover software vulnerabilities at machine speed.[45,46,48] No chip moved. No container was opened. A letter was sent, and within hours two of the most advanced artificial-intelligence systems on Earth went dark for every customer worldwide, because the provider could not reliably distinguish nationalities across a global user base in real time.[45,47] The restriction was lifted weeks later, but the precedent was set: the export-control apparatus of the United States had reached past the semiconductor, past the data center, and directly into the API endpoint. That episode, examined in detail in Section 3, is the clearest evidence yet that the sanctions frontier has moved.

That change is the starting point of this paper. The first modern phase of artificial-intelligence sanctions focused on semiconductor chokepoints: advanced graphics processing units, semiconductor manufacturing equipment, electronic-design software, high-bandwidth memory, and the fabrication ecosystem needed to produce leading-edge chips. Those controls were logical because AI chips are unusually concentrated, expensive, identifiable, and essential. Early policy research correctly recognized that specialized accelerators were becoming foundational inputs into both AI training and deployment.[1,32] The United States then translated that logic into the October 2022 export controls and successive revisions intended to restrict the People’s Republic of China’s access to advanced computing and semiconductor manufacturing capability.[2,3,4]

But hardware control was never the final destination. A chip is an instrument. The strategic object is the capability produced by the chip. That capability may appear as a trained model, a downloadable set of weights, a remotely served API, a cloud-hosted inference endpoint, a coding agent, a cyber-defense system, a biological-design assistant, a robotics controller, or a network of autonomous software agents that can plan and act. Once intelligence is delivered as a service, the traditional distinction between exporting a product and providing access to a capability begins to collapse.

This collapse is visible in the regulatory language that emerged between 2025 and 2026. Current U.S. rules and guidance no longer speak only about boxes and destinations. They address model weights classified under ECCN 4E091, remote end users, ultimate-parent relationships, Infrastructure-as-a-Service access, customer due diligence, storage security, and whether restricted parties receive remote access to an algorithm trained on controlled commodities.[9,10,11,12,13] The legal regime is still incomplete and contested, but its direction is unmistakable. The sanctions perimeter is moving upward through the AI stack.

“Compute is detectable, excludable, and quantifiable.”

— Girish Sastry, Lennart Heim, Yoshua Bengio, Diane Coyle, Gillian Hadfield, and coauthors [18]

This movement does not mean that chips have become unimportant. The opposite is true. Chips remain the material foundation of the system, and physical controls can slow acquisition, raise costs, complicate scaling, and preserve bargaining leverage. Yet the experience of 2022–2026 also revealed the limits of hardware-centered policy. Restricted actors may use downgraded chips more efficiently, acquire hardware through intermediaries, rent cloud capacity abroad, divide training across sites, distill capabilities from stronger models, or adopt open-weight systems that can be downloaded and operated outside the provider’s control.[27,28,29] A sanctions strategy that stops at the chip can become a border fence built around only the first entrance.

Inference is the second strategic frontier because inference is where a trained model becomes an economic, scientific, military, and political service. Training creates a capability; inference distributes it. Training may occur a limited number of times, in highly visible clusters, over weeks or months. Inference can occur continuously across millions of users, enterprises, devices, laboratories, robots, and agents. It can be sold by subscription, embedded into software, routed across cloud regions, combined with proprietary tools, and delegated authority to execute transactions. The Stanford AI Index documents the widening commercial and institutional diffusion of AI systems across the economy: generative AI reached fifty-three percent global population adoption within three years of its commercial debut, faster than the personal computer or the internet, and is now used in at least one business function at seventy percent of surveyed organizations.[33,53] In the long run, the geopolitical value of AI will not be measured only by who trains the largest model. It will also be measured by who can call it, where they can call it, what tools it can reach, and what actions it is permitted to take.

The term Inference Blockade therefore describes a post-chip sanctions architecture that denies, limits, conditions, observes, or revokes access to usable machine intelligence. Its objects are broader than inference in the narrow technical sense. The term includes access to training compute, model weights, fine-tuning, API outputs, agent tools, and autonomous execution because these layers are economically and operationally connected. A country that cannot import a frontier chip may still obtain frontier outputs through an API. A company that cannot download model weights may still automate high-value work through a hosted agent. A military laboratory that cannot train a model may still use a commercial model to accelerate coding, intelligence analysis, simulation, or cyber operations. The sanctionable object has become the delivered capability.

This paper develops the concept through a Six-Gate Inference Blockade. The gates are the Silicon Gate, Compute Gate, Weight Gate, API Gate, Tool Gate, and Action Gate. Each gate controls a different form of access. Each has different intermediaries, evidence, technical enforcement mechanisms, economic costs, and evasion risks. Together, they form a continuum from the physical accelerator to the autonomous action.

The six-gate approach also prevents a common policy mistake: treating every layer as though it can be governed in the same way. A chip is scarce and durable. Cloud access is metered and reversible. Model weights are copyable and difficult to recall after release. An API is centralized but can be replicated through fraudulent accounts. Tools may be harmless individually but dangerous when assembled into a system. Autonomous actions require identity, authorization, auditability, and liability rules that are closer to financial controls than to customs law. A credible regime must fit the enforcement instrument to the object being controlled.

The United States is simultaneously pursuing two policies that appear contradictory but are better understood as complementary. It seeks to restrict adversarial access to strategic compute and models while exporting “full-stack” American AI packages — hardware, cloud services, models, applications, and standards — to allies and partners.[7,8] This is not simply containment. It is network construction. The objective is to make the American stack widely available inside trusted relationships while making access more conditional outside them. China is developing its own reciprocal logic. Reporting on July 21, 2026 indicated that Chinese authorities were considering tighter controls on advanced models, model weights, training data, chip designs, and foreign acquisitions of strategic AI technology; those measures remained under review rather than final at the time of writing.[31,54]

Europe occupies a different position. The European Union’s AI Act is not an export-control regime, yet its rules for general-purpose AI create documentation, risk-management, cybersecurity, incident-reporting, and market-access obligations that can influence which models enter the European market and under what conditions.[16,17] On August 2, 2026 — twelve days after this paper’s date of writing — the European Commission’s enforcement and penalty powers over general-purpose model providers become applicable, including fines of up to three percent of global annual turnover.[52] This matters because the future sanctions environment will not be built by export-control agencies alone. It will be assembled across trade law, market regulation, cybersecurity standards, cloud contracts, financial compliance, procurement, and private platform governance.

The central argument is that an Inference Blockade is already emerging, but in fragmented form. The challenge for policymakers is not whether to invent it from nothing. The challenge is whether the pieces can be designed into a coherent, proportional, enforceable, and democratically legitimate system before emergency incidents produce a rushed and indiscriminate digital iron curtain. The June 2026 frontier-model order — issued abruptly, complied with in hours, contested in public, and partially reversed within weeks — is a preview of what governance by emergency looks like.[46,47,48]

Section 1 explains how the policy frontier moved from semiconductor hardware toward delivered intelligence, and measures the extraordinary economic mass now stacked behind each gate. Section 2 develops the Six Gates and the Capability Custody Chain. Section 3 examines the evolving legal architecture in the United States, Europe, and China, including the first live application of export controls to a deployed frontier model. Section 4 studies evasion, attribution, distillation, open weights, shell companies, and the Remote Beneficiary Problem. Section 5 evaluates the consequences for chipmakers, hyperscalers, model laboratories, startups, allied governments, and the Global South. Section 6 distills the analysis into five pillars for a governable regime. The conclusion returns to the opening scene. The container still matters. But the most consequential shipment may now arrive without a ship.


Section 1: From Chip Embargo to Inference Blockade

Every durable policy framework begins by identifying the object of control. In twentieth-century strategic trade, governments controlled weapons, machine tools, nuclear materials, cryptographic equipment, and dual-use technologies. The object was usually physical or embodied in technical documentation. Artificial intelligence destabilizes this tradition because the strategic object is distributed across a chain of inputs and services. An accelerator is useful because it performs computation. Computation is useful because it trains or serves a model. A model is useful because it can be integrated into workflows. A workflow is useful because it changes decisions and actions. Control at one layer may be bypassed through access at another.


1.1 The Hardware Chokepoint

The modern AI chip-control era did not appear suddenly in 2022. Research published in 2019 and 2020 had already identified advanced AI chips and semiconductor manufacturing equipment as concentrated chokepoints. Saif Khan and Alexander Mann explained that specialized accelerators were critical to training and deploying advanced algorithms, while allied control over equipment, design tools, and fabrication created leverage unavailable in many other digital technologies.[1,32] This was an unusually attractive policy target. The leading supply chain depended on a small number of firms across the United States, Taiwan, the Netherlands, Japan, and South Korea. The hardware was expensive, traceable, and difficult to produce without sophisticated capital equipment.

The October 7, 2022 U.S. controls converted this structural insight into policy. They restricted certain advanced computing items, semiconductor manufacturing equipment, and support by U.S. persons for specified advanced-node facilities in China. The October 2023 revisions expanded and clarified the performance thresholds, destination scope, end-use rules, and anti-circumvention provisions.[2,3] Additional measures in 2024 addressed high-bandwidth memory, software, equipment, and entities supporting advanced semiconductor production.[4] The objective was not to stop all Chinese computing. It was to slow the scale and sophistication of systems relevant to military modernization, surveillance, weapons development, and frontier AI.

Chris Miller of the Fletcher School at Tufts University, whose history of the semiconductor industry became the intellectual companion to the 2022 controls, has continued to frame the underlying policy dilemma in access terms rather than purely technical ones. Discussing renewed proposals in late 2025 to license advanced accelerators to Chinese customers, he warned that the decision is fundamentally about counterparties, not silicon:[49]

“[We have to be] very careful when deciding which countries and which companies we sell these to.”

— Chris Miller, The Fletcher School, Tufts University [49]

Miller’s caution reflects how contested the Silicon Gate had become by 2026. In late 2025 and early 2026, Washington shifted toward case-by-case licensing that permitted limited sales of Nvidia’s H200-class accelerators to approved Chinese buyers, provoking an unusual bipartisan backlash in Congress, where legislators from both parties supported statutory guardrails such as the proposed GAIN AI Act to prioritize American purchasers of advanced compute.[49,54] The political fight over a single product line demonstrated a deeper truth: once the chip is the only lever, every licensing decision becomes an existential argument, because there is no finer instrument available. The gates described in this paper exist, in part, to give policy more than one dial to turn.

These controls imposed real costs. They complicated procurement, limited access to the most advanced accelerators, increased dependence on domestic substitutes, and forced Chinese developers to devote more attention to hardware efficiency. Miller himself has argued that the material gap remains wide, estimating in mid-2026 that the United States and Taiwan were on track to produce roughly thirty times more quality-adjusted AI accelerators than China, and observing that Chinese hyperscaler capital spending on AI infrastructure had remained comparatively flat since 2022 even as American firms multiplied theirs.[51] Yet scarcity also creates adaptation. Researchers and companies learn to optimize memory, networking, quantization, mixture-of-experts architectures, inference scheduling, and distributed workloads. A system built with constrained hardware may still deliver significant capability if software and organizational efficiency improve. The hardware chokepoint is powerful, but it is not static.

The phrase “Whack-a-Chip,” used by Ritwik Gupta, Leah Walker, and Andrew Reddie, captures this problem. Their analysis argues that hardware-centered controls can become a repeated effort to close one technical pathway after another while model developers adapt to alternative accelerators, modified thresholds, and more efficient training.[27] Other researchers have reached related conclusions: restrictions may slow the frontier, but over time they can encourage domestic substitution, open-model ecosystems, and organizational innovation.[28,29] A blockade that focuses only on the physical item risks confusing the most visible bottleneck with the entire capability chain.


1.2 Training Was the First Digital Extension

The first extension beyond hardware focused on large-scale training. Training is attractive to regulators because frontier runs consume enormous quantities of compute and can be more observable than ordinary software activity. Compute-governance research argued that cloud providers could identify workload type and approximate computational consumption without inspecting all customer data. Providers could act as securers, record keepers, verifiers, and, under appropriate legal authority, enforcers.[19]

The policy analogy was Know Your Customer in finance. Banks do not inspect every purpose behind every dollar, but they identify customers, monitor high-risk activity, retain records, and report suspicious patterns. Janet Egan and Lennart Heim proposed applying a similar model to frontier compute. They emphasized an important difference between hardware ownership and cloud access: digital compute can be metered, limited, and suspended.[20] This reversibility makes cloud control more precise than a permanent prohibition on a chip purchase, but it also creates pressure to turn hyperscalers into compliance intermediaries.

“Digital access to compute offers more precise controls.”

— Janet Egan and Lennart Heim [20]

By 2025, U.S. regulation had moved toward model-weight and cloud-customer obligations. The January 2025 AI diffusion framework introduced controls on certain closed model weights alongside advanced chips, license exceptions, data-center authorizations, and security requirements.[5] The framework was subsequently rescinded in May 2025 before full implementation, but its conceptual contribution survived: the government had formally treated advanced model weights as export-controlled technology and had begun designing rules for their storage, transfer, and security.[6] Current EAR provisions continue to classify specified model weights under ECCN 4E091 and apply licensing and review policies based in part on the headquarters and ultimate parent of the recipient.[9,10]

The training layer matters because a frontier model is a capital asset that can be copied, fine-tuned, distilled, and deployed repeatedly. A state may accept a foreign data center containing controlled chips if the chips are physically secure, the customer is known, the training activity is disclosed, and the resulting weights remain inside approved custody. This converts the licensing problem from “Where did the chip go?” to “Who used it, what did they train, and where did the resulting capability travel?”


1.3 Why Inference Became the Missing Layer

Training controls alone leave a large gap. A restricted actor may not need to train a frontier model if a commercial provider will serve the model remotely. The user can purchase outputs rather than own the system. For many applications — code generation, intelligence analysis, translation, image analysis, scientific search, target recognition, cyber assistance, or robotic planning — API access may provide substantial value without control of the underlying weights.

This is the service substitution problem. When ownership of a strategic asset is restricted, access may substitute for ownership. A country without a domestic satellite constellation can buy imagery. A company without a supercomputer can rent high-performance computing. A laboratory without a frontier model can call an API. The economic logic of cloud computing is precisely to separate capability from ownership. Export-control logic built around possession must therefore confront business models designed around remote use.

Current U.S. rules reveal this transition. January 2026 license-review conditions required applicants to identify remote IaaS users in specified jurisdictions or connected through specified ultimate-parent relationships; to maintain Know Your Customer procedures; to prevent unauthorized remote access; and, in certain cases, to ensure that model weights trained on controlled items are not transferred to undisclosed users.[11,13] The rules also refer to remote access to an algorithm trained on controlled commodities. That language is an institutional bridge between chip control and inference control.

Yet the bridge is incomplete. Some current provisions expressly distinguish training from ordinary API inference, and not every form of model access is prohibited. This is an important legal and policy boundary. The Inference Blockade described here is emerging, not comprehensive. It consists of selective rules, licensing conditions, KYC expectations, provider policies, security requirements, and — since June 2026 — at least one direct licensing order against a deployed commercial model.[45,48] The paper’s purpose is not to claim that the United States already operates a universal API embargo. It is to explain why the architecture is moving in that direction for high-risk users and capabilities.


1.4 The Three Meanings of Inference

The word inference has three meanings in this paper. The first is technical: a trained model processes inputs and generates outputs. The second is commercial: inference is delivered as a metered service through APIs, applications, or embedded products. The third is geopolitical: inference is the practical availability of machine capability to a user, organization, or state.

The third meaning is the most important. A government does not ultimately care whether an adversary owns a particular file format. It cares whether the adversary can achieve a capability: accelerate weapons design, discover vulnerabilities, automate influence operations, analyze intelligence, operate drones, or scale surveillance. Technical inference becomes strategic inference when it changes the user’s capacity to act.

This distinction explains why models alone are not the correct regulatory endpoint. Michael Riegler and Inga Strümke argue that security policy should target systems rather than isolated models because scaffolding, tools, coordination, and multiple small agents can reproduce capabilities that appear absent from any single model.[26] The model is one component of a sociotechnical system. A small model with source-code access, a shell, a vulnerability scanner, persistent memory, and parallel agents may outperform a larger model with no tools. Therefore, a regime that controls only parameter count, training compute, or benchmark performance may miss the assembled system.

Inference Blockade broadens the object from model to capability system. It asks five questions: Who is the real beneficiary? What capability is being delivered? Through which infrastructure and credentials? What tools or data amplify it? What actions can the system execute? These questions move policy beyond the false choice between controlling chips and controlling speech. They focus on operational access.


1.5 The Economic Mass Behind the Gates

Before mapping the gates in detail, it is worth pausing to measure what now stands behind them, because the scale of the AI economy in mid-2026 explains both why governments want control points and why every control point is contested. The infrastructure being governed is no longer a research curiosity. It is the largest coordinated private capital deployment in modern industrial history, and its quarterly financial disclosures have become de facto geopolitical documents.

Consider the most recent corporate results available at the time of writing. Nvidia reported record revenue of $81.6 billion for its first quarter of fiscal 2027, the quarter ended April 26, 2026, up eighty-five percent year over year, with data-center revenue of $75.2 billion growing ninety-two percent; the company guided to $91 billion for the following quarter while explicitly assuming zero data-center compute revenue from China.[38,39] That last detail deserves emphasis: the world’s most valuable semiconductor company now writes the world’s second-largest economy out of its forward guidance as a matter of regulatory hygiene, having recorded no data-center Hopper shipments to China in the quarter against $4.6 billion in the comparable period one year earlier.[39] Export policy is no longer a footnote in these filings. It is a line item.

“The buildout of AI factories … is accelerating at extraordinary speed.”

— Jensen Huang, Founder and CEO, NVIDIA [39]

The foundry layer tells the same story from Taiwan. TSMC reported second-quarter 2026 revenue of $40.2 billion, up roughly thirty-six percent year over year, with gross margin reaching 67.7 percent and a record quarterly net profit; the company raised its full-year 2026 revenue-growth outlook to slightly above forty percent in U.S. dollar terms, citing demand from artificial-intelligence customers, and confirmed that advanced packaging capacity — the CoWoS process that binds accelerators to high-bandwidth memory — remains the binding physical constraint on global AI deployment, with lead times measured in years rather than months.[40,41] Chairman and Chief Executive C.C. Wei summarized the demand environment for analysts in a single sentence:[41]

“Our conviction in the multi-year AI megatrend remains very high.”

— C.C. Wei, Chairman and CEO, TSMC [41]

Above the silicon, the buyers are spending at a pace without industrial precedent. First-quarter 2026 earnings disclosures showed Alphabet, Amazon, Microsoft, and Meta collectively planning roughly $725 billion in capital expenditures for 2026, up approximately seventy-seven percent from the record $410 billion of 2025, with the overwhelming majority directed to AI data centers, accelerators, custom silicon, and power.[42] Microsoft set calendar-2026 capital spending at roughly $190 billion; Alphabet guided to as much as $190 billion; Amazon projected approximately $200 billion; and Meta raised its range toward $135 billion and beyond, citing memory-price inflation and additional data-center costs.[42,43,44] Alphabet’s Chief Financial Officer Anat Ashkenazi told investors the company was experiencing what she described as:[43]

“unprecedented internal and external demand for AI compute resources.”

— Anat Ashkenazi, Chief Financial Officer, Alphabet [43]

Meta’s Chief Financial Officer Susan Li framed the same commitment in strategic rather than financial terms, telling analysts that the company’s priority was to invest its resources to position itself “as a leader in AI” even at the cost of near-term free cash flow.[44] Skeptics who questioned whether these outlays could ever be recovered were answered bluntly by the sell side; Jefferies analyst Brent Thill dismissed the pessimists in a widely quoted assessment:[42]

“The AI economy is healthy … The bear thesis is garbage.”

— Brent Thill, Analyst, Jefferies [42]

The academic measurement confirms the corporate one. Stanford’s 2026 AI Index recorded global corporate AI investment of $581.7 billion in 2025, an increase of roughly 130 percent over the prior year, with U.S. private AI investment of $285.9 billion — more than twenty-three times China’s reported $12.4 billion, even allowing that the Chinese figure understates state-directed guidance funds.[33,53] The same report found that estimated U.S. consumer surplus from generative AI tools reached $172 billion annually by early 2026, that the performance gap between the best American and Chinese models had narrowed to under three percentage points, and that more than half of national AI strategies adopted since 2024 originated in emerging economies pursuing sovereignty over the very access this paper analyzes.[33,53]

These numbers carry three implications for the argument that follows. First, the sanctions surface now sits atop trillions of dollars of committed capital, which means every gate closure has an immediate, measurable market price — Nvidia’s excluded China guidance is the Silicon Gate priced in dollars. Second, the concentration of spending among a handful of hyperscalers and one foundry means the intermediaries a blockade must recruit are few, wealthy, and deeply exposed to government decisions, which makes them simultaneously capable enforcers and reluctant ones. Third, the sheer velocity of adoption — half the world’s population using generative tools within three years — means access restrictions are no longer felt only by militaries and laboratories. They are felt by economies. Any Inference Blockade will be constructed inside, not outside, this economic mass.


1.6 The Sanctions Surface

The sanctions surface is the collection of points where a state or authorized intermediary can condition capability. At the bottom of the stack are fabrication equipment, accelerators, memory, and networking. Above them are data centers, cloud tenants, schedulers, and training jobs. Above those are checkpoints, weights, adapters, and distilled models. Then come APIs, applications, tool connections, and autonomous actions. Payment systems, identity providers, software repositories, and telecommunications networks intersect the stack horizontally.

A wider sanctions surface creates more options, but also more risks. Multiple gates allow targeted intervention when one gate is ineffective. They also create overlapping jurisdiction, compliance cost, surveillance pressure, and the possibility that private firms become unaccountable geopolitical governors. The question is not how to maximize the number of control points. It is how to identify the narrowest effective intervention.

The six-gate framework developed in Section 2 is designed for that purpose. It does not assume that every gate should be closed. In many cases, access should remain open. The framework provides a vocabulary for distinguishing a chip shipment from cloud training, a model download from an API call, and a harmless output from an autonomous action. Without those distinctions, policy will either underreach and leave obvious substitutes available, or overreach and treat all machine intelligence as contraband.


Table 1. Evolution of the control perimeter, 2020–2026

PeriodPolicy or research shiftPrimary gateStrategic implication
2020AI chips and allied semiconductor chokepoints identified as concentrated policy levers.[1,32]SiliconControl begins with scarce physical inputs.
2022U.S. controls advanced computing items, supercomputing end uses, and semiconductor manufacturing support to China.[2]SiliconTechnical thresholds and end-use rules become central.
2023Rules expand and researchers propose KYC for frontier cloud compute.[3,20]Silicon / ComputeRemote capacity becomes an export-control concern.
2024High-bandwidth memory and additional production inputs are controlled; model-weight security research expands.[4,21]Silicon / WeightsThe protected asset moves beyond the accelerator.
Jan. 2025AI diffusion framework controls certain closed model weights and creates data-center security conditions.[5]Compute / WeightsA trained model is treated as portable strategic technology.
May–Jul. 2025Diffusion rule is rescinded, but the AI Action Plan calls for stronger enforcement and full-stack allied exports.[6,7,8]All gatesSelective network access replaces a simple global tier map.
Jan.–Jun. 2026Remote-user, ultimate-parent, and algorithm-access license conditions take effect; case-by-case H200 licensing to China debated; NIST advances agent identity.[11,13,14,49]Compute through ActionThe post-chip sanctions architecture becomes visible.
Jun.–Jul. 2026Commerce briefly restricts foreign access to two deployed U.S. frontier models; EU GPAI enforcement powers approach; China consults industry on reciprocal model, weight, data, and chip-design controls.[31,45,46,52,54]Weights / APIsDeployed intelligence itself becomes a licensable export.

Table 2. Traditional chip embargo versus Inference Blockade

DimensionTraditional chip embargoInference BlockadePolicy consequence
Controlled objectPhysical accelerator or manufacturing itemUsable capability across hardware, compute, weights, APIs, tools, and actionsThe object changes form across the lifecycle.
BorderCustoms destination and in-country transferCredential, cloud region, repository, endpoint, tool permission, and agent authorityTerritorial law must govern remote access.
IntermediaryExporter, distributor, freight and consigneeChipmaker, data center, cloud, model provider, identity provider, and agent platformCompliance responsibility becomes distributed.
EnforcementLicense, inspection, seizure, entity listingKYC, capacity limit, secure custody, rate limit, revocation, scoped tools, action approvalService-layer controls can be granular and reversible.
Main evasionSmuggling, misclassification, transshipmentShell companies, resellers, fraudulent accounts, distillation, open weights, agent decompositionIdentity and behavior matter as much as location.

Section 2: The Six Gates of Inference

The Six-Gate Inference Blockade begins from a simple proposition: intelligence is not delivered at one point. It is produced and transmitted through a sequence. Each stage changes the form of the asset and the type of control that is technically possible. A chip can be counted. A training job can be metered. A weight file can be hashed. An API credential can be revoked. A tool permission can be scoped. An autonomous action can be authorized, logged, or stopped. The regulatory design should follow the changing form of the capability.


2.1 Gate One: The Silicon Gate

The Silicon Gate controls the physical accelerators and manufacturing ecosystem. Its instruments include export classifications, destination rules, end-use restrictions, entity listings, licensing, serial-number verification, performance thresholds, high-bandwidth-memory controls, and restrictions on semiconductor manufacturing equipment. This gate is the most mature because it fits established trade law.

Its strengths are concentration and materiality. Leading accelerators and advanced manufacturing tools are produced by a small group of firms. Large clusters require thousands of identifiable devices, substantial capital, advanced packaging, networking, and power. TSMC’s mid-2026 disclosures underline how narrow the physical funnel remains: the company controls roughly three-quarters of the advanced foundry market, and its CoWoS advanced-packaging lines — booked out into 2027 — are the single constraint that determines how many frontier accelerators exist in the world at all.[40,41] The Silicon Gate can therefore increase cost, delay scaling, and complicate maintenance. It is especially effective against actors trying to assemble very large, reliable clusters quickly.

Its weakness is substitution. Hardware can be smuggled, leased through intermediaries, assembled from lower-performing devices, or used more efficiently. Threshold-based rules also create incentives to design chips immediately below regulatory limits. Enforcement must evolve as architectures change, but constant revisions can create uncertainty for manufacturers and allied customers. Hardware controls are indispensable, but they are a foundation rather than a complete wall.


2.2 Gate Two: The Compute Gate

The Compute Gate controls access to processing capacity, especially through cloud and new cloud providers. The central objects are accounts, tenants, clusters, workload scale, geographic location, and customer identity. The enforcement tools include KYC, beneficial-ownership verification, logging, workload attestations, capacity thresholds, remote-user disclosure, location controls, and account suspension.

This gate is more flexible than the Silicon Gate. A provider can permit limited inference while blocking large-scale training. It can cap the number or class of accelerators, restrict regions, require enhanced due diligence, or terminate access if ownership changes. It can also preserve evidence about who used the infrastructure and when. The cloud is therefore both a route around hardware controls and a potential enforcement layer.[35]

The danger is centralized surveillance. Cloud providers host sensitive research, commercial secrets, and personal data. A system that requires providers to inspect customer workloads can chill legitimate innovation and create powerful databases of scientific activity. Heim and coauthors emphasize that governance should rely where possible on non-confidential metadata rather than customer content.[19] The principle should be visibility sufficient for enforcement, not general-purpose monitoring.


2.3 Gate Three: The Weight Gate

The Weight Gate controls the trained parameters that encode model capability. Weights are unusual strategic assets. They are enormously expensive to create but cheap to copy. Once distributed, they can be stored offline, fine-tuned, quantized, merged, or deployed without the original provider. Unlike API access, a released weight file cannot be reliably revoked.

RAND’s work on securing frontier model weights describes the problem as one of theft prevention, access control, confidential computing, cyber defense, and institutional security.[21] Current U.S. rules classify specified advanced model weights and impose conditions on their export, reexport, transfer, and storage.[9,10] These provisions recognize that weights are not merely software artifacts. At sufficient capability, they may represent a portable strategic asset.

The Weight Gate is difficult to define. Parameter count is an inadequate measure because architectures vary. Training compute is measurable but does not perfectly predict capability. Benchmark performance can be manipulated and changes rapidly. The 2025–2026 regulatory approach attempts to connect control to computational thresholds and comparative capability, but any threshold will require regular revision. A better long-term design may combine training compute, capability evaluations, system risk, and deployment context.[34]

The open-weight debate is especially difficult. Open models support research, competition, language diversity, local deployment, audit, and innovation. They also diffuse irreversibly. A policy that treats every open release as a security threat would entrench the largest closed providers and undermine scientific collaboration. A policy that treats openness as automatically benign would ignore the fact that weights can be adapted for harmful purposes without provider oversight. The correct question is not open versus closed in the abstract. It is what capabilities, safeguards, documentation, and release conditions are appropriate for a particular system.[36,37]


2.4 Gate Four: The API Gate

The API Gate controls remotely delivered model outputs. It is the most commercially developed gate because it coincides with ordinary platform management. Providers already authenticate users, meter tokens, set rate limits, enforce acceptable-use policies, investigate abuse, and terminate accounts. OpenAI’s threat reports show how a provider can identify and disrupt state-affiliated or malicious use of hosted models.[24,25] Anthropic’s 2026 report on alleged industrial-scale distillation campaigns shows the same gate being used to defend model capability and regional access restrictions.[23] And in June 2026, the API Gate ceased to be theoretical for government policy: the Commerce Department’s licensing order against two deployed frontier models — analyzed in Section 3.4 — was executed entirely at this layer, through the provider’s own access controls, in a matter of hours.[45,48]

API control is attractive because it is reversible and granular. A provider can block one user without banning an entire country, reduce rate limits, disable dangerous tools, require stronger verification, or permit lower-risk models while withholding more capable ones. It can respond to behavior rather than nationality alone. These characteristics make the API Gate potentially more precise than physical embargoes.

But the API Gate is also porous. Users can create fraudulent accounts, use resellers, rotate payment methods, access third-party wrappers, route through virtual private networks, or query several models to distribute their activity. Outputs can be harvested for distillation. If one provider blocks access, another may not. A national API restriction may therefore require coordinated identity standards and provider cooperation, raising questions about due process and global market fragmentation. The June 2026 episode also exposed a granularity problem in the other direction: when the government demanded nationality-based screening that the provider’s identity systems could not perform in real time, the only compliant option was global suspension — the bluntest possible outcome from the gate that was supposed to be the most precise.[45,47]

The API Gate also creates a profound political issue: private companies become the first-line arbiters of international access to intelligence. Their usage policies, risk classifications, and enforcement systems may determine which researchers, businesses, journalists, or governments can use frontier systems. Provider discretion can move faster than law, but it lacks the procedural protections of public regulation. An Inference Blockade cannot simply outsource foreign policy to terms of service.


2.5 Gate Five: The Tool Gate

The Tool Gate controls the external systems that amplify a model. A language model without tools generates text. The same model with access to a code repository, shell, browser, vulnerability scanner, laboratory system, financial account, customer database, or robotics interface can change the world outside the chat window.

This gate is increasingly important because agentic systems are assembled from connections. Model Context Protocol servers, plugins, function calls, enterprise connectors, and application permissions transform general language capability into domain-specific action. The security boundary is therefore not only the model endpoint. It is the authorization boundary around each tool.

A tool-centered approach can be more proportional than banning a model. A researcher may be permitted to use a frontier model for literature review but not to access a restricted biological-design tool. A foreign subsidiary may use a coding model against public repositories but not against source code for critical infrastructure. A general-purpose model may remain available while high-risk connectors require licensing, human approval, or verified institutional identity.

The Tool Gate is also where system-level policy becomes necessary. Riegler and Strümke’s 2026 analysis demonstrates that coordinated agents and scaffolding can produce capabilities not predicted by the base model alone.[26] Control must therefore evaluate the assembled system: model, prompts, memory, tools, data, parallelism, and evaluation loop. The sanctionable object may be the integration, not the model.


2.6 Gate Six: The Action Gate

The Action Gate controls what an autonomous or semi-autonomous system is authorized to do. This is the final and least developed gate. It includes deployment of code, transfer of funds, procurement, operation of machinery, execution of cyber actions, modification of infrastructure, and communication with other agents.

NIST’s 2026 work on software and AI agent identity emphasizes identification, authorization, auditing, non-repudiation, and controls against prompt injection.[14,15] These concepts are foundational to an Action Gate because an autonomous system needs a verifiable principal. Who authorized the agent? Which organization does it represent? What scope of authority was delegated? Which actions require human approval? Can the action be attributed and reversed?

AI agents can “autonomously perform tasks.”

— National Institute of Standards and Technology [14]

The Action Gate changes the nature of sanctions. Traditional sanctions prevent a transaction between people or companies. Agentic sanctions may need to prevent an autonomous system from executing a prohibited transaction, even when the human principal is distant or hidden. Financial institutions already screen counterparties and transactions. Future agent platforms may need machine-readable policy checks before executing sensitive actions.

This gate should be approached with caution. An overbroad action-control system could become a universal surveillance layer for software. The strongest justification exists where actions are high consequence, irreversible, or directly connected to controlled sectors. Low-risk personal productivity should not require state authorization. The principle is graduated authority: the greater the consequence, the stronger the identity, logging, and approval requirements.


2.7 The Capability Custody Chain

The six gates are connected by what this paper calls the Capability Custody Chain. The chain follows strategic intelligence from hardware manufacture to final action:


Figure 1. The sanctions perimeter moves upward from physical hardware to autonomous action.

GATE 1: SILICON  →  GATE 2: COMPUTE  →  GATE 3: WEIGHTS

GATE 4: APIs  →  GATE 5: TOOLS  →  GATE 6: ACTIONS

Accelerators → Cloud clusters → Trained parameters → Hosted outputs → Amplifying connectors → Autonomous execution

Each gate has different intermediaries, evidence, enforcement levers, and evasion risks.


Silicon manufacturer → data-center operator → cloud tenant → training workload → model checkpoint → weight custodian → API provider → authenticated user → agent identity → tool authorization → external action.

At each transition, the holder changes. Evidence can be lost, identities can be obscured, and responsibility can be divided. The physical owner of a chip may not know the ultimate user. The cloud tenant may be a subsidiary. The weight custodian may license an API to a reseller. The end user may delegate tasks to an agent. The agent may call tools operated by third parties. The chain is only as strong as its weakest identity and recordkeeping transition.

This does not imply universal traceability of every prompt. The objective is capability custody, not content surveillance. The relevant records are ownership, authorization, workload scale, model identity, endpoint access, tool permissions, and high-consequence actions. A well-designed regime would minimize content collection and retain only what is necessary for compliance, security, and investigation.


Table 3. The Six Gates of Inference

GateControlled objectPrimary intermediaryControl leverTypical evasion
1. SiliconAccelerators, memory, equipmentManufacturers and exportersClassification, license, serial and location verificationSmuggling, threshold design, third-country diversion
2. ComputeCloud clusters and workloadsHyperscalers and neocloudsKYC, capacity cap, logging, region restrictionResellers, shell tenants, distributed training
3. WeightsCheckpoints, parameters, adaptersLabs, repositories, custodiansClassification, secure storage, access control, provenanceTheft, mirrors, fine-tuning, model merging
4. APIsHosted outputs and model servicesModel providers and platformsAuthentication, rate limits, monitoring, revocationFraudulent accounts, wrappers, extraction, distillation
5. ToolsCode, data, browsers, lab and enterprise connectorsAgent platforms and tool operatorsScoped credentials, allowlists, approvalsTool substitution and multi-agent decomposition
6. ActionsTransactions and physical or digital effectsEnterprises, banks, infrastructure operatorsAgent identity, authorization, audit, human checkpointDelegation chains and cross-platform automation

Figure 2. The Capability Custody Chain tracks responsibility as intelligence changes form.

Silicon manufacturer → Data-center operator → Cloud tenant → Training workload

Model checkpoint → Weight custodian → API provider → Authenticated user

Agent identity → Tool authorization → External action

The chain is only as strong as its weakest identity and recordkeeping transition.


2.8 Training and Inference Are Different Control Problems

Training and inference should not be collapsed. Training creates a reusable capability and often requires concentrated compute. Inference distributes capability and may occur at much smaller scale per request. Training is easier to detect through aggregate compute but harder to interpret without workload context. Inference is easier to control through centralized APIs but harder when weights are open or locally hosted.

The distinction also affects proportionality. Blocking a massive frontier training run by a military-linked entity is different from blocking a student from using a translation model. Restricting the download of exceptionally capable weights is different from limiting access to a safety-filtered API. The Six Gates permit differentiated controls rather than a binary open-or-closed policy.

The next section examines how existing institutions are beginning to occupy these gates. The result is not yet a single blockade. It is a patchwork. But the patchwork already contains the legal vocabulary of the future: ultimate parent, remote user, model weights, IaaS, cybersecurity, systemic risk, agent identity, and revocable access.


Section 3: The Emerging Legal and Institutional Regime

The law of Inference Blockade is emerging from several different traditions. Export controls govern items, technology, software, destinations, end uses, and end users. Sanctions law governs transactions with designated persons and jurisdictions. Cloud regulation governs identity, cybersecurity, data, and service access. AI regulation governs models, systems, risks, and market placement. Corporate policies govern platform use. None of these systems was designed to carry the entire burden of regulating strategic machine intelligence. Their convergence is therefore uneven.


3.1 The United States: From Destination to Ultimate Parent

The U.S. system remains anchored in the Export Administration Regulations. The October 2022 and October 2023 rules targeted advanced computing and semiconductor manufacturing capability through technical thresholds, destination-based controls, end-use provisions, and the Foreign Direct Product framework.[2,3] These rules treated the chip and the facility as the central objects. The next regulatory phase added questions of ownership and remote access.

The concept of the ultimate parent is especially important. A customer may be incorporated in Singapore, the United Arab Emirates, Malaysia, or another commercial hub while being controlled by a company headquartered in a restricted jurisdiction. If the law examines only the immediate buyer, corporate layering can defeat destination controls. Current EAR provisions therefore consider the headquarters of the recipient and, in specified circumstances, the headquarters of its ultimate parent.[9,10,11]

This change creates what might be called corporate genealogy for compute. Providers and exporters must understand not only the legal entity signing the contract but the chain of ownership above it. Beneficial-ownership verification becomes a national-security function. The compliance problem begins to resemble anti-money-laundering practice, where shell companies, nominees, and complex ownership structures can hide the real beneficiary.

The January 2026 licensing conditions extended this logic to remote IaaS users. Applicants may need lists of remote users in named jurisdictions or entities connected through ultimate-parent relationships. Providers may need procedures to prevent unauthorized remote access and to ensure that weights trained on controlled hardware are not transferred to undisclosed users.[11,13] The physical destination of the accelerator remains relevant, but it no longer answers the decisive question.

The license conditions address parties receiving “remote access to any algorithm trained” on controlled commodities.

— U.S. Export Administration Regulations [11,13]

The policy significance of this phrase is larger than its immediate legal scope. It recognizes that algorithmic access can be a controlled benefit derived from controlled hardware. The same insight can be extended further: if an API provides a strategic capability trained on controlled infrastructure, the policy concern may follow the capability even after the training job is complete.


3.2 Model Weights as Controlled Technology

The classification of certain model weights under ECCN 4E091 is a major institutional innovation. It treats a trained model as technology capable of export, reexport, or in-country transfer. Current rules include storage and security conditions and apply a presumption of denial in specified circumstances for recipients outside listed destinations.[9,10]

This does not mean every model is controlled. The classification targets exceptionally capable models through technical criteria. The practical challenge is that capability evolves faster than rulemaking. A threshold calibrated to the frontier in one year may capture ordinary commercial models a few years later. Benchmark definitions can become obsolete. Training runs can be concealed or distributed. Models can be fine-tuned or merged after release.

A model-weight regime therefore requires a living measurement system. The government needs technical expertise, secure evaluation capacity, industry reporting, and mechanisms for rapid clarification. Classification requests under current EAR provisions indicate one possible process, but a sustainable system must also provide predictable safe harbors. Developers need to know when research publication, academic collaboration, red-team access, or ordinary deployment does not create a licensing obligation.

The deeper issue is that model weights combine characteristics of software, intellectual property, and strategic capital. They are not consumed when used. They can be copied at near-zero marginal cost. They may contain capabilities not fully understood by the developer. Their value depends on inference infrastructure and system integration. They are therefore difficult to fit inside legal categories inherited from physical trade.


3.3 The Rescinded Diffusion Rule and the Surviving Idea

The January 2025 AI diffusion rule attempted a broad global framework for advanced chips and certain closed model weights. It created destination groupings, license exceptions, data-center authorizations, and security requirements.[5] The rule was controversial because of its complexity, geographic tiers, compliance burden, and potential effects on U.S. cloud competitiveness. The Department of Commerce rescinded it in May 2025 and announced that a replacement approach would be developed.[6]

Rescission is not the same as abandonment of the underlying problem. The subsequent AI Action Plan called for strengthened export enforcement, location verification, diversion monitoring, and coordination with allies.[7] Current rules retained model-weight classifications and remote-access conditions. The institutional lesson is that the first comprehensive design may fail while its concepts migrate into narrower, more targeted instruments.

This pattern is common in technology policy. A broad proposal reveals the regulatory object, generates industry feedback, and is replaced by a sequence of specific measures. Inference Blockade is likely to evolve this way. Governments may avoid a single dramatic “AI services ban” and instead build capability controls through licenses, user screening, security certifications, procurement restrictions, cloud-region rules, and provider reporting.


3.4 The June 2026 Frontier-Model Order: A Live Test of the API Gate

In June 2026 the abstract became operational. In April of that year, Anthropic had released its most capable system, Mythos 5, not to the public but to a few dozen vetted cybersecurity vendors, major technology firms, and the U.S. government under a controlled-access program called Project Glasswing, after internal testing showed the model could discover exploitable software vulnerabilities — including flaws in widely used systems that had escaped human researchers for decades — at unprecedented speed.[48] A general-availability derivative, Fable 5, followed in early June with additional safeguards. Within days of that public release, and after external research described a jailbreak pathway that could redirect the model’s vulnerability-discovery capability, Commerce Secretary Howard Lutnick sent Anthropic a letter requiring a license for the export, reexport, or transfer of both models to any foreign national anywhere, including foreign nationals on U.S. soil.[46,48,52]

The company reportedly received ninety minutes to comply.[48] Because its identity systems could not verify nationality across a global customer base in real time, Anthropic disabled both models for every user worldwide — the largest single withdrawal of deployed AI capability in the industry’s history — while stating publicly that the underlying concern appeared to rest on a misunderstanding it hoped to resolve:[45]

“[W]e must abruptly disable Fable 5 and Mythos 5 for all our customers to ensure compliance.”

— Anthropic, corporate statement, June 2026 [45]

The order lasted roughly two and a half weeks. By June 30, 2026, the restrictions had been lifted: Fable 5 returned to general availability, and Mythos 5 remained gated behind Project Glasswing’s vetted-access structure for critical-infrastructure defenders while the company negotiated broader domestic and international access.[46] The legal basis of the directive was immediately contested — commentators questioned whether existing export-control authority reached a hosted service used by foreign nationals inside the United States — and the political context was inseparable from a broader dispute between the administration and the company over AI-safety positions and military use policies.[46,48]

Whatever one concludes about the order’s wisdom or lawfulness, it is the single most instructive event in the short history of the Inference Blockade, because it demonstrated every dynamic this paper describes within a single month. It showed that the API Gate can be closed by a letter rather than a statute, and closed in hours rather than years. It showed the granularity failure at the heart of nationality-based controls: absent verified identity infrastructure, a targeted restriction collapses into a global one, punishing allied hospitals, universities, and enterprises alongside the intended targets. It showed the credibility cost: analysts at the Council on Foreign Relations warned that abrupt, opaque restrictions on American models teach allies and customers that access to the U.S. stack is a political variable, undermining the very full-stack export strategy the same government promotes.[47] And it showed the substitution effect in real time: in the weeks surrounding the order, demand and investment surged toward open-weight alternatives that no government can switch off, including a record funding round for a leading Chinese laboratory.[55]

The episode also previewed the future division of labor between the Weight Gate and the API Gate. Mythos 5’s continued restriction through Project Glasswing is, functionally, a voluntary licensing regime for exceptional capability: identity-verified institutions, defined use cases, contractual conditions, and revocable access. That structure — capability tiering with verified custody, rather than blanket nationality bans — is much closer to the governable architecture proposed in Section 6 than the ninety-minute directive that preceded it. The lesson is not that model-level controls are impossible. It is that they must be designed in advance, with identity, process, and proportionality built in, because improvised versions will be simultaneously overbroad and short-lived.


3.5 Full-Stack Exports and Selective Blockade

The United States is not pursuing technological isolation. Executive Order 14320 and the American AI Action Plan promote exports of full-stack American AI packages that include chips, data centers, cloud services, models, cybersecurity, and applications.[7,8] The strategy is expansive: deploy the American stack globally, deepen allied dependence on U.S. technology, and compete with Chinese alternatives.

This produces a two-track system. Trusted partners may receive integrated packages with financing, security standards, and continuing vendor relationships. Restricted actors face heightened scrutiny or denial. The policy is therefore better described as selective network access than universal containment.

The model resembles earlier technology alliances but with greater operational dependence. A country that adopts an American AI stack may depend on U.S. chips, cloud management, model updates, security patches, and developer tools. This creates benefits: interoperability, reliability, and access to leading systems. It also creates leverage. Service access can be conditioned or withdrawn. The foreign-policy significance of the stack is not only what is exported, but the continuing relationship embedded in the service.

The same architecture can create anxiety among allies. A government may worry that a future U.S. administration could restrict updates, inference, or cloud access during a political dispute — a worry the June 2026 order converted from hypothetical to historical, as the European Commission opened an examination of the order’s practical consequences for users across the bloc within days.[45,47] Sovereign AI strategies are partly responses to this fear. An effective alliance policy must therefore combine security conditions with predictability, consultation, and due process.


3.6 Europe: Market Access Rather Than Blockade

The European Union’s AI Act approaches the problem from market regulation rather than strategic trade. The Act entered into force in 2024; general-purpose AI obligations began applying in August 2025; and major enforcement powers become effective in August 2026.[16,17] Providers of general-purpose models must prepare technical documentation, maintain copyright policies, publish information about training content, and provide information to downstream deployers. Providers of models with systemic risk face additional risk assessment, mitigation, incident reporting, and cybersecurity obligations.

The August 2, 2026 threshold deserves particular attention because it converts paper obligations into enforceable ones. From that date, the European Commission’s AI Office may request documentation, conduct technical evaluations of models, order corrective and risk-mitigation measures, restrict or withdraw a general-purpose model from the Union market, and impose fines of up to fifteen million euros or three percent of global annual turnover, whichever is higher; prohibited-practice violations carry penalties up to thirty-five million euros or seven percent of worldwide turnover.[16,52] The Digital Omnibus package agreed in spring 2026 deferred certain high-risk system obligations to late 2027, but the general-purpose enforcement machinery proceeds on schedule.[52] For the world’s model providers, the practical consequence is that the European market now contains a public authority with the legal power to do what only private terms of service and American emergency letters had done before: switch a model off.

These rules are not an Inference Blockade in the sanctions sense. They do not primarily deny adversaries access. Yet they create a model-governance perimeter around the European market. A provider that cannot or will not meet the obligations may face enforcement or lose practical access to EU customers. Market regulation can therefore condition capability without invoking national-security export law.

Europe also illustrates the importance of lifecycle governance. The obligations extend beyond training into documentation, deployment, incidents, and downstream integration. This is closer to the system-oriented approach needed for APIs and agents. The European framework may become a complementary layer in allied coordination, but differences in privacy, competition, open-source policy, and national security will make harmonization difficult.

The global regime may ultimately contain three overlapping logics: American strategic access, European market accountability, and Chinese technological sovereignty. Companies operating worldwide will have to comply with all three, sometimes simultaneously.


3.7 China: From Target to Rule Maker

China has long been the principal target of U.S. advanced-computing controls, but it is also a major rule maker. It has used export controls on critical minerals and technologies, cybersecurity review, data rules, procurement, and industrial policy to protect domestic capability. Reporting in July 2026 indicated that Chinese regulators were considering adding advanced AI models, training data, downloadable weights, chip designs, and strategic acquisitions to restricted technology categories.[31,54]

The July 2026 reporting is worth examining in detail because it sketches the first draft of a Chinese Weight Gate. According to accounts of the consultations, the Ministry of Commerce, working with the National Development and Reform Commission, held weeks of discussions with Alibaba, ByteDance, Zhipu, and other leading developers on two central questions: whether critical training data should be permitted to leave the country, and whether foreign users should continue to be able to freely download the weights of Chinese open models — while access through APIs and cloud services would remain available.[31,54] A tiered regime was reportedly floated: filing requirements for less capable open releases, security reviews for stronger systems, and possible prohibition of public release for the most capable models, with theft or leakage of proprietary AI technology potentially treated as a national-security offense.[54] Officials also sought views on preventing foreign foundries and chipmakers, including TSMC and Qualcomm, from manufacturing advanced semiconductors based on designs from Huawei, Alibaba, and ByteDance, and on screening foreign acquisitions of strategic AI firms; the measures could enter the next revision of China’s catalogue of technologies prohibited or restricted from export.[31]

The proposals were not final when this paper was written, and their timing was striking — they surfaced within days of high-level Chinese endorsements of open-source AI and international cooperation.[54] Their significance lies in symmetry. The United States seeks to prevent strategic U.S. capability from strengthening Chinese competitors and military systems. China increasingly faces the same concern as its models improve and its companies expand abroad. Chinese open-weight families — Qwen, GLM, DeepSeek’s releases — have become global commodity infrastructure precisely because they are freely downloadable; walling off the next generation would trade diffusion-based influence for capability retention, mirroring the American dilemma in reverse.[54] A world in which both powers restrict weights, data, designs, and acquisitions would mark the transition from semiconductor rivalry to reciprocal intelligence control.

China also promotes open models and AI cooperation with developing countries. This is not necessarily contradictory. A state can favor broad diffusion of selected models while protecting the most advanced weights, training data, or industrial know-how. The distinction between public diplomacy and strategic reserve will become central. Governments may release capable open models to build ecosystems while retaining their frontier systems behind controlled APIs — which is, notably, the exact structure the reported Chinese proposals would codify: open access to services, restricted custody of weights.


3.8 Private Governance as De Facto Foreign Policy

Governments do not operate the API Gate directly. Model providers do. OpenAI, Anthropic, Google, Microsoft, Amazon, Meta, xAI, and other companies control accounts, endpoints, rate limits, model versions, safety systems, and usage logs. Their decisions can deny access faster than a formal regulatory process.

OpenAI’s reports on state-affiliated threat actors and malicious campaigns show a private provider investigating and terminating accounts linked to cyber or influence activity.[24,25] Anthropic’s distillation report describes fraudulent account networks allegedly used to extract model capabilities in violation of terms and regional restrictions.[23] These actions resemble sanctions enforcement at the platform layer, even when they arise from private contracts rather than government designation.

Private enforcement has advantages. Providers possess technical visibility and can adapt quickly. But private rules are not a substitute for law. Companies may have commercial conflicts, inconsistent standards, limited appeal procedures, or incentives to over-block users in high-risk regions. Governments should define the public objectives and legal boundaries while providers implement proportionate technical controls. The June 2026 order illustrated the inverse hazard as well: when government acts through a provider without process, the provider inherits the diplomatic fallout, the customer anger, and the legal exposure for a decision it did not make.[47,48]


3.9 The Institutional Map

An Inference Blockade cannot be administered by one agency. Commerce departments and export-control bureaus understand trade and licensing. Treasury departments understand sanctions and beneficial ownership. Homeland-security agencies understand identity and infrastructure. Intelligence agencies understand adversary behavior. Standards organizations understand authentication and cybersecurity. Competition authorities understand concentration. Foreign ministries manage alliances. Courts protect rights and review government action.

The architecture therefore requires coordination without creating an unaccountable security bureaucracy. One option is a specialized interagency office for strategic AI access, supported by a technical advisory board and an appeals mechanism. Another is to preserve existing agency authorities but establish common definitions, data standards, and escalation procedures. Either design must include legislative oversight and transparent reporting.


Table 4. Three emerging governance logics

JurisdictionCore logicPrincipal instrumentsStrengthMain risk
United StatesStrategic access and alliance networksEAR, licensing, entity controls, ultimate-parent and remote-user conditions, model-level orders, full-stack export programLeverage over chips, clouds, and modelsUnilateralism, complexity, and service fragmentation
European UnionMarket accountability and lifecycle governanceAI Act, GPAI documentation, systemic-risk duties, cybersecurity, AI Office enforcement and fines from August 2026Large market and regulatory reachCompliance burden and uneven coordination with security policy
ChinaTechnological sovereignty, domestic substitution, and reciprocal controlIndustrial policy, data and cybersecurity rules, critical-input controls, proposed model, weight, data, and chip-design restrictionsScale, domestic ecosystem, and strategic state coordinationReduced openness and international technology separation

Figure 3. The United States, European Union, and China are developing different but overlapping governance logics.

UNITED STATES: Strategic Access — who may use the American stack, and on what conditions

EUROPEAN UNION: Market Accountability — what any model must document and withstand to be sold

CHINA: Technological Sovereignty — what domestic capability may leave, and in what form

Overlap zone: model weights, cloud custody, identity, and systemic-risk evaluation.


Section 4: Evasion, Attribution, and the Remote Beneficiary Problem

Every blockade creates an economy of avoidance. The more valuable the restricted capability, the greater the incentive to acquire it through intermediaries, technical workarounds, or domestic substitutes. Enforcement must therefore distinguish between ordinary global commerce and deliberate circumvention. It must also recognize that perfect denial is rarely achievable. The realistic objectives are to raise cost, delay scale, preserve visibility, and reduce access for the highest-risk actors.


4.1 The Remote Beneficiary Problem

The Remote Beneficiary Problem arises when the entity legally receiving a service is not the actor that ultimately benefits from the capability. The immediate customer may be a cloud reseller, foreign subsidiary, research partner, contractor, or shell company. Engineers in another jurisdiction may receive credentials. Outputs may be forwarded automatically. A model may be integrated into a product sold to an undisclosed end user.

Traditional end-user certification is strained by this structure. The provider may know the corporate customer but not every person or agent using the account. A multinational company may have legitimate teams across several countries. A university may collaborate internationally. A software platform may serve thousands of downstream clients. Excessive requirements can make ordinary services impossible.

The solution is risk-tiered identity. Low-capability, low-volume services may require ordinary account verification. Large training clusters, frontier-model access, high-risk tools, or unusual usage patterns may require beneficial-ownership checks, institutional verification, named administrators, and disclosure of remote regions. High-risk access should be periodically revalidated because ownership and control can change after onboarding.

The ultimate-parent rules in current U.S. regulation are an early response.[9,11,13] Yet corporate ownership is only one dimension. A legally independent company can still act on behalf of a state organization. A contractor can serve a military customer without common ownership. A researcher can share outputs informally. Attribution therefore requires multiple signals: ownership, control, funding, personnel, network behavior, billing, workload patterns, and downstream integration.


4.2 Chip Smuggling and Cluster Reconstruction

Physical diversion remains the most visible evasion pathway. Chips may be routed through third countries, misdeclared, split into smaller shipments, installed in remote facilities, or resold after legitimate purchase. Serial-number tracking, distributor due diligence, post-shipment verification, and location attestation can reduce leakage, but each introduces cost and privacy concerns.

Large clusters are harder to conceal than individual chips. They require power, cooling, networking, maintenance, and facilities. This creates opportunities for infrastructure intelligence. Utilities, data-center operators, equipment vendors, and cloud providers may detect unusual concentrations. Hardware-rooted location verification and firmware-based licensing have been proposed as technical enforcement mechanisms.[30]

Such mechanisms require careful governance. A remotely disableable chip could strengthen export enforcement, but it could also create cybersecurity vulnerabilities, supply-chain distrust, and fears of extraterritorial control. Allied customers may resist hardware that can be disabled by a foreign government or vendor. Any location-verification system needs transparent governance, narrowly defined triggers, independent security evaluation, and protection against misuse.


4.3 Cloud Resellers and Account Laundering

Cloud access can be laundered through resellers. A provider sells capacity to an authorized enterprise, which embeds it in a managed service or grants credentials to downstream users. The original provider may see only the reseller’s account. The downstream beneficiary may appear as traffic rather than a named customer.

This is why KYC obligations cannot stop at the first contract. High-risk resellers need Know Your Customer and, in selected cases, Know Your Customer’s Customer procedures. The requirement should not extend infinitely through every software layer. It should attach where the reseller provides meaningful access to controlled compute or models.

Account laundering is easier at the API Gate. Fraudulent identities, prepaid cards, compromised accounts, virtual private networks, and automation can create large pools of access. Anthropic reported that alleged distillation campaigns used roughly 24,000 fraudulent accounts and more than 16 million exchanges.[23] By April 2026, the issue had reached the highest levels of policy: the White House Office of Science and Technology Policy publicly accused foreign entities of industrial-scale campaigns to distill American frontier systems, and OpenAI separately warned lawmakers about extraction attempts against its models.[55] The reports illustrate both the scale of evasion and the provider’s ability to detect patterns across accounts.

Identity alone will not solve this problem. Behavioral detection is essential. Coordinated query patterns, high-volume extraction, repeated account creation, shared infrastructure, and unusual model-to-model prompting can reveal capability harvesting. Providers should share threat indicators under legal safeguards, much as financial institutions and cybersecurity companies share information about fraud and malicious infrastructure.


4.4 Distillation as Capability Transfer

Distillation is legitimate and widely used. A stronger model produces outputs that train a smaller or cheaper model. The technique can reduce cost, improve specialized performance, and make AI more accessible. It can also transfer capability without transferring the original weights.

This creates a major gap in weight-based controls. A user may comply with a prohibition on downloading weights while systematically querying an API to approximate the model’s behavior. The resulting model is not a copy in the simple file-transfer sense, but it may reproduce valuable capabilities. The transfer occurs through outputs.

A legal regime must avoid defining ordinary learning from outputs as prohibited extraction. Users routinely build products, fine-tune models, and analyze API responses. The relevant distinction is industrial-scale, deceptive, or contractually prohibited extraction intended to reproduce restricted capability. Evidence may include fraudulent accounts, coordinated volume, deliberate evasion of regional restrictions, and training-oriented query patterns.

The broader lesson is that capability moves through information, not only artifacts. A Weight Gate without an API Gate can be bypassed through distillation. An API Gate without open-model policy can be bypassed through public weights. A Silicon Gate without Compute Gate controls can be bypassed through remote clusters. The six gates are substitutes and complements.


4.5 Distributed Training and the Geography Problem

Training no longer needs to occur in one visible data center. Distributed systems can coordinate across clusters, regions, or providers. Network latency and communication costs limit some forms of distribution, but technical progress continues. A developer may divide experiments across sites, perform pretraining in one jurisdiction and fine-tuning in another, or combine checkpoints produced by different organizations.

This complicates thresholds based on a single training run. If each provider sees only part of the workload, no one may observe the total compute. Regulators could require aggregation across commonly controlled entities or declared projects, but undeclared coordination remains difficult to detect. Hardware attestation and model provenance may help, but they are not complete solutions.

The policy response should prioritize the highest-risk scale rather than attempt to count every operation globally. Large frontier projects leave organizational traces: financing, talent, data procurement, cluster reservations, and model releases. International intelligence and industry cooperation will remain necessary. Purely automated compute accounting cannot replace investigation.


4.6 Open Weights and Irreversibility

Open-weight systems create the hardest enforcement problem because there is no central provider to revoke access. Once a model is mirrored across repositories, private servers, and personal devices, a sanctions order can remove official distribution channels but cannot recall every copy. Open systems can also be quantized to run on smaller hardware or combined with specialized tools.

At the same time, open weights are strategically important for universities, startups, public-interest research, and countries seeking alternatives to U.S. or Chinese proprietary platforms. Research on openness emphasizes reproducibility, auditability, competition, and collaboration.[36,37] Stanford’s transparency work also shows why independent scrutiny matters: average foundation-model transparency declined substantially in 2025, and major developers disclosed little about several dimensions of development and deployment.[22]

A governable regime should therefore avoid a blanket open-weight prohibition. It should distinguish ordinary capable models from systems that cross clearly justified risk thresholds. It should encourage documentation, provenance, safety evaluations, and staged release. It should also recognize that closed APIs concentrate power and create dependence on a few companies. Security and pluralism must be balanced.

The unintended-consequence literature is important here. Wang Jin and coauthors argue that U.S. restrictions may have increased the strategic value of open, locally adaptable AI ecosystems in China.[29] Controls that make proprietary U.S. systems inaccessible can accelerate investment in open alternatives. The June 2026 frontier-model order supplied a natural experiment in miniature: within weeks of two American frontier models going dark, demand for downloadable Chinese alternatives surged, and DeepSeek reportedly closed a record funding round of roughly $7.4 billion — with the irrevocability of open weights functioning as the marketing pitch.[55] This does not prove that controls are ineffective. It shows that the response of the target — and of the uncommitted global market watching from the sidelines — is part of the policy outcome.


4.7 Tool Substitution and Agent Swarms

A provider may block a dangerous request, but a user can decompose the task across models and tools. One model creates a plan, another writes code, an open-source scanner tests it, and multiple agents search for alternatives. Each individual action may appear benign. The system-level objective emerges from coordination.

This is why the Tool and Action Gates cannot rely solely on content filters. Authorization should constrain what the system can reach. High-risk tools may require stronger identity, scoped credentials, rate limits, and human approval. Logs should preserve the chain of delegated action without capturing unrelated private content.

Agent swarms also challenge nationality-based access rules. An agent may call services in several countries, use accounts owned by different entities, and operate continuously. Machine-to-machine authentication will become as important as human identity. NIST’s work on agent identity provides a foundation, but international standards are still immature.[14,15]


4.8 The False-Positive Problem

Aggressive enforcement can harm legitimate users. Names may resemble sanctioned entities. Corporate ownership data may be outdated. Researchers may generate suspicious technical queries for defensive work. Diaspora communities and companies in transit jurisdictions may face disproportionate scrutiny. Automated risk systems can encode political bias.

Due process is therefore not an optional ethical addition. It is an operational requirement. Providers need mechanisms to explain restrictions, accept additional evidence, correct identity errors, and restore access. Government designations should be reviewable. Emergency blocks may be necessary, but permanent denial should require a stronger evidentiary basis.

A blockade that routinely excludes legitimate users will encourage them to migrate to alternative ecosystems. Overblocking is not only unfair; it can weaken strategic influence. The June 2026 global suspension — in which every customer of two models, in every allied country, lost access because a nationality screen did not exist — is the canonical example of a false-positive event at planetary scale.[45,47] The most effective regime is credible because it is targeted.


4.9 Enforcement as Delay, Not Omnipotence

No export-control system achieves perfect denial. The correct measure is marginal effect: How much cost is added? How much delay is created? How much scale is prevented? How much visibility is gained? Which high-risk actors lose reliable access?

Liu and Lee describe a strategic stalemate in which controls constrain access but also stimulate self-reliance and legal conflict.[28] The lesson is not to abandon controls. It is to avoid claims of technological containment that cannot be sustained. Inference Blockade should be understood as risk management under competition, not a promise to freeze another country’s technological development.


Table 5. Evasion pathways and proportionate defensive controls

Evasion pathwayWhy it worksDefensive controlGuardrail
Third-country chip diversionPhysical destination obscures final userDistributor due diligence, serial tracking, end-use checksProtect lawful resale and allied commerce
Cloud reseller or shell tenantProvider sees intermediary rather than beneficiaryRisk-tiered KYC and beneficial-ownership verificationDo not impose infinite downstream liability
Fraudulent API accountsAccess is divided across identities and payment methodsBehavioral detection, identity escalation, shared threat indicatorsAppeal and correction for false positives
DistillationOutputs transfer capability without weight transferExtraction detection, rate controls, contractual and legal remediesPreserve ordinary product development and research
Distributed trainingNo provider sees the full compute totalAggregation across common control, attestation, investigationAvoid universal workload inspection
Open-weight mirrorsCopies remain after official removalThresholded release policy, provenance, repository cooperationProtect open science and lower-risk models
Agent decompositionBenign sub-tasks combine into harmful system capabilityTool permissions, action limits, agent identity and auditControl high-consequence action, not ordinary assistance

Figure 4. The Remote Beneficiary Problem: the legally authorized customer may not be the strategic beneficiary.

PROVIDER  →  sees →  Authorized reseller (allied jurisdiction)

↓ credentials flow onward, unseen ↓

Shell subsidiary → Contract engineers → Undisclosed ultimate parent (restricted jurisdiction)

Legal destination answers less and less; ownership, control, and behavior answer more.


Section 5: Corporate, Economic, and Geopolitical Consequences

The Inference Blockade is not only a national-security policy. It is a reorganization of the AI economy. It changes which firms carry compliance obligations, which business models remain scalable, which countries receive frontier services, and which ecosystems become trusted or distrusted. Because the AI stack is vertically interconnected, restrictions at one gate redistribute power across the others.


5.1 Nvidia and the Hardware Providers

Nvidia, AMD, Intel, and specialized accelerator firms sit at the Silicon Gate. Their immediate obligations concern classification, licensing, distributors, destinations, and end use. But as controls move upward, chipmakers may also be expected to support serial-number verification, location attestation, security testing, firmware controls, or reporting about large clusters.

The financial record of 2025–2026 shows what occupying the Silicon Gate costs and pays. In April 2025, new licensing requirements on the H20 product line forced Nvidia to record a $4.5 billion charge for excess inventory and purchase obligations in a single quarter.[38] One year later, the company’s first quarter of fiscal 2027 contained no data-center Hopper shipments to China at all, against $4.6 billion in the comparable prior-year period, and management excluded China data-center compute from forward guidance entirely — even as total revenue reached a record $81.6 billion and free cash flow reached $48.6 billion.[38,39] The message of those numbers is double-edged. Export policy has demonstrably redirected billions of dollars of demand; and the affected firm has grown so fast elsewhere that the redirection, however large in absolute terms, has become a rounding error against an AI-factory buildout its chief executive describes as the largest infrastructure expansion in human history.[39] Political leverage over the Silicon Gate is real, but the leverage flows in both directions: Washington’s willingness in early 2026 to consider case-by-case H200 licenses for approved Chinese buyers, and the congressional revolt it provoked, showed that the gate’s operators — public and private — no longer agree on which way it should swing.[49,54]

This creates tension between market access and strategic policy. Chipmakers want broad global sales, predictable rules, and product roadmaps that are not redesigned around changing regulatory thresholds. Governments want leverage over the most capable hardware. Repeatedly modifying performance thresholds can encourage product segmentation, but it can also create a complex market of compliant, downgraded, and region-specific accelerators.

The long-term question is whether hardware companies remain sellers of components or become participants in an ongoing compliance relationship. Location verification and secure firmware would make the manufacturer relevant after sale. That may improve enforcement, but it could expose the company to political retaliation and customer distrust. The semiconductor industry could be transformed from a manufacturing supply chain into a governed service layer.


5.2 Hyperscalers as Intelligence Intermediaries

Amazon Web Services, Microsoft Azure, Google Cloud, Oracle, and specialized AI clouds occupy the Compute Gate. They are positioned to identify customers, allocate accelerators, monitor aggregate usage, restrict regions, and preserve logs. Academic research describes them as potential securers, record keepers, verifiers, and enforcers.[19] Current regulation is beginning to formalize parts of that role.[11,12,13]

The scale of what these intermediaries now steward magnifies both their enforcement value and their exposure. With combined 2026 capital budgets around $725 billion — Amazon near $200 billion, Microsoft and Alphabet near $190 billion each, Meta well above $125 billion — the four largest American hyperscalers are constructing, in a single year, more strategic compute than most nations will ever possess, backed by contracted demand such as Google Cloud’s reported backlog exceeding $460 billion.[42,43,44] Every gate obligation imposed on these firms therefore governs an enormous share of the world’s usable intelligence in one administrative stroke; and every obligation they resist, they resist with balance sheets larger than most national budgets. The Compute Gate’s promise of precise, reversible control is inseparable from this concentration.

Cloud providers will resist obligations that require intrusive workload inspection or make them liable for undisclosed customer behavior. They will support rules that create common standards, reduce uncertainty, and prevent competitors from offering noncompliant access. The design of safe harbors will therefore shape market structure. A provider that follows approved identity, logging, and escalation procedures should not face unlimited liability for sophisticated deception it could not reasonably detect.

The compliance burden may favor the largest providers. They can build specialized legal, security, and intelligence teams. Smaller clouds may struggle with beneficial-ownership checks, threat detection, and global sanctions screening. Regulation intended to prevent adversarial access could unintentionally consolidate the cloud market. Public standards, shared utilities, and proportionate requirements are needed to prevent security from becoming an entry barrier.

Cloud providers also face jurisdictional conflict. One government may require disclosure of remote users while another restricts data transfer. A multinational customer may be lawful in one region and restricted in another. Providers will increasingly partition services by geography, ownership, model class, and tool access. The cloud may become less universal and more treaty-like.


5.3 Model Laboratories and the Politics of Access

OpenAI, Anthropic, Google DeepMind, Meta, xAI, Mistral, Chinese model developers, and other laboratories occupy the Weight and API Gates. Their release choices determine whether capability is centralized, downloadable, auditable, revocable, or widely reproducible. Their security programs protect weights; their trust-and-safety teams enforce access; their enterprise systems manage identity and permissions.

Model providers have already become geopolitical actors. When they terminate a state-linked account, block a region, or restrict a tool, they exercise a form of platform sovereignty. Their public threat reports show that model access is monitored and can be withdrawn.[24,25] Anthropic’s activation of stronger model-weight protections, its controlled-access Glasswing structure for its most capable cyber-relevant system, and its work on confidential inference demonstrate the technical investment required to secure frontier capability.[21,23,46,48]

The events of June 2026 revealed the other side of that position: the laboratory as object, not subject, of geopolitics. A single directive converted a company’s flagship commercial products into controlled items overnight, placed its foreign employees outside the permission boundary of its own models, and made its access decisions a matter of intergovernmental concern from Brussels to Beijing.[45,46,47] Laboratories now carry a strategic profile once reserved for defense contractors, without the procurement relationships, security clearances, or statutory frameworks that historically accompanied that profile. Closing that institutional gap — through formal consultation channels, classified threat sharing, defined emergency procedures, and clear legal authority — is among the most urgent tasks identified in this paper.

The problem is legitimacy. A model company may be asked to decide whether a foreign research institution is military-linked, whether a request is dual use, or whether an account is part of a distillation campaign. These are not ordinary customer-service decisions. Companies need government intelligence and clear legal standards, while governments need provider expertise and evidence. Formal public-private procedures are preferable to informal pressure.


5.4 Open Models, Startups, and the Innovation Tradeoff

Startups benefit from open models, affordable APIs, and rented compute. They are also vulnerable to compliance cost. A large laboratory can maintain export-control counsel and global identity systems; a five-person startup cannot. If every model integration requires geopolitical due diligence, the result will be fewer entrants and greater dependence on incumbent platforms.

The policy design should separate providers of strategic access from ordinary downstream developers. A company reselling large-scale frontier compute or offering a high-risk agent platform may need enhanced obligations. A small business using an API for customer service should not become an export-control intermediary. Responsibilities should attach to control and visibility.

Open models complicate this allocation. A startup can deploy weights locally and become the provider of record. This increases autonomy and competition but shifts security responsibility downstream. Documentation, provenance, standardized model cards, and safety tooling can help. Prohibiting open models would solve neither malicious use nor international competition; it would primarily transfer power to closed providers.


5.5 The Global South and Intelligence Dependency

Countries outside the leading U.S.–China–European technology blocs face a difficult choice. They need AI for education, health, agriculture, government, language technology, and industry. Building frontier models domestically is expensive. Imported full-stack packages can accelerate adoption, but they create dependency on foreign chips, cloud services, model updates, and policy decisions.

An Inference Blockade could divide the world into trusted, conditional, and excluded intelligence zones. Countries may be evaluated through political alignment, cybersecurity, ownership rules, and their willingness to prevent diversion. This can resemble financial-risk grading, but with access to machine capability rather than capital.

The stakes of that grading are rising quickly because adoption in the Global South is not waiting for policy. Stanford’s 2026 AI Index found generative-AI adoption rates in some emerging and middle-income economies exceeding those of the United States, and reported that more than half of all national AI strategies adopted since 2024 originated in emerging economies, many explicitly framed around sovereignty and local capacity.[33,53] A blockade architecture designed as though frontier demand were a rich-country phenomenon will misread its own strategic environment.

If the process is opaque or humiliating, countries will seek alternatives. Chinese open models and infrastructure packages may be attractive because they offer local deployment or fewer access conditions — although the reported July 2026 Chinese proposals suggest that Beijing’s most capable systems may themselves retreat behind API-only access, narrowing that alternative at the top end.[31,54] European providers may compete on privacy. Regional powers may build sovereign clouds. The global contest will be shaped not only by benchmark performance but by the perceived reliability of access — and June 2026 demonstrated to every procurement ministry on Earth that reliability of access to American frontier models is now a variable to be priced, not a constant to be assumed.[47]

The United States and its allies should therefore treat trusted access as a development strategy, not merely a policing strategy. Secure regional compute, academic access, local-language models, transparent licensing, and technical assistance can reduce incentives for diversion. A blockade without an inclusion policy will push undecided countries toward competing ecosystems.


5.6 Allied Coordination and the Weakest-Link Problem

Advanced semiconductors depend on allied supply chains. Cloud and model controls will also require coordination. If one jurisdiction restricts an actor but another offers equivalent compute or API access, the control loses effect. Yet full harmonization is unrealistic because allies differ in commercial interests, privacy law, relations with China, and views on open source.

The objective should be interoperability rather than identical law. Allies can agree on high-risk definitions, ownership data, license-recognition, incident reporting, and evidence standards while retaining different domestic procedures. Multilateral hardware coordination identified in 2020 remains relevant, but the agenda must now include model and service access.[32]

A trusted network also requires assurance against unilateral disruption. Partners will hesitate to build critical systems on foreign APIs if access can be revoked without consultation. Long-term agreements, transition periods, appeal channels, and continuity provisions can make strategic access more credible. Trust is an enforcement asset.


5.7 China, Adaptation, and Reciprocal Controls

The U.S.–China competition will not produce a one-way blockade. China is investing in domestic accelerators, model efficiency, open ecosystems, cloud platforms, and potential controls over its own technology. Research published in 2026 suggests that U.S. restrictions may have accelerated Chinese emphasis on open and locally adaptable AI.[29] Other work warns that controls can create strategic stalemate and stronger self-reliance.[28]

The material asymmetry nonetheless remains large, and measuring it honestly matters for policy calibration. Miller’s mid-2026 assessment holds that Chinese hyperscalers have collectively underinvested in AI infrastructure relative to their American counterparts since 2022, and that allied production of quality-adjusted accelerators may exceed Chinese production by an order of magnitude or more.[51] Stanford’s investment data point the same direction: U.S. private AI investment in 2025 exceeded China’s reported figure by a factor of more than twenty, even acknowledging that state guidance funds make the Chinese total larger than private data capture.[33,53] Yet the capability gap at the model layer tells a different story — the best Chinese systems trailed the best American one by under three percentage points on composite benchmarks as of March 2026 — which is precisely the pattern the efficiency-adaptation literature predicts: constrained inputs, converging outputs.[33] The two measurements together define the strategic problem. Hardware controls are visibly working at the input layer and visibly insufficient at the capability layer, which is why the frontier of control keeps moving up the stack.

This does not make the controls pointless. Delay has strategic value. Denying reliable access to the largest clusters can affect experimentation, scale, and deployment. But policy should be assessed against realistic counterfactuals. The relevant question is not whether China continues to innovate. It is how the speed, cost, reliability, and direction of innovation change.

Reciprocal model and data controls would also create bargaining opportunities. The United States and China may eventually negotiate around categories of access, safety testing, military end use, or incident communication. A fully fragmented world would increase duplication and reduce scientific exchange. Selective agreements may be possible even amid competition, especially for preventing loss-of-control incidents or unauthorized military use. The mirrored structure now visible — each side treating frontier models as national assets, each contemplating weight custody rules, each preserving API access as the diplomatic release valve — creates, for the first time, a shared vocabulary in which such negotiation could occur.[31,54]


5.8 Economic Fragmentation and the Price of Intelligence

A fragmented AI market will raise costs. Providers will maintain separate infrastructure, compliance systems, model versions, and regional policies. Customers will pay for sovereignty, locality, and auditability. Smaller markets may receive older or less capable models. Cross-border research will slow.

At the same time, fragmentation can create new industries: identity verification for agents, model provenance, secure inference, trusted execution environments, compute auditing, sanctions analytics, and sovereign cloud infrastructure. The compliance layer around AI may become as economically significant as cybersecurity became around the internet.

The distributional effects will matter. Large countries can negotiate access and build domestic capacity. Small countries and startups may become price takers. Universities may lose access to frontier systems if licensing is designed around commercial users. Humanitarian, medical, climate, and academic safe harbors should therefore be built into the architecture from the beginning.


5.9 The Risk of an Intelligence Iron Curtain

The most serious danger is that a targeted security regime expands into a broad separation of scientific and commercial communities. Once every model interaction is framed as potential strategic transfer, openness becomes suspicious. Researchers avoid collaboration. Providers over-block. Governments use national security to protect domestic firms. Citizens lose access based on nationality rather than conduct.

An Inference Blockade should not become an intelligence iron curtain. The framework is justified only where capability, user, and context create a credible risk. Controls should be narrow enough to preserve ordinary education, communication, research, and commerce. Their effectiveness should be reviewed against measurable objectives rather than symbolic toughness.


Table 6. Corporate exposure across the post-chip sanctions stack

ActorPrimary gateNew responsibilityStrategic opportunityKey risk
ChipmakersSiliconClassification, customer diligence, possible location or firmware supportTrusted hardware and compliant-market accessThreshold churn and customer distrust
Hyperscalers / neocloudsComputeKYC, beneficial ownership, workload records, remote-user controlsPreferred trusted infrastructure intermediarySurveillance pressure, liability, concentration
Model laboratoriesWeights / APIsWeight security, account enforcement, capability evaluation, emergency compliancePremium secure access and allied deploymentPrivate foreign-policy power, inconsistent due process, abrupt government orders
Agent and tool platformsTools / ActionsScoped credentials, agent identity, approvals, logsEnterprise trust and regulated automationCross-platform delegation and prompt injection
Startups and universitiesDownstream useDocumentation and compliance at proportionate scaleOpen models, safe harbors, research accessIncumbent advantage and overblocking
Governments and alliesAll gatesCommon definitions, licenses, intelligence sharing, appealsTrusted AI network and bargaining leverageFragmentation and retaliation

Section 6: What Have We Learned? Five Pillars for a Governable Inference Blockade

The preceding sections reveal a policy problem larger than export licensing. Artificial intelligence is becoming a cross-border service that can be owned, rented, copied, queried, integrated, and delegated. A governable regime must follow capability across these forms without turning the global digital economy into a permissioned security zone.

The five pillars below are not a blueprint for closing every gate. They are tests for deciding when a gate should be used, by whom, with what evidence, and under what limits. Each pillar is drawn from the failures and near-misses documented above: the granularity collapse of June 2026, the distillation campaigns of 2025–2026, the corporate-genealogy problem of the ultimate-parent rules, the allied-trust deficit exposed by abrupt orders, and the irreversibility of open weights.


Pillar 1 — Identity Before Access

The first pillar is Identity Before Access. Strategic AI services should not be governed through nationality labels alone, but high-risk access cannot remain effectively anonymous. The provider needs confidence about the organization, beneficial owner, ultimate parent, administrators, remote regions, and delegated agents associated with the account.

Identity should be graduated. Ordinary consumer use can rely on ordinary fraud controls. Frontier training, unusually high-volume extraction, controlled model weights, dangerous tools, and high-consequence autonomous actions require stronger verification. The greater the capability and irreversibility, the higher the assurance level. The June 2026 order failed on exactly this dimension: a nationality-based restriction was imposed on an infrastructure that had never been asked to verify nationality, and the only lawful response was global shutdown.[45,47] Identity infrastructure is not surveillance for its own sake; it is the precondition for precision, and precision is the precondition for keeping legitimate access open.

Identity must extend to software agents. NIST’s 2026 work correctly emphasizes identification, authorization, auditing, and non-repudiation for agentic systems.[14,15] An agent should have a verifiable principal, defined authority, scoped credentials, and a record of actions. Machine identity will become part of sanctions compliance because an agent can transact or operate across borders without a human login at each step.

Identity systems should minimize data and provide correction procedures. A false match should not become permanent exclusion. Beneficial-ownership records should be current and protected. High-risk determinations should be reviewable. Identity is a gate, not a presumption of guilt.


Pillar 2 — Capability-Proportional Controls

The second pillar is Capability-Proportional Controls. Regulation should respond to the actual strategic value and risk of the system, not merely the presence of the word AI. A small translation model, a frontier cyber agent, and a multimodal scientific platform should not be treated alike.

The six gates provide different levels of intervention. A government may deny a chip export, cap cloud capacity, restrict weight downloads, require enhanced API monitoring, disable a dangerous tool, or require human approval for an autonomous action. The narrowest effective gate should be preferred. The Glasswing structure that survived the June 2026 episode — exceptional capability held behind verified institutional access while a safeguarded derivative serves the general market — is an early working example of capability tiering done at the provider layer.[46,48]

Capability measurement should combine multiple signals: training compute, evaluations, system scaffolding, tool access, deployment scale, and end-use context. Riegler and Strümke’s system-level argument is important because small models can become powerful when coordinated with tools and agents.[26] Thresholds based on a single model characteristic will be gamed or become obsolete.

Proportionality also requires safe harbors. Defensive cybersecurity, academic evaluation, humanitarian use, low-volume research, and clearly bounded commercial services may deserve exemptions or streamlined licenses. A regime without safe harbors will drive legitimate users toward unregulated alternatives.


Pillar 3 — Custody and Provenance Across the Lifecycle

The third pillar is Custody and Provenance Across the Lifecycle. The Capability Custody Chain should record who controls the strategic asset as it changes form: chip, compute, checkpoint, weight, API, agent, and action.

Provenance does not require recording every prompt or exposing training data. It requires durable identifiers and auditable transitions. Hardware can use serials and attestations. Workloads can use account and cluster records. Weights can use hashes, access logs, and secure storage. APIs can use verified accounts and rate records. Agents can use signed identity and authorization tokens. High-consequence actions can use transaction logs.

Model provenance is particularly important after fine-tuning, merging, or distillation. A derivative system may not be identical to the original, but its lineage can help determine obligations and risk. Provenance standards should be interoperable and protect trade secrets. They should support investigation without creating a public registry of sensitive research.

The principle is continuity of responsibility. When capability moves to a new custodian, obligations should not disappear. The transfer should be explicit, documented, and proportionate to the asset.


Pillar 4 — Allied Interoperability and Trusted Access

The fourth pillar is Allied Interoperability and Trusted Access. No national blockade can govern a global AI stack alone. Hardware supply chains, cloud regions, repositories, models, and agent platforms cross jurisdictions. The weakest uncoordinated gateway can become the route of diversion.

Allies do not need identical laws. They need interoperable definitions, due-diligence standards, licensing evidence, security requirements, and information-sharing procedures. A customer verified in one trusted jurisdiction should not repeat every process unnecessarily. A high-risk determination should be explainable and reviewable across the network.

Trusted access must be a positive offer. The United States’ full-stack export strategy recognizes that technological leadership depends on diffusion among partners.[7,8] Secure infrastructure financing, local-language models, academic programs, and continuity guarantees can make compliance attractive rather than coercive.

This pillar also limits unilateralism. Allies need consultation before major access rules change. Companies need transition periods. Countries need confidence that ordinary political disagreements will not suddenly terminate essential services. The June 2026 order — issued without allied consultation and examined by the European Commission within days — showed how quickly unilateral action at the API Gate converts partners into skeptics.[45,47] Predictable access strengthens the legitimacy and durability of the network.


Pillar 5 — Reversibility, Auditability, and Democratic Oversight

The fifth pillar is Reversibility, Auditability, and Democratic Oversight. The great advantage of service-layer controls is reversibility. An account can be limited or restored. A tool permission can be scoped. A license can be modified. This flexibility should be used to avoid permanent, nationality-wide bans.

Reversibility must work in both directions. Governments and providers need the ability to stop a dangerous use quickly. Legitimate users need a path to regain access when facts change or errors are corrected. Emergency measures should expire unless renewed with evidence. The June 2026 episode was, in this one respect, a success: the restriction was imposed in hours and substantially lifted within weeks once evidence and negotiation caught up with the emergency.[46] A well-designed regime would make that reversibility a formal property of the instrument rather than an accident of political pressure.

Auditability means that important decisions leave records: who designated the user, what authority applied, which evidence supported the restriction, how the provider implemented it, and whether the control achieved its objective. Aggregate public reporting can improve accountability without exposing sensitive intelligence.

Democratic oversight is essential because the Inference Blockade can influence speech, research, competition, and foreign policy. Legislatures should define authorities and receive regular effectiveness reports. Courts or independent review bodies should hear challenges. Technical advisory groups should include security experts, economists, civil-liberties specialists, academics, and industry representatives.

The purpose of oversight is not to make enforcement impossible. It is to prevent an emergency architecture from becoming permanent and unbounded.


Table 7. Inference Blockade Decision Form

Decision testQuestions for policymakersEvidence requiredReview safeguard
1. Strategic capabilityWhat material capability or harm is at issue? Is the concern the model, system, tool, or action?Evaluations, threat intelligence, deployment contextPublish a non-sensitive statement of rationale
2. Real beneficiaryWho owns, controls, funds, administers, and ultimately benefits from access?Ownership records, institutional identity, remote regions, behaviorCorrection and appeal process
3. Narrowest gateCan risk be addressed at a narrower gate than complete denial?Substitution analysis across the Six GatesPrefer reversible and scoped measures
4. ProportionalityAre low-risk users, research, humanitarian work, and ordinary commerce protected?Impact assessment and safe-harbor designIndependent civil-liberties and competition review
5. Allied feasibilityCan partners implement compatible controls, or will access shift elsewhere?Allied consultation and provider capabilityMutual recognition and transition periods
6. Technical enforceabilityCan the measure be verified without universal content surveillance?Metadata, attestation, provenance and audit designData minimization and security testing
7. EffectivenessWhat delay, cost, scale reduction, or visibility gain is expected?Baseline, metrics, counterfactual and evasion monitoringRegular public effectiveness report
8. Sunset and remedyWhen does the measure expire, and how can access be restored?Defined trigger, expiration date and restoration criteriaLegislative oversight and judicial or independent review

6.1 A Practical Policy Sequence

The five pillars suggest a practical sequence for decision makers. First, identify the capability and harm, not merely the technology label. Second, identify the real beneficiary and context. Third, choose the narrowest effective gate. Fourth, assess substitution and evasion. Fifth, provide safe harbors, appeal, and review. Sixth, coordinate with allies and providers. Seventh, measure the result.

This sequence should be documented through the Inference Blockade Decision Form included in this paper. The form asks whether the capability is materially strategic, whether the identity evidence is sufficient, whether a less restrictive measure is available, whether allies can implement the control, whether privacy and competition effects have been assessed, and when the measure will be reviewed.

The discipline of writing down these answers matters. National-security technology controls are often introduced during moments of urgency. Had a form of this kind governed the June 2026 order, the questions it would have forced — Is the narrowest gate global suspension? What identity evidence supports nationality screening? What is the sunset trigger? — are precisely the questions that ended up being litigated in public, after the fact, at maximum cost to every party.[47,48] A structured test forces policymakers to explain the causal chain from access to harm and from intervention to expected effect.


Conclusion: The Border Has Moved

Return to the two scenes from the introduction. In the first, the customs officer stands beside a container. The shipment has an exporter, consignee, destination, and manifest. The law can inspect the item because the item moves through a physical chokepoint.

In the second, the accelerators remain inside an approved data center. The capability moves through credentials, ownership relationships, model files, API calls, tools, and autonomous actions. There may be no crate to open. There may be no ship to stop. Yet intelligence has crossed the border.

This is why the main title of this paper is Inference Blockade. The phrase names the moment when strategic technology policy moves beyond possession of the machine and begins to govern access to the intelligence produced by the machine. The post-chip regime does not replace the semiconductor embargo. It extends its logic upward.

The extension is no longer hypothetical. U.S. rules address model weights, ultimate-parent relationships, remote IaaS users, KYC, storage security, and access to algorithms trained on controlled hardware.[9,10,11,12,13] The American AI strategy seeks both stronger enforcement and full-stack exports to partners.[7,8] In June 2026, the United States briefly converted two deployed frontier models into licensable exports and then substantially reversed itself, teaching every government and boardroom that the API Gate can be closed — and how much it costs to close it badly.[45,46,47] Europe’s AI Office acquires the power to evaluate, restrict, and fine general-purpose models in the European market from August 2, 2026.[16,52] China is consulting its champions on reciprocal restrictions covering models, data, weights, designs, and acquisitions.[31,54] NIST is building standards for the identity and authority of software agents.[14,15] Private providers are already terminating accounts, detecting industrial-scale extraction, and governing access at the platform layer.[23,24,25] Behind all of it, three-quarters of a trillion dollars of annual infrastructure spending is pouring concrete around the gates.[38,40,42]

These developments are fragments of one future architecture. The Six Gates — Silicon, Compute, Weights, APIs, Tools, and Actions — show how the sanctions surface expands as AI becomes a service and then an actor. The Capability Custody Chain shows why responsibility must continue across transitions. The Remote Beneficiary Problem shows why legal destination is no longer enough. The five pillars show how the system can be governed without treating every user as an adversary.

The strongest case for an Inference Blockade is narrow. It is the case for preventing clearly identified high-risk actors from obtaining exceptional capabilities that materially strengthen military, cyber, surveillance, weapons, or coercive power. The weakest case is broad. It is the temptation to convert all advanced knowledge into a controlled commodity and all cross-border model use into suspicion.

A successful regime must live between these extremes. It must be technically literate enough to understand substitution, restrained enough to preserve science and commerce, and adaptable enough to follow capability from hardware to action. It must recognize that open models have public value, that cloud providers should not become universal surveillance agencies, that allies require predictable access, and that controls can stimulate the very alternatives they seek to constrain — as the record fundraising that followed the June 2026 suspension made unmistakably clear.[55]

The decisive policy question is no longer simply, “Did the chip leave the country?” It is, “Who received the capability, through which gate, under whose authority, with which tools, and toward what action?”

The border has moved. It now runs through the data-center account, the model repository, the API credential, the agent identity, and the permission to act. The institutions that learn to govern that border with precision will shape the next era of artificial intelligence. Those that govern it indiscriminately may discover that a blockade can also become a wall around their own influence.


Notes and Sources:

[1] Saif M. Khan and Alexander Mann, “AI Chips: What They Are and Why They Matter,” Center for Security and Emerging Technology, Georgetown University, April 2020. https://cset.georgetown.edu/publication/ai-chips-what-they-are-and-why-they-matter/

[2] U.S. Department of Commerce, Bureau of Industry and Security, “Commerce Implements New Export Controls on Advanced Computing and Semiconductor Manufacturing Items to the People’s Republic of China,” October 7, 2022. https://www.bis.gov/press-release/commerce-implements-new-export-controls-advanced-computing-semiconductor-manufacturing-items-peoples

[3] U.S. Department of Commerce, Bureau of Industry and Security, “Commerce Strengthens Restrictions on Advanced Computing Semiconductors, Semiconductor Manufacturing Equipment, and Supercomputing Items,” October 17, 2023. https://www.bis.gov/press-release/commerce-strengthens-restrictions-advanced-computing-semiconductors-semiconductor-manufacturing-equipment

[4] U.S. Department of Commerce, Bureau of Industry and Security, “Commerce Strengthens Export Controls to Restrict China’s Capability to Produce Advanced Semiconductors for Military Applications,” December 2, 2024. https://www.bis.gov/press-release/commerce-strengthens-export-controls-restrict-chinas-capability-produce-advanced-semiconductors-military

[5] U.S. Department of Commerce, Bureau of Industry and Security, “Regulatory Framework for the Responsible Diffusion of Advanced Artificial Intelligence Technology,” January 13, 2025. https://www.bis.gov/press-release/biden-harris-administration-announces-regulatory-framework-responsible-diffusion-advanced-artificial

[6] U.S. Department of Commerce, Bureau of Industry and Security, “Rescission of the Biden-Era Artificial Intelligence Diffusion Rule and Strengthened Chip-Related Export-Control Guidance,” May 13, 2025. https://www.bis.gov/press-release/department-commerce-announces-rescission-biden-era-artificial-intelligence-diffusion-rule-strengthens

[7] The White House, “Winning the Race: America’s AI Action Plan,” Executive Office of the President, July 2025. https://www.whitehouse.gov/wp-content/uploads/2025/07/Americas-AI-Action-Plan.pdf

[8] The White House, “Executive Order 14320: Promoting the Export of the American AI Technology Stack,” Executive Office of the President, July 23, 2025. https://www.whitehouse.gov/presidential-actions/2025/07/promoting-the-export-of-the-american-ai-technology-stack/

[9] U.S. Department of Commerce, Bureau of Industry and Security, “Export Administration Regulations, Part 742: Control Policy — CCL Based Controls,” current through July 2026. https://www.bis.gov/regulations/ear/742

[10] U.S. Department of Commerce, Bureau of Industry and Security, “Export Administration Regulations, Part 740: License Exceptions,” current through July 2026. https://www.bis.gov/regulations/ear/740

[11] U.S. Department of Commerce, Bureau of Industry and Security, “Export Administration Regulations, Part 748: Applications, Classification Requests, and Advisory Opinions,” current through July 2026. https://www.bis.gov/regulations/ear/748

[12] U.S. Department of Commerce, Bureau of Industry and Security, “Know Your Customer Guidance and Red Flag 28 for Infrastructure-as-a-Service Providers,” 2025–2026. https://www.bis.gov/node/1533

[13] U.S. Department of Commerce, Bureau of Industry and Security, “Revision to License Review Policy for Advanced Computing Commodities,” Federal Register, January 15, 2026. https://www.federalregister.gov/documents/2026/01/15/2026-00789/revision-to-license-review-policy-for-advanced-computing-commodities

[14] Harold Booth, William Fisher, Ryan Galluzzo, and Joshua Roberts, “Accelerating the Adoption of Software and Artificial Intelligence Agent Identity and Authorization,” National Institute of Standards and Technology, February 5, 2026. https://csrc.nist.gov/pubs/other/2026/02/05/accelerating-the-adoption-of-software-and-ai-agent/ipd

[15] National Institute of Standards and Technology, “AI Agent Standards Initiative and Software and AI Agent Identity and Authorization Project,” NIST and NCCoE, 2026. https://www.nccoe.nist.gov/projects/software-and-ai-agent-identity-and-authorization

[16] European Commission, “AI Act: Regulatory Framework and Application Timeline,” Shaping Europe’s Digital Future, updated July 2026. https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai

[17] European Commission, “Guidelines for Providers of General-Purpose AI Models,” AI Office, European Commission, April 28, 2026. https://digital-strategy.ec.europa.eu/en/policies/guidelines-gpai-providers

[18] Girish Sastry, Lennart Heim, Haydn Belfield, Markus Anderljung, Miles Brundage, Gillian K. Hadfield, Yoshua Bengio, Diane Coyle, and coauthors, “Computing Power and the Governance of Artificial Intelligence,” arXiv, 2024. https://arxiv.org/abs/2402.08797

[19] Lennart Heim, Tim Fist, Janet Egan, Sihao Huang, Stephen Zekany, Robert Trager, Michael A. Osborne, and Noa Zilberman, “Governing Through the Cloud: The Intermediary Role of Compute Providers in AI Regulation,” arXiv, 2024. https://arxiv.org/abs/2403.08501

[20] Janet Egan and Lennart Heim, “Oversight for Frontier AI through a Know-Your-Customer Scheme for Compute Providers,” arXiv, 2023. https://arxiv.org/abs/2310.13625

[21] RAND Corporation, “Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models,” 2024. https://www.rand.org/content/dam/rand/pubs/research_reports/RRA2800/RRA2849-1/RAND_RRA2849-1.pdf

[22] Stanford Center for Research on Foundation Models, “Foundation Model Transparency Index 2025,” Stanford University, December 2025. https://crfm.stanford.edu/fmti/

[23] Anthropic, “Detecting and Preventing Distillation Attacks,” February 23, 2026. https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks

[24] OpenAI and Microsoft Threat Intelligence, “Disrupting Malicious Uses of AI by State-Affiliated Threat Actors,” February 14, 2024. https://openai.com/index/disrupting-malicious-uses-of-ai-by-state-affiliated-threat-actors/

[25] OpenAI, “Disrupting Malicious Uses of AI,” February 25, 2026. https://openai.com/index/disrupting-malicious-ai-uses/

[26] Michael A. Riegler and Inga Strümke, “Position: AI Security Policy Should Target Systems, Not Models,” arXiv, May 2026. https://arxiv.org/abs/2605.09504

[27] Ritwik Gupta, Leah Walker, and Andrew W. Reddie, “Whack-a-Chip: The Futility of Hardware-Centric Export Controls,” arXiv, November 2024. https://arxiv.org/abs/2411.14425

[28] Jingwen Liu and Jyh-An Lee, “Strategic Stalemates: The Paradox of Export Controls in the U.S.-China AI Race,” arXiv, May 2026. https://arxiv.org/abs/2605.23475

[29] Wang Jin, Nadav Kunievsky, Bowen Lou, Tianshu Sun, and James Evans, “U.S. Policies Unintentionally Accelerated China’s Open AI Ecosystems,” arXiv, June 2026. https://arxiv.org/abs/2606.15999

[30] James Petrie, “Near-Term Enforcement of AI Chip Export Controls Using a Firmware-Based Design for Offline Licensing,” arXiv, April 2024. https://arxiv.org/abs/2404.18308

[31] Reuters, “China Considers Tighter Export Controls on AI Models and Chips, FT Reports,” July 21, 2026. https://www.reuters.com/world/asia-pacific/china-considers-tighter-export-controls-ai-models-chips-ft-reports-2026-07-21/

[32] Carrick Flynn and Saif M. Khan, “Multilateral Controls on Hardware Chokepoints,” Center for Security and Emerging Technology, Georgetown University, September 2020. https://cset.georgetown.edu/publication/multilateral-controls-on-hardware-chokepoints/

[33] Stanford Institute for Human-Centered Artificial Intelligence, “The 2026 AI Index Report,” Stanford University, 2026. https://hai.stanford.edu/ai-index/2026-ai-index-report

[34] Dan Hendrycks, Mantas Mazeika, Thomas Woodside, and coauthors, “Model Evaluation for Extreme Risks,” arXiv, May 2023. https://arxiv.org/abs/2305.15324

[35] Lennart Heim and Governance of AI research team, “Accessing Controlled AI Chips via Infrastructure-as-a-Service,” Centre for the Governance of AI, 2023. https://cdn.governance.ai/Accessing_Controlled_AI_Chips_via_Infrastructure-as-a-Service.pdf

[36] Elizabeth Seger, Aviv Ovadya, Ben Garfinkel, Divya Siddarth, and coauthors, “Open-Sourcing Highly Capable Foundation Models,” Centre for the Governance of AI, 2023. https://cdn.governance.ai/Open-Sourcing_Highly_Capable_Foundation_Models_2023_GovAI.pdf

[37] Matt White and coauthors, “The Model Openness Framework: Promoting Completeness and Openness for Reproducibility, Transparency, and Usability in Artificial Intelligence,” arXiv, 2024. https://arxiv.org/abs/2403.13784

[38] CNBC, “Nvidia (NVDA) Q1 2027 Earnings Report: Live Updates,” May 20, 2026. https://www.cnbc.com/2026/05/20/nvidia-nvda-earnings-report-q1-2027.html

[39] NVIDIA Corporation via GlobeNewswire / StockTitan, “NVIDIA Announces Financial Results for First Quarter Fiscal 2027,” May 20, 2026 (record $81.6 billion revenue; $75.2 billion data-center revenue; no China data-center compute in outlook; Jensen Huang statement). https://www.stocktitan.net/news/NVDA/nvidia-announces-financial-results-for-first-quarter-fiscal-fq78amc9h84m.html

[40] Investing.com, “Earnings Call Transcript: TSMC Lifts 2026 Outlook as AI Demand Stays Hot in Q2 2026,” July 16, 2026 ($40.2 billion revenue; 67.7% gross margin; full-year growth outlook raised above 40%). https://www.investing.com/news/transcripts/earnings-call-transcript-tsmc-lifts-2026-outlook-as-ai-demand-stays-hot-in-q2-2026-93CH-4794777

[41] TechTimes, “TSMC Posts Record Quarter as AI Chip Demand Pushes Full-Year Growth Outlook Past 40%,” July 16, 2026 (record net profit; CoWoS constraint; C.C. Wei statement). https://www.techtimes.com/articles/320696/20260716/tsmc-posts-record-quarter-ai-chip-demand-pushes-full-year-growth-outlook-past-40.htm

[42] Tom’s Hardware, “Google, Microsoft, Meta, and Amazon Capex Spending to Hit $725 Billion in 2026, Up 77% From Last Year,” April 30, 2026 (Financial Times earnings compilation; Brent Thill, Jefferies, statement). https://www.tomshardware.com/tech-industry/big-tech/big-techs-ai-spending-plans-reach-725-billion

[43] Fortune, “Microsoft, Meta, and Google Just Announced Billions More in AI Spending. Only Google Convinced Investors It’s Paying Off,” April 29, 2026 (Alphabet capex guidance; Google Cloud backlog; Anat Ashkenazi statements). https://fortune.com/2026/04/29/microsoft-meta-google-ai-capex-spending-billions/

[44] CNBC, “Tech AI Spending Approaches $700 Billion in 2026, Cash Taking Big Hit,” February 6, 2026 (hyperscaler capex projections; Susan Li, Meta CFO, statement). https://www.cnbc.com/2026/02/06/google-microsoft-meta-amazon-ai-cash.html

[45] CNN Business, “Anthropic Suspends All Access to Mythos Model After US Government Bans Foreign Nationals Use,” June 13, 2026 (Anthropic corporate statement; Commerce Department directive). https://www.cnn.com/2026/06/13/business/anthropic-mythos-model-national-security

[46] Alexander Martin, “US Lifts Export Controls on Anthropic’s Frontier Cybersecurity AI Models,” The Record, Recorded Future News, June 30, 2026 (restoration of Fable 5; Mythos 5 gated through Project Glasswing). https://therecord.media/us-lifts-export-controls-anthropic-cyber-models

[47] Matthew Ferren, “The U.S. Is Losing the AI Credibility War — to Itself,” Council on Foreign Relations, June 22, 2026. https://www.cfr.org/articles/the-u-s-is-losing-the-ai-credibility-war-to-itself

[48] Herbert Smith Freehills Kramer, “License to Model: Emerging US Rules Impact Global Access to Frontier AI,” July 2026 (Commerce directive timeline; ninety-minute compliance window; Project Glasswing; legality debate). https://www.hsfkramer.com/insights/2026-07/license-to-model-emerging-us-rules-impact-global-access-to-frontier-ai

[49] Chris Miller (The Fletcher School, Tufts University), interview on CNBC Squawk Box, “’Chip War’ Author Chris Miller on the Battle of AI Chip Export Controls,” December 2025, via StartupHub.ai summary. https://www.startuphub.ai/ai-news/ai-video/2025/chip-war-author-chris-miller-on-the-battle-of-ai-chip-export-controls/

[50] Carnegie Mellon Institute for Strategy and Technology, “Chips and Chokepoints: Chris Miller on the Geopolitics of the AI Supply Chain,” Carnegie Mellon University, March 2026. https://www.cmu.edu/cmist/news-archive/news/2026/march/chips-and-chokepoints-chris-miller-on-the-geopolitics-of-the-ai-supply-chain.html

[51] Crypto Briefing, “China’s AI Spending Lags Behind the US by a Staggering Margin, Says ‘Chip War’ Author Chris Miller,” June 10, 2026 (quality-adjusted accelerator production estimates; Chinese hyperscaler capex trends). https://cryptobriefing.com/china-ai-spending-lags-chip-war-miller/

[52] EU Artificial Intelligence Act (Regulation (EU) 2024/1689), “Article 99: Penalties,” and related enforcement provisions applicable from August 2, 2026, AI Act Explorer. https://artificialintelligenceact.eu/article/99/

[53] Stanford Institute for Human-Centered Artificial Intelligence, “The 2026 AI Index Report — Economy Chapter,” Stanford University, 2026 (investment, adoption, consumer-surplus, and capex data). https://hai.stanford.edu/ai-index/2026-ai-index-report/economy

[54] Unite.AI, “China Weighs Walling Off Its Best AI and Chip Designs,” July 21, 2026 (Ministry of Commerce consultations; tiered open-weight regime; H200 case-by-case licensing context; June 2026 Anthropic order context). https://www.unite.ai/china-weighs-walling-off-its-best-ai-and-chip-designs/

[55] TechJournal, “US AI Export Controls 2026: The Anthropic Ban Explained,” June 2026 (OSTP industrial-scale distillation accusation, April 2026; DeepSeek record funding round; open-weight substitution effects). https://techjournal.org/us-ai-export-controls-anthropic-ban-2026