Introduction: The Idea That Crossed the Boundary

For most of the past decade, the direction of enterprise computing seemed irreversible. Servers left corporate buildings. Applications became software-as-a-service. Databases migrated into hyperscale clouds. Companies were told, in boardroom after boardroom and analyst report after analyst report, that owning hardware was inefficient, that maintaining private infrastructure was expensive, and that the future belonged to Amazon Web Services, Microsoft Azure, Google Cloud, and an ever-expanding universe of externally hosted software. The governing principle of enterprise technology became remarkably simple, so simple that it hardly needed to be defended anymore: move computing away from the corporation and consume it as a service. The corporate datacenter came to be described in the past tense, the way one might describe the mainframe room or the typewriter pool—an artifact of a less enlightened age of information technology.

Artificial intelligence may complicate that trajectory, and the reason begins with something far more sensitive than the location of a server. The reason begins with a question that no chief information officer of 2015 ever had to ask: what happens to an idea after someone types it into a prompt?

Consider how thoroughly the most valuable thinking inside modern organizations now passes through frontier models. Scientists use them to reason through experimental results that have not yet been published. Engineers ask models to analyze proprietary designs whose competitive value depends on their secrecy. Lawyers paste portions of contracts, privileged memoranda, and litigation strategies into AI systems in order to check reasoning, tighten language, or anticipate objections. Executives use models to evaluate merger possibilities before a single banker has been retained. Entrepreneurs test business plans, product names, patent concepts, pricing strategies, and fundraising presentations against a machine that has read more of the world than any human advisor ever could. Venture capitalists ask models to evaluate companies that have not yet publicly disclosed what they are building. In each of these situations, the prompt itself contains information whose commercial value depends precisely upon somebody else not knowing it yet, and in each of these situations the act of asking the question moves that information across a computational boundary that the organization does not fully control.

Consider a scientist who has discovered an unusual experimental result but has not yet published a paper or filed a patent. The scientist asks a frontier model whether the finding appears novel and, in order to receive a useful answer, pastes the most important details into the prompt—the anomaly, the conditions under which it appeared, the hypothesis that might explain it. Or consider the founder of an early-stage company who asks a model to improve a confidential business proposal containing the company’s product architecture, customer strategy, prospective partnerships, and financing requirements, a document that constitutes very nearly the entire enterprise value of the firm. Even when nothing improper occurs—and in the overwhelming majority of cases nothing improper occurs—the organization has created a new information-governance question that did not exist before generative AI: a highly valuable idea has crossed the company’s own computational boundary, and the company must now reason carefully about where that idea went, who could theoretically observe it, how long it persists in logs, and which jurisdiction’s law governs it.

The same problem can be illustrated through the domain-name industry, where the calendar of 2026 supplied an unusually concrete example. ICANN’s 2026 New Generic Top-Level Domain application window closed at 23:59 UTC on August 12, 2026, after receiving more than 1,600 primary applications, more than 1,100 of which included replacement strings as fallbacks [1]. The final list of applied-for strings remains sealed until Reveal Day, expected in the autumn of 2026, which means that during the entire fifteen-week application window every prospective applicant’s list of candidate strings constituted genuinely valuable competitive intelligence [2]. Now imagine a legal team representing a prospective applicant in the months before that deadline, holding a confidential list of strings the client was considering proposing. The legal team might reasonably have wanted an AI model to rank those strings, estimate commercial demand for each, improve the business case, analyze the likelihood of objections, or draft portions of an application that carried a $227,000 evaluation fee per string [2]. Yet the prospective list itself was exactly the kind of information a rival would pay to see, because a competitor who knew the strings before submission could alter its own application strategy, prepare contention-set tactics, or file overlapping applications. The tool that would have been most useful for preparing the application was also the tool whose careless use could have compromised it.

The concern should be stated carefully, because it is easy to state it incorrectly and thereby discredit it. It would be inaccurate—and unfair to the major providers—to suggest that enterprise AI services simply take one company’s confidential prompt and reveal it to another user. OpenAI states that business and API inputs and outputs are not used to train its models by default, and offers retention controls including zero-data-retention arrangements for qualifying use cases [8]. Anthropic makes a similar commitment for its commercial products, stating that inputs and outputs from its enterprise offerings are not used for training by default [9]. Microsoft states that prompts and completions processed through its managed Azure AI services are neither made available to other customers nor shared with the underlying model providers [10]. AWS states that Amazon Bedrock inputs and outputs are not shared with model providers and are not used to improve foundation models [11], and Google publishes similar enterprise-data protections for Vertex AI, including commitments that customer prompts and tuning data are not used to train Google’s models or anyone else’s [12]. These are serious commitments made by serious companies, backed by contracts, certifications, and audit regimes.

But contractual assurance does not eliminate the strategic issue; in an important sense it clarifies it. For a corporation protecting an unpublished scientific discovery, a trade secret, a privileged legal communication, an acquisition strategy, a defense design, a pharmaceutical compound, proprietary source code, or an unannounced domain application, the question increasingly becomes larger than “does the provider promise not to train on this?” Management may want to know exactly where information is processed and in which physical facility; how long logs exist and under what deletion guarantees; which personnel or subprocessors could theoretically access them under what circumstances; which jurisdiction governs them and what happens when a foreign court issues a demand; which security controls apply and who audits them; what happens to the information during legal discovery in litigation the company is not even a party to; whether an external API remains available during an outage or a geopolitical crisis; how prompt-injection attacks against agentic systems are contained; and whether the company can continue operating if a provider suddenly changes its pricing, its models, its terms of service, or its access policy. Each of these questions has a contractual answer, but the sum of the questions is architectural, and architecture is precisely what a pure consumption relationship does not let the customer control.

There is an old biblical metaphor that unexpectedly captures part of this emerging corporate instinct. Matthew 6:3 reads: “But when you give to the needy, do not let your left hand know what your right hand is doing.” The verse concerns humility in charitable giving, not cybersecurity, and it would be irreverent to pretend otherwise. Yet as a metaphor for information compartmentalization it is remarkably contemporary. In the AI economy there will increasingly be situations in which a corporation concludes that its “left hand”—an external model, a cloud service, a contractor, an API, or even another internal department—does not need to know everything its “right hand” knows. The objective is not secrecy for secrecy’s sake. It is selective disclosure, practiced deliberately and engineered into the computing architecture itself. Public information can be sent freely to public frontier systems. Moderately sensitive information can run through enterprise cloud environments protected by contract, encryption, and private networking. Highly sensitive information may never leave a dedicated corporate AI environment at all. The more capable AI becomes, paradoxically, the more valuable this distinction becomes, because companies will want models to reason over categories of information that previously would never have been exposed to any external computing service under any terms.

The empirical urgency of the problem is no longer hypothetical, because employees have already voted with their keyboards. Research summarized in 2025 and 2026 found that more than 80 percent of workers use AI tools their employers never approved, and IBM’s 2025 Cost of a Data Breach study found that one in five organizations had already experienced a breach linked to unsanctioned “shadow AI” [13]. Cyberhaven’s analysis of the actual usage patterns of roughly seven million workers found that 34.8 percent of the corporate data employees feed into AI tools is now classified as sensitive, up from 10.7 percent only two years earlier, with source code, research-and-development material, and sales data leading the categories of exposure [14]. The canonical cautionary tale remains Samsung’s semiconductor division, where in the spring of 2023 engineers pasted proprietary source code, meeting transcripts, and chip-yield test sequences into a consumer chatbot within weeks of being permitted to use it, prompting a company-wide ban [15]. The lesson of these numbers is not that employees are careless; it is that frontier intelligence is so useful that information flows toward it under almost any policy regime, which means the only durable governance strategy is to control where that intelligence runs.

A revealing example of what that control can look like appeared on September 10, 2026. Latham & Watkins, one of the world’s largest law firms with approximately $8.3 billion in revenue, disclosed through reporting first published by the Financial Times an unusually ambitious approach to enterprise AI: the firm has purchased Nvidia GPU servers over the past several years, is fine-tuning Nvidia’s Nemotron open-weight models with its own machine-learning engineers, employs a technology organization of more than 900 specialists, and operates its private infrastructure in secure datacenter space that only its own staff can access [3]. The Financial Times described the move as the first public example of a major law firm buying its own AI hardware and customizing models, a significant departure from the prevailing Big Law pattern of relying primarily on external AI vendors [3]. Security of highly confidential client information is one rationale; the firm’s chief information officer, Rene Mendoza, was explicit that certain client information is so sensitive that the firm does not want to place it with any cloud vendor at all [3]. But flexibility and the economics of rapidly growing AI consumption are equally important, as Mendoza made clear in a single sentence that could serve as the thesis of this entire paper:

“We are not hitching our wagon to one particular company.”

— Rene Mendoza, Chief Information Officer, Latham & Watkins [4]

Latham’s decision matters far beyond the legal industry, and it matters precisely because of what Latham is not. A law firm is not a hyperscaler. It is not an AI laboratory. It does not sell software, chips, or datacenter capacity. Its principal products are legal judgment, advice, negotiation, and representation—activities as far from silicon as professional life gets. Yet this firm has concluded that artificial-intelligence infrastructure may become sufficiently important to justify directly controlling GPUs and customized models, accepting in exchange the burdens of cybersecurity, server operations, and specialist staffing that the cloud era taught corporations to shed [7]. That conclusion should attract the attention of banks, pharmaceutical companies, defense contractors, insurers, engineering firms, hospitals, industrial manufacturers, scientific laboratories, energy companies, governments, and any corporation whose proprietary information constitutes a large fraction of its enterprise value. Nor is Latham obviously alone: Kirkland & Ellis, the largest American firm by revenue, posted job listings in 2026 seeking AI infrastructure directors to manage on-premise GPU clusters, suggesting that the pattern is already propagating through the profession [6].

Technology vendors themselves are preparing for this possibility, and none more explicitly than the company whose chips power the entire buildout. Nvidia now markets enterprise “AI factory” architectures designed for deployment in company datacenters and colocation facilities, built around Blackwell-generation GPUs, enterprise networking, accelerated storage, Kubernetes orchestration, Nvidia AI Enterprise software, and inference microservices, and it describes these systems as capable of running generative, agentic, industrial, and scientific AI workloads on infrastructure controlled by the enterprise itself [18]. Its founder and chief executive, Jensen Huang, has spent 2026 telling world leaders at Davos that artificial intelligence should be treated as essential infrastructure in the same category as electricity and roads, arguing that every country—and by extension every serious institution—should refine and control its own intelligence rather than merely import it [18]. Latham’s reported hardware of choice, the Nvidia H200, with the firm actively evaluating Blackwell and Vera Rubin successors, shows how quickly this vendor strategy is finding non-technology buyers [6].

This development creates an apparent historical contradiction that deserves to be stated plainly. The cloud era was built on the argument that enterprises should stop owning computers. The AI era may encourage some of the world’s largest enterprises to start owning important computers again. That does not mean the cloud disappears—quite the opposite, as the capital-expenditure figures examined throughout this paper make overwhelming clear. AWS, Microsoft, Google, Oracle, and the other hyperscalers will remain indispensable providers of enormous training clusters, elastic inference capacity, data services, distributed storage, cybersecurity, model APIs, and agentic platforms, and their combined 2026 capital spending of roughly $760 billion represents the largest private infrastructure investment program in the history of capitalism [20]. But the architecture is changing from cloud-first to placement-aware. Instead of asking “why shouldn’t this workload go to the cloud?”, enterprises are beginning to ask “where should this particular model, dataset, agent, or inference request live?” That question—modest in phrasing, enormous in consequence—is the beginning of Compute Repatriation.


Why I Chose the Title “Compute Repatriation”

I chose the term Compute Repatriation because repatriation describes the return of something important to the jurisdiction, ownership, or control from which it previously departed, and that is precisely the movement this paper documents. During the cloud era, corporations deliberately exported much of their computing infrastructure to outside providers, and they did so for reasons that were entirely rational at the time: elasticity was worth more than ownership, operating expenditure was preferable to capital expenditure, and the hyperscalers could operate general-purpose infrastructure more efficiently than any individual enterprise ever could. Artificial intelligence now creates reasons to bring selected portions of that capability back under direct enterprise control—not necessarily into an office building, and not necessarily onto hardware the company physically owns, but inside the corporation’s technical, contractual, security, and governance perimeter, where the corporation rather than its vendor holds the decision rights. Sensitive inference, proprietary models, high-value datasets, autonomous agents, GPU capacity, and institutional knowledge can therefore be “repatriated” without requiring the enterprise to abandon cloud computing, in the same way that a nation can repatriate its gold reserves without abandoning international trade.

The title also captures the paper’s central paradox, which is that both halves of the story are true simultaneously. AI is driving the largest hyperscale cloud expansion in history—combined capital expenditure by Microsoft, Amazon, Alphabet, and Meta is expected to reach approximately $760 billion in 2026, up from roughly $413 billion in 2025 [20]—and at the very same time AI is making private computing strategically valuable again, with 86 percent of CIOs telling the Barclays survey that they plan to move at least some workloads out of the public cloud, the highest share ever recorded [24]. Compute Repatriation therefore does not predict the death of AWS, Azure, or Google Cloud, a prediction that would be refuted by every earnings report of 2026. It describes the emergence of a hybrid equilibrium in which public cloud, private cloud, dedicated GPU infrastructure, open-weight models, edge systems, and frontier-model APIs coexist, each assigned the workloads for which it is genuinely superior. The question for 2030 may no longer be whether a corporation is “in the cloud.” The more consequential question will be: which intelligence does the corporation rent, and which intelligence does it insist on owning or controlling itself?


Section 1: The Cloud-to-AI Reversal

The first task of this paper is to explain why artificial intelligence, alone among the workloads of the past twenty years, has the power to bend the trajectory of enterprise infrastructure. Databases did not do it. Enterprise resource planning did not do it. Video, mobile, analytics, and the entire software-as-a-service revolution did not do it; each of these made the case for centralization stronger. To understand why AI is different, it helps to see that AI changes not only the economics of computing but the meaning of computing—what a computer is doing when it touches corporate information, and therefore what it means to let someone else’s computer do it.


1.1 From Server Repatriation to Intelligence Repatriation

Traditional cloud repatriation, the phenomenon that infrastructure analysts have tracked for a decade, generally meant moving an application or a database out of a public cloud because its operating cost, performance profile, or regulatory requirements no longer justified external hosting. That older form of repatriation is itself accelerating and is worth pausing on, because it forms the economic backdrop against which AI repatriation is unfolding. The Barclays CIO Survey found that 86 percent of chief information officers planned to move at least some workloads from public cloud back to private cloud or on-premises environments, the highest rate the survey has ever recorded [24]. Flexera’s State of the Cloud research found that roughly one-fifth of workloads and data originally moved to the public cloud have already been pulled back [26], and its 2026 report placed 73 percent of organizations on deliberately hybrid estates [25]. Yet the same body of research shows that only around 8 percent of organizations intend to repatriate entire workload portfolios [25], and Gartner’s forecast of public cloud spending growing past $723 billion demonstrates that the cloud itself continues to expand vigorously [24]. The correct reading of these superficially contradictory statistics is not that enterprises are fleeing the cloud but that they have stopped treating the cloud as a default—they are re-sorting workloads one by one according to cost, control, and sensitivity, which is exactly the behavioral precondition for what follows.

AI expands this re-sorting exercise into something categorically larger, because what is being reconsidered is no longer merely the location of software. It is the location of reasoning. When an enterprise owns or controls the infrastructure running an AI model, the boundary around its intellectual property changes in kind rather than in degree. Proprietary documents can be retrieved locally, so that the retrieval index—itself a compressed map of everything the company knows—never leaves the building. Model prompts can remain inside controlled networks, so that the questions executives ask, which often reveal strategy more clearly than any document, are never observable by an outside party. Fine-tuning datasets can be isolated, so that the distilled essence of decades of institutional experience is not entrusted to a third party’s custody chain. Internal agents can access privileged corporate systems without every transaction being routed through an external inference service that logs, meters, and mediates it. Compute, in short, becomes part of information governance, and the chief information officer’s org chart begins to merge with the general counsel’s.

This reframing explains why the a16z analysis of cloud economics, published by Sarah Wang and Martin Casado in 2021 under the title “The Cost of Cloud, a Trillion Dollar Paradox,” reads today as a prophecy that arrived one technology cycle early. Wang and Casado calculated that cloud spending was consuming roughly half the cost of goods sold at scaled public software companies and estimated that the resulting margin suppression was weighing on hundreds of billions of dollars of aggregate market value, concluding that companies should plan for repatriation before lock-in forecloses the option [27]. In 2021 that argument concerned commodity compute and storage. In 2026 the same arithmetic applies to inference—except that the numerator is growing far faster than commodity workloads ever did, and the workload being priced is not file storage but the company’s own thinking.


1.2 Training Created the Cloud Boom; Inference Could Produce a More Distributed Architecture

Frontier-model training naturally favors enormous centralized clusters, and nothing in this paper disputes that centralization or predicts its end. Training leading models requires tens or hundreds of thousands of accelerators connected by exotic high-bandwidth networks, gigawatt-class power commitments, and capital budgets that only a handful of institutions on Earth can sustain; Jensen Huang has told policymakers that a single gigawatt of AI infrastructure now costs roughly $50 to $60 billion to build [19]. The scale of the 2026 buildout confirms how thoroughly training and frontier inference have concentrated: Amazon guided toward roughly $200 billion of capital expenditure for 2026, Alphabet raised its ceiling toward $205 billion at its second-quarter earnings, Microsoft is tracking toward $190 billion on a calendar basis, and Meta raised its range to $130–145 billion, for a combined figure near $760 billion, up approximately 77 percent from 2025 [20], on top of Oracle’s roughly $50 billion and the half-trillion-dollar Stargate program running alongside the hyperscaler budgets [21]. Analysts at Barclays and elsewhere have begun modeling negative free cash flow for some of these companies in 2027 and 2028 as the buildout front-runs the revenue it is meant to serve [22]. Few corporations will ever reproduce that infrastructure internally, and none should try.

Inference is different, and the difference is architectural rather than incidental. A bank does not need 100,000 GPUs to run a specialized fraud-detection model; it needs predictable, low-latency capacity close to its transaction systems. A law firm does not need a frontier training cluster to operate an internal legal assistant over its document archive; Latham & Watkins is doing precisely this work on a handful of multi-GPU servers [3]. A pharmaceutical company can run specialized scientific models on hundreds or thousands of accelerators rather than hundreds of thousands. An industrial company can operate smaller models adjacent to its factories, and a hospital network can deploy clinical models entirely within environments it controls. The governing distinction of the coming architecture can be stated in one line: training centralizes AI, while inference can decentralize it. And inference is winning the budget. Gartner’s August 2026 forecast marked the first year in which enterprise inference spending surpassed training spending in AI-optimized cloud infrastructure, with $23.3 billion of a $42 billion market flowing to inference [39], while Morgan Stanley projects that inference will constitute 70 to 80 percent of all AI compute spending by 2027 [41]. As inference becomes the dominant and permanent source of AI consumption, the architecture of artificial intelligence could become considerably more distributed than the architecture that initially created frontier models—just as electricity generation centralized into utilities while electricity consumption spread into every wall socket in the world.


1.3 Open-Weight Models Change the Make-or-Buy Decision

Compute Repatriation becomes significantly more plausible—arguably it becomes possible at all—because corporations no longer have to build foundation models from scratch in order to operate intelligence privately. An enterprise can take an open-weight or commercially licensable model, customize it, fine-tune adapters around proprietary datasets, connect it to retrieval systems over internal archives, impose corporate guardrails, and operate the resulting system on infrastructure it controls, all without spending a single dollar on pre-training. Stanford’s 2026 AI Index documents how narrow the capability sacrifice has become: the gap between the best closed models and the best open-weight models, which nearly vanished in 2024, stood at only about 3.3 percent on leading evaluations in 2025, and the gap between the top American and Chinese models—many of the latter released with open weights—has effectively closed to within roughly 2.7 percent [32]. For an enormous range of enterprise tasks, an open model fine-tuned on the company’s own data is not a compromise; it is the better tool, because it knows things no frontier model was ever shown.

That changes the economic calculation at the heart of enterprise AI strategy. The original enterprise AI decision was: which frontier-model API should we buy? The emerging decision is: for which tasks should we buy frontier intelligence, and for which tasks should we operate intelligence ourselves? Recent enterprise activity suggests this distinction is becoming commercially decisive rather than merely theoretical. Reporting on the open-model ecosystem in September 2026 found that closed models’ share of queries routed through the OpenRouter aggregation platform had fallen from roughly three-fifths at the start of the year to about one-quarter; that DoorDash, Siemens, and Airbnb had shifted portions of their workloads toward open models for cost savings; and that AT&T’s open-model share of internal AI workloads had risen from 20 percent in May to 40 percent, with a path to 60 percent and reported savings of up to 80 percent on the affected workloads [45]. Latham & Watkins’s fine-tuning of Nvidia’s Nemotron open-weight family [5] belongs to the same movement: the open-weight model has become the vehicle by which intelligence travels from the frontier laboratories into buildings the laboratories will never see.


1.4 Latham & Watkins as an Early Signal

Latham & Watkins provides an unusually clean case study because its decision cannot be dismissed as an AI company’s ordinary infrastructure spending, a startup’s marketing stunt, or a technology vendor talking its own book. The firm began developing its Nvidia server infrastructure roughly three years before the September 2026 disclosure, has acquired multiple H200 GPU systems housed in a third-party datacenter accessible only to Latham personnel, and is actively evaluating Nvidia’s newer Blackwell and Vera Rubin generations [6]. Its technology organization now exceeds 900 people, of whom roughly 100 work on AI specifically, including machine-learning engineers, AI engineers, and lawyers with coding expertise [6]. The reported cost of buying and operating GPU infrastructure at this scale, together with the specialists required to run, maintain, and protect it, can reach into the tens of millions of dollars per year, potentially rising toward hundreds of millions as the technology advances [7]—a sum that a partnership answerable to its own equity holders does not spend on aesthetics.

What makes the case genuinely instructive, however, is what Latham did not do. It did not cancel its commercial AI relationships; the firm continues to use tools from Harvey and Legora as well as offerings from OpenAI and Anthropic, and it explicitly stated that its server buildout is not intended to reduce reliance on its legal-AI vendors [6]. Latham is not choosing between all cloud and no cloud. It is creating optionality. Sensitive workloads can remain inside a controlled environment whose door only its employees can open. Other tasks can be routed to commercial models where those models are superior. Models can be selected task by task according to security, capability, latency, and cost, and—crucially—the existence of a credible internal alternative becomes negotiating leverage over every outside AI provider the firm deals with, particularly as vendors move away from subsidized token pricing toward consumption-based billing [3]. This is the template: not exit, but optionality; not ideology, but placement. The remainder of this paper argues that what one law firm has assembled in a locked cage of a colocation facility is a small-scale model of what the sophisticated enterprise of 2030 will operate at every scale.


1.5 The Five-Layer AI Economy Moves Inside the Corporation

Compute Repatriation also connects naturally to the five-layer structure of the AI economy that Jensen Huang sketched publicly at Davos in January 2026, when he described artificial intelligence as an economic stack whose bottom layer is energy, above which sit chips, then datacenters and cloud infrastructure, then models, and finally applications and agents at the top [18]. During the first phase of the AI era, the ordinary corporation participated in this economy at exactly one layer: it was a customer at the top, consuming applications built by others upon infrastructure owned by others. Compute Repatriation is unusual—and unusually consequential—because it drives the corporation downward through almost the entire stack simultaneously.

At Layer 1, energy, major private AI installations create new power and cooling requirements that corporate real-estate and facilities organizations have not confronted since the mainframe era, and that increasingly require liquid cooling, high-density electrical service, and long-term power procurement. At Layer 2, chips, corporations become direct purchasers or long-term lessees of Nvidia, AMD, or specialized accelerators, exposed for the first time to allocation queues, generational transitions, and the supply commitments—Nvidia’s own purchase obligations reached $279 billion in mid-2026, much of it memory for the Vera Rubin generation—that shape the silicon market [16]. At Layer 3, datacenters, companies require colocation capacity, secure private cages of the kind Latham leases [7], private networking, and AI-class storage. At Layer 4, models, businesses choose among frontier proprietary models, open-weight models, customized fine-tunes, and small specialized models, becoming model portfolio managers rather than single-vendor licensees. And at Layer 5, applications and agents, proprietary business workflows become the final and least copyable source of competitive differentiation. The corporation ceases to be merely the consumer at Layer 5 and begins selectively internalizing Layers 2, 3, and 4—which is why Compute Repatriation is best understood not as an IT procurement trend but as a structural migration of the AI economy itself into the enterprise.


Section 2: Security, Economics, Latency, and Data Sovereignty—The Four Engines of Repatriation

If Section 1 described what Compute Repatriation is, this section examines why it happens—the four distinct forces that, individually or in combination, push particular workloads back inside the enterprise perimeter. It matters that there are four forces rather than one, because the four operate on different workloads, answer to different corporate officers, and mature on different timescales. Security concerns move privileged and trade-secret workloads first and answer to the general counsel and the chief information security officer. Economics moves high-volume, predictable inference and answers to the chief financial officer. Latency moves machine-speed and physical-world workloads and answers to the heads of product and operations. Sovereignty moves regulated and jurisdictionally sensitive workloads and answers to regulators who do not work for the company at all. A repatriation thesis built on any single one of these forces would be fragile; a thesis built on all four is structural.


2.1 Intellectual Property Becomes an Infrastructure Question

Artificial intelligence changes the meaning of corporate cybersecurity because the information being protected increasingly includes reasoning context—the live, in-motion representation of what the organization is thinking about—rather than only data at rest. Traditional security protects files, passwords, databases, and communications, and an entire industry of controls has grown up around those objects. AI security must additionally protect prompts, retrieval context, embeddings, vector databases, model adaptations and fine-tuned weights, system instructions, agent memory, tool permissions, internal reasoning workflows, and machine-generated intermediate outputs, none of which existed as security objects five years ago and several of which—embeddings especially—are lossy but substantial reconstructions of the underlying confidential material. An employee may never download a confidential file or email it outside the company; yet if that employee pastes its most sensitive contents into an unauthorized AI service, the information has still crossed an organizational boundary, and it has done so through a channel that most data-loss-prevention tooling, built to inspect attachments rather than conversational text typed into an encrypted browser session, cannot see.

The scale of this exposure is now well documented, and the numbers deserve to be internalized rather than skimmed. More than 80 percent of workers report using unapproved AI tools, and one in five organizations has already suffered a breach connected to shadow AI [13]. The share of sensitive material within employee AI inputs has more than tripled in two years to 34.8 percent, led by source code and research material [14]. IBM’s 2025 breach-cost research attributes an additional $670,000 of average breach cost to organizations with high levels of shadow AI, and found that 97 percent of AI-related incidents occurred at companies lacking proper AI access controls [15]. The rise of shadow AI—employees independently adopting consumer models because they are more convenient and often more capable than approved corporate tools—is therefore precisely analogous to the shadow-IT problem of the previous generation, with one aggravating difference: the unsanctioned Dropbox account of 2013 stored corporate data, while the unsanctioned chatbot of 2026 reads it, reasons about it, and remembers the conversation. The organizational response that actually works, as the Samsung episode and dozens of quieter successors have shown, is not prohibition, which employees route around, but substitution: giving the workforce sanctioned intelligence, inside a governed environment, that is good enough to make the ungoverned alternative unattractive [15]. Building that governed environment is, in miniature, the entire project of Compute Repatriation.


2.2 Enterprise AI Providers Are Strengthening Privacy—But That Does Not End Repatriation

The hyperscalers and frontier laboratories understand this concern intimately, and it would be a serious analytical mistake to portray them as indifferent to it. OpenAI states that its enterprise and API customer data is excluded from model training by default and offers administrative retention controls up to zero-data-retention arrangements for qualifying endpoints [8]. Anthropic states that inputs and outputs from its commercial offerings are not used for training by default [9]. Microsoft documents that prompts and completions submitted to its managed Azure OpenAI services are not available to other customers, are not shared with OpenAI, and are not used to improve the underlying models [10]. AWS states that Amazon Bedrock does not share customer inputs or outputs with third-party model providers and does not use them to train foundation models [11], and Google makes equivalent commitments for Vertex AI enterprise customers [12]. Beyond these defaults, the providers now sell dedicated capacity, customer-managed encryption keys, private networking, regional processing guarantees, and compliance attestations that would have seemed exotic in 2023. These protections are real, they are contractually enforceable, and for a large majority of corporate AI workloads they are sufficient.

But the strengthening of provider privacy does not end the repatriation question; it sharpens it, because it clarifies what remains unpurchasable. The argument of this paper is emphatically not that cloud AI is insecure—by most measures the hyperscalers operate the most sophisticated security organizations on Earth. The argument is that different organizations will assign different values to maximum control, and that for a specific class of information, contractual privacy and architectural isolation are not the same good. A contract is a promise backed by remedies; an architecture is a physical and logical fact. For certain workloads, the promise is plainly sufficient. For others—the unpublished discovery, the privileged litigation strategy, the weapons design, the board’s acquisition deliberations—boards, regulators, clients, governments, or customers may demand the fact. Latham’s CIO articulated the distinction with the bluntness of a practitioner rather than the hedging of a policy paper, explaining that some client information is simply too sensitive to place with any cloud vendor under any terms [3]. Privacy, in short, becomes a spectrum rather than a binary, and the enterprise’s task becomes mapping each information class to the correct point on that spectrum—which is an architecture problem, not a procurement problem.


2.3 The Economics of Inference May Favor Ownership at Scale

The second force is financial, and in the long run it may prove the most powerful, because it operates silently on every invoice. Cloud computing is attractive because organizations can buy computing incrementally, converting capital expenditure into operating expenditure and risk into flexibility; owning infrastructure requires capital, technical staff, power, maintenance, utilization discipline, and replacement cycles. Nothing about AI repeals those trade-offs. What AI changes is the magnitude and the shape of the consumption to which they apply, because AI inference produces recurring computational demand of a kind no previous enterprise workload has ever generated.

The numbers describing this demand explosion have become genuinely startling. Goldman Sachs Research projects that agentic AI will multiply global token consumption roughly 24-fold between 2026 and 2030, to approximately 120 quadrillion tokens per month [38]. Research from Microsoft and Stanford’s Digital Economy Lab found that agentic tasks consume roughly one thousand times the tokens of a standard chat interaction, because a single delegated task can trigger dozens or hundreds of model calls as an orchestrator decomposes work, invokes tools, validates outputs, and retries failures [40]. EY’s cost analysis found that a single customer-service interaction rose from roughly $0.04 under the linear chatbot architectures of 2023 to roughly $1.20 under the orchestrated agentic architectures of 2026—a thirty-fold increase per interaction even as the price per token collapsed [39]. Gartner forecasts worldwide AI spending reaching approximately $2.59 trillion in 2026, and Morgan Stanley expects inference to absorb 70 to 80 percent of AI compute spending by 2027, converting AI from a one-time capital project into a permanent, usage-driven operating cost [41]. The paradox at the center of these figures—unit prices falling 60 to 70 percent per year per token while total bills rise relentlessly [38]—was named by Microsoft’s chief executive on the day a low-cost Chinese model appeared:

“Jevons Paradox strikes again.”

— Satya Nadella, Chief Executive Officer, Microsoft [40]

At sufficiently high and predictable utilization, this consumption profile changes the rent-versus-own calculation in exactly the way the a16z analysis predicted for commodity cloud a technology cycle ago [27]. A corporation operating a handful of AI queries per day has no business owning GPUs, and never will. A corporation whose employees, customers, software agents, robots, factories, and scientific instruments together generate billions of inference transactions—steady, forecastable, around-the-clock demand—may discover that owning or leasing dedicated accelerators provides materially lower marginal inference cost than perpetually paying retail API prices, particularly for the routine majority of tasks that do not require frontier capability. Enterprise cost data already validates the routing logic underlying this conclusion: a 2026 analysis of 2.4 billion enterprise API calls found that organizations running tiered model architectures achieved a blended cost of $2.31 per million tokens, while organizations routing everything to frontier models paid $18.40—an eight-fold difference produced purely by placement decisions [39]. The decision begins to resemble the classic economic distinction between renting capacity and owning productive capital: rent the peaks, own the baseload, and know precisely where the crossover sits, because at agentic scale the crossover arrives years earlier than chatbot-era budgeting assumed.


2.4 Latency Turns Compute Placement Into Product Design

The third force is physical, and it is the one force that no contract, no price cut, and no privacy commitment can address, because it is imposed by the speed of light and the depth of the software stack. Agentic AI makes latency compound in a way that chatbot AI never did. A human waiting four seconds for an answer tolerates the delay and may not even notice it; a software agent that must execute fifty sequential inference operations converts that same four seconds into more than three minutes of dead time, and an agent pipeline that fans out into sub-agents multiplies the penalty again. Industrial robots, autonomous vehicles, trading systems, manufacturing control loops, cybersecurity response agents, and interactive voice systems require response times measured in tens of milliseconds, budgets that a round trip to a distant cloud region consumes before the model has generated its first token. For these applications, the geographic distance between data, models, tools, and accelerators is not an infrastructure detail; it is a product specification, and compute placement becomes product design.

This is why Compute Repatriation does not terminate at the corporate datacenter, and why the four-zone architecture developed in Section 5 requires an edge zone at all. The same logic that pulls sensitive inference from the public cloud into a private environment pulls time-critical inference from the private environment outward to factories, laboratories, warehouses, hospitals, retail locations, vehicles, and end-user devices, where small and mid-sized models—whose quality, as the Stanford AI Index documents, now approaches frontier quality on bounded tasks [32]—can answer in the time the application actually has. The endpoint of the latency force is an enterprise whose intelligence is layered like its caching: frontier reasoning far away and rarely, specialized reasoning nearby and often, reflexive reasoning on-device and constantly.


2.5 Data Sovereignty Evolves Into Intelligence Sovereignty

The fourth force is jurisdictional, and unlike the first three it is being driven as much by governments as by enterprises. Data sovereignty historically asked where information was stored, and an entire regulatory apparatus—GDPR residency provisions, national data-localization statutes, sectoral rules for health and financial records—grew up around the location of bits at rest. AI introduces a deeper issue that this apparatus was never designed to reach: where is information interpreted? A company may lawfully store its data in Germany while sending selected excerpts elsewhere for inference; another may keep a database entirely on-premises while an externally hosted agent continuously reasons over it through an API. In both cases the bits at rest never moved, and in both cases the intelligence applied to them crossed a border. Future regulation is therefore likely to distinguish among data residency, model residency, inference residency, agent residency, training residency, and human-access residency, and the corporate objective will increasingly be not merely to control where bits are located but to control where intelligence is applied to those bits.

The market has begun pricing this evolution with remarkable speed. In January 2026, AWS launched the AWS European Sovereign Cloud, a physically and logically separate cloud located entirely within the European Union, operated under independent governance, backed by a planned investment exceeding €7.8 billion in Germany alone, with sovereign Local Zones announced for Belgium, the Netherlands, and Portugal [42]. Microsoft had announced its own Sovereign Public Cloud across all European regions in June 2025, with data processed inside the EU Data Boundary under the control of European personnel, alongside national partner clouds such as Bleu in France and Delos Cloud in Germany, and Gartner projects European sovereign-cloud spending to triple to roughly $23.1 billion by 2027 [44]. Skeptics correctly note the unresolved legal question of whether any U.S.-owned sovereign offering can fully insulate European data from American legal process under the CLOUD Act and FISA [43]—an unresolved question that itself pushes the most cautious institutions one ring further inward, toward infrastructure they control outright. At the level of nations, the same instinct has acquired a name and an evangelist: Nvidia’s chief executive has spent three years telling governments that intelligence is a sovereign resource, framing the argument at Davos in the simplest possible terms:

“AI is infrastructure.”

— Jensen Huang, Founder and Chief Executive Officer, Nvidia, World Economic Forum, Davos, January 2026 [18]

What Huang urges upon nations—produce your own intelligence rather than merely importing it, because your language, your data, and your institutional knowledge are assets no one else should hold [47]—is structurally identical to what this paper describes corporations beginning to practice. Sovereign AI is Compute Repatriation at the scale of the state; Compute Repatriation is sovereign AI at the scale of the firm. The two movements share vendors, architectures, and rhetoric, and they reinforce each other: every sovereign-cloud region, evaluated open model, and enterprise AI factory built for one constituency lowers the cost curve for the other.


Section 3: The Industries Most Likely to Become Private AI Operators

Repatriation will not arrive uniformly across the economy, and pretending otherwise would weaken the thesis. It will arrive first and hardest in industries where three conditions coincide: the information is extraordinarily valuable, the value depends on nondisclosure or is protected by regulation, and the volume of AI reasoning applied to that information is large enough to justify dedicated infrastructure. This section examines six settings where those conditions already hold—law, finance, healthcare, defense, pharmaceutical and scientific research, and, perhaps surprisingly, startups—and asks in each case what form private AI operation is likely to take.


3.1 Law: Privilege Becomes Computational Architecture

Law firms are natural early adopters of Compute Repatriation because confidentiality is not a feature of the product they sell; it is the product. A major law firm is a concentrated repository of other institutions’ most dangerous information: litigation strategies, merger plans before announcement, internal investigations, unpublished regulatory arguments, privileged communications, and corporate facts that could move markets if disclosed a day early. The attorney-client privilege and the work-product doctrine—legal constructs centuries in the making—exist precisely to wall this information off from the world, and courts have repeatedly held that careless handling can waive the wall entirely. Once AI becomes sufficiently useful that lawyers want models to read, summarize, cross-reference, and reason over enormous quantities of this material—and 2026 practice confirms that they emphatically do—the question of which computers perform that reasoning stops being an IT question and becomes a question of professional responsibility. Privilege, in other words, becomes computational architecture.

This is why the Latham & Watkins disclosure examined in Section 1 is best read not as an anomaly but as an early indicator of where the profession’s logic leads. The firm’s configuration—private H200 servers in a locked colocation cage, open-weight Nemotron models fine-tuned by an internal team of roughly 100 AI specialists, continued parallel use of Harvey, Legora, OpenAI, and Anthropic for tasks where external tools are superior [5][6]—is a working implementation of privilege-aware routing: the most sensitive client matters stay on infrastructure only Latham employees can touch, while less sensitive work flows to whichever external tool does it best. Kirkland & Ellis’s recruitment of AI infrastructure directors for on-premise GPU clusters suggests the second-largest firm’s largest rival is drawing the same map [6], and other elite firms have chosen adjacent points on the spectrum—A&O Shearman with Harvey, Freshfields with Anthropic [4]—confirming that the industry is converging not on a single answer but on the practice of deliberate placement, which is the essence of the repatriation thesis.


3.2 Finance: Proprietary Information Has Immediate Monetary Value

Banks, hedge funds, asset managers, private-equity firms, insurers, and trading organizations represent the second obvious category, because financial information differs from other confidential information in one crucial respect: its value is immediate, quantifiable, and adversarial. A leaked litigation strategy injures a client eventually; a leaked trading strategy is arbitraged away by lunchtime. The proprietary datasets of a modern financial institution—trading strategies and their telemetry, risk models, credit decisions, acquisition pipelines, client portfolios, underwriting methods, market forecasts, internal research, and nonpublic deal information—constitute the institution’s entire edge, and every one of them is exactly the kind of material that AI systems are now uniquely good at reasoning over. The temptation to apply frontier intelligence to this material is enormous, and so is the institutional resistance to letting that material travel.

The predictable resolution is the internal model router: an automated policy layer that decides, transaction by transaction, which information may leave the enterprise and which may not, and dispatches each task accordingly—public-data research to the cheapest capable external model, moderately sensitive analysis to a contractually protected enterprise cloud, and anything touching positions, pipelines, or material nonpublic information to models the institution operates itself. Financial institutions will still consume frontier cloud models extensively, and the eight-fold cost advantage of tiered routing documented in Section 2 [39] means the router pays for itself even before its confidentiality function is counted. The consequence worth anticipating is an epistemic one: the highest-value financial AI of the coming decade may be almost entirely invisible to the public, because it will never operate through consumer-facing systems, never appear on a leaderboard, and never be benchmarked by anyone outside the institution that owns it. The observable AI economy and the actual AI economy will diverge, and finance will be where they diverge first.


3.3 Healthcare: Privacy Meets Clinical Intelligence

Healthcare presents a different version of the same structure, in which the protected asset is not competitive advantage but human dignity codified into law. Hospitals and health systems will increasingly want AI to reason across patient histories, medical imaging, laboratory results, genomic information, clinical notes, medication records, and real-time monitoring streams—in effect, to hold the entire longitudinal patient in working memory—because that breadth is precisely where clinical AI creates value that narrow tools cannot. Yet this is the most heavily regulated information in civilian life, governed by HIPAA in the United States, the GDPR and national health codes in Europe, and a thickening layer of AI-specific rules, and it is information whose subjects—patients—did not choose their data’s custodian and cannot practically revoke it.

Compliant cloud services for healthcare AI exist and will proliferate, and much clinical AI will run on them. But the direction of medical AI development cuts toward repatriation for a structural reason: the more personalized medicine becomes, the more the valuable computation involves joining a specific institution’s complete records with models tuned to that institution’s population, practices, and equipment—a workload that is high-volume, latency-sensitive at the bedside, permanent rather than experimental, and catastrophic to breach. Some health systems will conclude, workload by workload, that high-volume clinical inference belongs within dedicated environments they control, with de-identified or aggregate tasks flowing outward to cloud and frontier services. Healthcare thereby illustrates a general principle of the four-zone architecture: the zones are assigned not by industry but by information class, and a single institution will legitimately operate in all four at once.


3.4 Defense and National Security: The Air Gap Returns in AI Form

Defense represents the limiting case that proves the framework, the setting in which the repatriation logic runs to completion because the adversary is not a competitor but a nation-state and the failure mode is not embarrassment but strategic defeat. Militaries cannot assume that every intelligence analysis, targeting workflow, weapons simulation, classified sensor feed, autonomous-system decision, or battlefield agent should depend on a public internet connection to a commercial API, and no contractual privacy commitment, however sincere, addresses an environment in which connectivity itself is contested. The defense AI architecture of the 2030s will therefore be deliberately layered: frontier models developed by commercial laboratories at the unclassified outer ring; government-modified and classified models within accredited enclaves; local inference clusters at commands and installations; mobile tactical compute in vehicles and forward positions; and autonomous systems performing inference onboard, where cloud connectivity is impossible by design or by jamming. The Cold War concept of the air-gapped network evolves into the air-gapped intelligence environment—entire reasoning systems, from model weights to retrieval stores to agent frameworks, that live and die inside a classified boundary.

Two features of the defense case radiate outward to the civilian economy. First, defense procurement normalizes and funds the tooling of private AI operation—accredited model registries, offline fine-tuning pipelines, hardened inference stacks—that then becomes commercially available to banks and hospitals, exactly as classified networking and cryptography once did. Second, defense articulates without euphemism the principle that civilian institutions apply in softened form: that certain reasoning is itself a strategic asset whose location is a matter of security policy. Nvidia’s chief executive has made the national version of the argument explicit, contending that sovereign AI capability now outranks traditional strategic arsenals in importance, on the ground that every nation needs intelligence infrastructure of its own [47].


3.5 Pharmaceuticals and Scientific Research: Discovery Before Disclosure

The scientist in this paper’s opening anecdote illustrates what may be the most intellectually interesting case, because science monetizes a temporal asymmetry: a result can be enormously valuable before publication and nearly worthless as property afterward, when it becomes, by design, public knowledge. AI is increasingly capable of contributing to every stage of the discovery pipeline—molecule generation, protein analysis, materials science, experimental design, simulation, scientific coding, literature synthesis, and hypothesis generation—and the 2026 Stanford AI Index records models meeting or exceeding human performance on PhD-level science questions [30]. But every one of those contributions requires showing the model the crown jewels: the unpublished result, the proprietary assay data, the compound library, the negative results that competitors would pay to learn. A pharmaceutical company’s pre-patent research corpus is a portfolio of billion-dollar options, and the industry’s history of espionage and litigation demonstrates how contested that portfolio is.

Pharmaceutical companies therefore face perhaps the purest incentive in the economy to combine frontier-class intelligence with proprietary experimental data while preventing premature disclosure of either the data or the resulting hypotheses—and the resolution, once open-weight scientific models made it feasible, is the founding maneuver of Compute Repatriation: bring the model closer to the secret rather than continuously sending the secret toward the model. A fine-tuned open model running inside the research perimeter can be shown everything, including the failures and dead ends that constitute much of a laboratory’s real institutional knowledge, without any disclosure occurring at all. The frontier API remains available for what it is uniquely good at—general reasoning over public literature—while the discovery itself gestates on machines the company controls, until the patent filing or the publication converts secrecy from an asset into a formality.


3.6 Startups: The Smallest Companies May Possess the Most Concentrated Secrets

Large corporations are not the only candidates for repatriation, and the startup case is worth taking seriously precisely because it seems, at first hearing, absurd. A startup may own almost no physical assets while possessing one highly valuable idea, dataset, algorithm, customer relationship, or technical discovery; its intellectual-property concentration is therefore extreme—the entire enterprise value may reside in information that fits in a single prompt. In the early years of generative AI, startups accepted external APIs without hesitation because building private infrastructure was unrealistic at seed-stage budgets, and because speed mattered more than anything the APIs could conceivably leak. Both conditions are now weakening. Powerful small open models, workstation-class accelerators, private inference appliances, and lower-cost enterprise GPU servers have collapsed the minimum viable scale of private AI operation, and tools that let a two-person team run competent local models on a single desktop have moved from hobbyist novelty to engineering practice [15].

The result could be a surprisingly democratized version of Compute Repatriation, in which capabilities that in 2024 belonged only to governments and the Fortune 500 diffuse to any company that decides its secret is worth a server. A startup’s rational architecture comes to mirror a bank’s in structure while differing by orders of magnitude in scale: frontier APIs for general product features, an enterprise cloud tier for customer-facing inference, and a locally operated open model for the core algorithm, the training data, and the strategy documents—the ten percent of information that is the company. That such an architecture is now affordable at seed stage is itself evidence for this paper’s larger claim that placement-awareness, not cloud-first, is becoming the default mental model of enterprise computing at every scale.


Section 4: What Compute Repatriation Means for AWS, Azure, Google Cloud, and the Hyperscalers

A thesis about computing moving inward must confront, honestly and at length, the most spectacular fact of the 2026 technology economy: computing investment moving outward at a scale never before witnessed. The four largest hyperscalers reported second-quarter 2026 earnings that lifted their combined projected capital expenditure for the year toward $760 billion, up from roughly $413 billion in 2025 [20]; Nvidia reported quarterly revenue of $96.2 billion, up 106 percent year over year, with $89.0 billion from datacenters alone, and guided the following quarter to $108 billion [16]. Any account of Compute Repatriation that cannot coexist with these numbers is wrong. This section argues that the two phenomena not only coexist but require each other, and that the hyperscalers’ own conduct through 2026 shows they understand this better than their skeptics do.


4.1 Repatriation Does Not Mean De-Clouding

The easiest mistake would be to interpret this paper as forecasting the decline of cloud computing, and that interpretation should be closed off decisively. The world’s demand for AI compute is expanding so rapidly that public and private infrastructure can grow simultaneously for years without competing for the same workloads. Training frontier models will remain intensely centralized, for the capital and engineering reasons detailed in Section 1. Temporary and unpredictable inference surges will always favor elastic cloud capacity, because no enterprise rationally buys hardware for its worst-case hour. Smaller companies will continue to prefer fully managed AI for the same reasons they preferred managed everything. Global consumer applications require geographically distributed cloud presence that no single enterprise could replicate. And an enormous share of enterprise AI consumption will arrive embedded invisibly inside SaaS products whose infrastructure decisions belong to the vendor, not the customer. The cloud’s own leadership is, accordingly, unembarrassed by the repatriation conversation: Amazon’s chief executive told shareholders in 2026 that AWS could plausibly become a trillion-dollar-revenue business over time [23], and his spending posture matched the words:

“We’re not going to be conservative in how we play this – we’re investing to be the meaningful leader.”

— Andy Jassy, Chief Executive Officer, Amazon, 2026 Letter to Shareholders [20]

The more plausible future is therefore hybridization—the same conclusion the pre-AI repatriation literature reached, now amplified. Flexera places 73 percent of organizations on hybrid estates already [25]; the Barclays survey’s 86 percent of CIOs planning partial repatriation coexists with Gartner’s forecast of public cloud spending growing 21 percent past $723 billion [24], and the reconciliation of those figures is that both describe the same behavior: selective, workload-level placement, executed continuously, in both directions.


4.2 The Cloud Becomes the Overflow Market: Baseload and Peaking Compute

One useful way to picture the 2030 architecture borrows from a century of electricity economics. Power systems distinguish baseload generation—steady, predictable demand served by owned plants running near constant utilization—from peaking capacity, purchased from the market only during surges, at prices that would be ruinous if paid around the clock. The distinction exists because ownership wins wherever utilization is high and predictable, while markets win wherever demand is spiky and uncertain, and no sane utility chooses only one. Enterprise AI consumption is acquiring exactly this shape. A corporation’s sensitive, high-volume, forecastable inference—the document pipelines, the compliance surveillance, the customer-service agents, the background monitoring workloads that Gartner identifies as the fastest-growing inference category [39]—is baseload, and it migrates naturally toward owned or dedicated capacity where marginal cost is lowest and governance is total. The variable remainder—seasonal peaks, experimental workloads, frontier-model reasoning, burst training—flows to hyperscale clouds, which function as the peaking market. Private GPUs become the enterprise’s baseload compute; cloud GPUs provide elastic overflow. Far from injuring the cloud, this division can improve the economics of both sides at once: the enterprise raises the utilization of what it owns, and the hyperscaler sells its scarce capacity into the demand segments—elasticity and frontier capability—where its advantage is genuinely unassailable.


4.3 Hyperscalers Can Sell the Repatriation They Appear to Be Competing Against

AWS, Microsoft, and Google are not passive victims of this transition, and their 2025–2026 product catalogs read as evidence that they intend to own it. The list of what a hyperscaler can profitably sell to a repatriating customer is long: dedicated and physically isolated infrastructure; private networking; managed Kubernetes for on-premises clusters; model catalogs and one-click deployment of open weights; security and identity systems; colocation partnerships; sovereign-cloud offerings; private AI appliances; enterprise agent platforms; hybrid orchestration control planes; model monitoring; and burst capacity contracts for exactly the overflow architecture described above. AWS emphasizes private model customization and network isolation within Bedrock and has now built an entire physically separate European Sovereign Cloud, representing more than €7.8 billion of committed investment, to serve customers whose sovereignty requirements its standard regions could not satisfy [42]. Microsoft operates managed-model environments in which customer prompts are segregated from external model providers [10] and has extended a Sovereign Public Cloud across every European region alongside national partner clouds [44]. Google offers enterprise data controls and zero-retention configurations for its generative services [12]. The pattern in these offerings is unmistakable: the hyperscalers are unbundling control from location, selling ever-stronger isolation guarantees at ever-finer granularity, and thereby evolving from places where AI runs into control planes that manage AI running everywhere—including inside buildings they do not own.


4.4 Nvidia Could Benefit From Both Outcomes

Compute Repatriation is particularly interesting for Nvidia because the company occupies the one position in the AI economy that is indifferent to where the computing happens. If enterprises rely entirely on hyperscalers, Nvidia sells accelerators into hyperscale datacenters, as its record $89.0 billion of quarterly datacenter revenue attests [16]. If enterprises begin operating private AI infrastructure, Nvidia sells the same accelerators—packaged as enterprise AI factories with networking, software, and inference microservices—into corporate datacenters and colocation cages instead, which is precisely the configuration Latham & Watkins bought [5]. Nvidia’s own commentary in 2026 makes clear that management regards the broadening of its customer base beyond a handful of hyperscalers as a strategic objective: the company described accelerating demand across AI clouds, industrial users, and enterprises alongside the hyperscalers [16], and its chief executive framed the quarter’s significance in language that doubles as a mission statement for the inference economy:

“AI has reached its inflection point. … Now, compute is revenue.”

— Jensen Huang, Q2 Fiscal 2027 Earnings Announcement, August 26, 2026 [16]

A partial migration away from cloud-only AI therefore does not represent a migration away from Nvidia; it may broaden Nvidia’s addressable market from several gigantic buyers into thousands of corporate AI operators, each individually small but collectively enormous, and each stickier than a hyperscaler that designs its own silicon on the side. The strategic risk the buildout does concentrate—memory scarcity driven by the AI expansion itself, which Nvidia warned will compress its margins into early fiscal 2027 [17]—binds cloud and enterprise buyers alike, and underscores that the constraint on the entire five-layer economy is shifting downward, from model capability toward physical supply.


4.5 Frontier Laboratories Face a Different Strategic Problem

OpenAI, Anthropic, and the other frontier laboratories face a more nuanced challenge than the hyperscalers, because their product—maximum capability—sits at exactly the point in the routing hierarchy that repatriation squeezes. Their most powerful models will likely remain superior to anything corporations can operate privately; the Stanford AI Index shows the closed-versus-open gap persisting even as it narrows [32]. But enterprises are learning to reserve those models for the hardest tasks while routing repetitive, predictable, sensitive, or high-volume inference toward cheaper or private alternatives, and the open-model ecosystem’s rising share of routed queries [45] shows the squeeze operating in real time. A future corporate AI router might make decisions of the following kind, continuously and automatically:


Task CategoryRouting DestinationDeciding Factor
Routine summarization and draftingInternal open-weight modelCost at volume; no external disclosure
Confidential contract analysisPrivate fine-tuned model (Zone 3)Privilege and trade-secret protection
Extremely difficult scientific or strategic reasoningFrontier model via API (Zone 1)Maximum capability worth the premium
Customer-facing low-latency querySpecialized local or edge model (Zone 4)Response-time budget
Sensitive board materialIsolated enterprise model, restricted enclaveArchitectural isolation, not contract
Nonconfidential research and searchCheapest acceptable providerPure price per token

This routing table changes the competitive question of the model market from “which model does the company use?” to “which model receives each individual task?”—a shift from winner-take-all platform competition toward continuous, per-transaction auction, mediated by routing software that becomes one of the most consequential enterprise AI markets of the late 2020s. For the laboratories, the strategic responses are already visible: sell capability where capability is the binding constraint; sell enterprise privacy tiers that reduce the security motive for routing away [8][9]; license or release open weights to participate in the private tier rather than ceding it; and move up the stack into agents and applications where integration, not raw model quality, is the moat. The laboratories that treat the router as their true customer will fare better than those that assume the enterprise buys models the way it once bought databases—exclusively and forever.


Section 5: The 2030 Enterprise Architecture—Public Cloud + Private Compute + Open Models + Proprietary Data

The preceding sections argued from forces and cases; this section assembles the destination. What does the enterprise intelligence architecture of 2030 actually look like, if the four engines of Section 2 keep running and the industry patterns of Section 3 propagate? The answer developed here has five components: the end of the single-model enterprise; a four-zone operational architecture; the ascendancy of proprietary data as models commoditize; the decisive influence of autonomous agents; and the return of the corporate datacenter in a form its 2010 predecessor would not recognize—followed by the policy problem this architecture creates for governments.


5.1 The End of the Single-Model Enterprise

By 2030, sophisticated corporations are unlikely to run on one AI model; they may operate dozens, and the portfolio will be as deliberately diversified as an investment book. Some will be frontier models supplied through APIs and reserved for frontier problems. Others will be open-weight models running privately, fine-tuned into specialists for the company’s documents, products, and vocabulary. Some will be narrow experts—for code, science, finance, speech, vision, cybersecurity, or robotics—chosen because a small specialist beats a large generalist on bounded tasks at a fraction of the cost. Smaller models still will run directly on employee devices and embedded systems. The Stanford AI Index’s finding that 88 percent of surveyed organizations already use AI in at least one function while agent deployment remains in the single digits [31] describes an adoption wave that has reached breadth but not yet depth; as depth arrives, heterogeneity arrives with it, because no single model optimizes simultaneously for capability, cost, latency, and confidentiality. AI architecture therefore becomes inherently heterogeneous, and “which model do you use?” becomes as naive a question to ask an enterprise as “which computer do you use?”


5.2 A Four-Zone Corporate Intelligence Architecture

A useful way to conceptualize the emerging enterprise is through four operational zones, distinguished not by technology but by who holds control and how much of it, with information classes mapped onto zones by policy and the router of Section 4.5 enforcing the map in real time.


ZoneNameDescriptionTypical Workloads
Zone OnePublic Frontier IntelligenceThe most capable commercial frontier models, accessed through APIs, where maximum reasoning capability matters more than computational sovereignty.Hardest reasoning tasks; public-data research; frontier coding and analysis
Zone TwoPrivate Hosted IntelligenceDedicated or logically isolated infrastructure operated by a hyperscaler, colocation provider, sovereign cloud, or specialized AI cloud, governed under strict enterprise controls.Regulated workloads; moderately sensitive inference; burst capacity; sovereign-residency workloads
Zone ThreeEnterprise-Owned IntelligenceAccelerators, open-weight models, proprietary retrieval systems, internal agents, and sensitive datasets owned or directly controlled by the corporation.Trade secrets; privileged material; core proprietary data; baseload high-volume inference
Zone FourEdge IntelligenceSmaller models operating at factories, laboratories, stores, hospitals, vehicles, robots, and employee devices, where latency, connectivity, or privacy makes centralized inference undesirable.Real-time control; on-device assistance; disconnected operation; machine-speed loops

Two properties of this architecture deserve emphasis. First, the zones are permeable by policy rather than sealed by technology: a task’s classification, not its type, determines its zone, and the same document-summarization request lands in Zone One or Zone Three depending entirely on whose document it is. Second, the architecture is dynamic: the enterprise continuously re-routes intelligence across zones as prices move, models improve, regulations tighten, and threats evolve, which is why the routing layer—and the telemetry, evaluation, and cost-attribution systems that feed it—constitutes the real crown jewel of the 2030 stack. The zones are where intelligence lives; the router is how the corporation thinks about its thinking.


5.3 Proprietary Data Becomes More Important as Models Commoditize

If model performance converges—and the AI Index’s measurements of narrowing gaps between leading closed, open, American, and Chinese models suggest convergence is the trend at the top of the market [32]—then the durable strategic asset shifts away from the model itself toward what the model is given to reason over. Every major corporation possesses information unavailable on the public internet and therefore absent from every foundation model’s training corpus: decades of customer interactions; industrial telemetry; scientific experiments, including the unpublished failures; financial histories; internal engineering documents; legal archives; supply-chain records; employee expertise; failure reports; and operational procedures. This corpus is the raw material from which differentiated corporate intelligence is constructed, and it is the one input to the AI value chain that competitors cannot buy, benchmark, or replicate at any price. Consequently, Compute Repatriation is not merely a defensive doctrine about protecting data; it is an offensive doctrine about turning proprietary data into a productive AI asset without surrendering control over it—fine-tuning on it, retrieving over it, and letting agents act upon it, all within a perimeter that keeps the resulting intelligence as proprietary as the data that produced it. In the commoditized-model world, the corporations that win are not those with access to the best model, since everyone has that; they are those whose private data, private fine-tunes, and private agents compound advantages no one outside the perimeter can observe, which is also why the economics literature’s emphasis on organizational complements to technology—the restructuring, training, and process redesign that Brynjolfsson’s productivity J-curve describes [33]—applies with full force here: the data asset pays only for organizations rebuilt to exploit it.


5.4 Agents Make Repatriation More Important Than Chatbots Ever Did

The chatbot era involved humans consciously deciding what to type, which meant the disclosure surface of enterprise AI was bounded by human judgment, one prompt at a time. The agentic era is categorically different. Agents autonomously access email, databases, payment systems, internal messaging, contracts, enterprise-resource-planning systems, customer records, development environments, and any other software their permissions reach; a chatbot sees the information a human chooses to provide, while an agent can potentially see everything its credentials allow it to retrieve. That distinction dramatically expands the security boundary at precisely the moment the volume expands with it—agentic workflows consuming on the order of a thousand times the tokens of chat interactions [40], token demand projected to grow twenty-four-fold by 2030 [38], and background agents running continuously against every event in the enterprise [39]. Erik Brynjolfsson has projected that within a generation most workers will command fleets of AI agents larger than today’s biggest corporate workforces, performing design, coding, negotiation, and experimentation continuously [34]. An agent is, in effect, a tireless employee with a photographic memory and root access, and the question of which infrastructure that employee works on—who hosts its reasoning, who logs its actions, who could subpoena its memory, who patches the framework that mediates its permissions—becomes among the most consequential security questions the corporation faces. Prompt-injection attacks, in which hostile content hijacks an agent’s instructions, make the point vividly: the blast radius of a compromised chatbot is one bad answer, while the blast radius of a compromised agent is every system the agent can touch. As AI shifts from answering questions to taking actions, corporations will care intensely about where their agents run, which models control them, where their memory resides, and what infrastructure mediates their permissions—and each of those concerns is a placement decision, which is to say each of them is Compute Repatriation stated in the vocabulary of the agentic economy. This is the deepest reason the thesis strengthens with time: repatriation was optional for chatbots and becomes existential for agents.


5.5 The Corporate Datacenter Returns—But It Is Not the Old Datacenter

The corporate datacenter of 2030 will not reproduce the corporate datacenter of 2010, and mistaking the former for the latter is the surest way to misjudge this entire transition. The earlier facility primarily stored applications and databases; it was a warehouse for software, judged on uptime and cost per rack. The new facility is an intelligence factory, judged on the value of the reasoning it produces, and its bill of materials reads accordingly: GPUs and specialized accelerators; high-bandwidth, low-latency networking; vector databases and AI-optimized storage; model registries and evaluation harnesses; agent orchestration frameworks; enterprise knowledge graphs; security guardrails and inference gateways; identity systems extended to non-human actors; and the proprietary models themselves, which are assets in the balance-sheet sense as much as the technical one. Its power density, cooling design, and physical security requirements resemble a hyperscale pod far more than a 2010 server room, which is why most organizations will not construct these facilities themselves: they will lease locked, dedicated space in colocation campuses—exactly Latham’s arrangement [7]—or procure managed private infrastructure from the hyperscalers themselves, per Section 4.3. The physical building matters far less than the governance boundary. What returns to the enterprise is not necessarily the real estate; it is control over the intelligence stack that the real estate houses, and that distinction is why this paper speaks of repatriating compute rather than repatriating datacenters.


5.6 Public Policy Must Distinguish Cloud Concentration From Enterprise AI Concentration

Compute Repatriation finally presents governments with a genuinely new regulatory geometry, and it arrives just as policymakers were settling into the old one. For several years the dominant policy anxiety has been concentration: the fear that a handful of hyperscalers and frontier laboratories could dominate AI infrastructure, with all the leverage over prices, access, and political power that domination implies. Private enterprise AI partially answers that anxiety by decentralizing the architecture—thousands of firms operating their own models is, structurally, an antitrust regulator’s dream relative to five firms operating everyone’s. Yet the same development creates the mirror-image problem: thousands of private AI installations are thousands of security perimeters, incident-response teams, and model-governance regimes of wildly varying quality, largely invisible to the oversight mechanisms—API monitoring, provider reporting, cloud choke points—that regulators were counting on. Regulators will accordingly need standards governing private-model cybersecurity; model provenance and supply-chain integrity; agent permissions and audit trails; AI in critical infrastructure; logging and incident disclosure; data and inference residency; high-risk autonomous systems; accelerator export controls that now reach corporate rather than only national buyers; and the private deployment of highly capable open-weight models, which is the point where the open-model debate documented by Stanford HAI’s researchers [28] collides with enterprise practice. The tension can be stated as a single trade-off that will organize AI-infrastructure policy from 2027 to 2030: centralized AI is easier to observe but concentrates power, while distributed private AI reduces concentration but makes oversight harder. Good policy will refuse to pick a side and will instead price the externalities of each—demanding transparency from the centralized tier and baseline security from the distributed one—because the architecture that is coming, on all the evidence assembled here, is irreversibly both.


Section 6: What Have We Learned? Seven Pillars

A paper of this length owes its reader a distillation, and this section provides one in the form of seven pillars—seven conclusions that survive contact with the evidence of 2020 through 2026 and that, together, constitute the operating doctrine of Compute Repatriation. The original architecture of this argument contained five pillars; the research assembled here compels two additions, one economic and one epistemic, which appear as Pillars Six and Seven.


Pillar 1 — The AI Era Will Not Simply Extend the Cloud Era

The first lesson is that technological transitions rarely move permanently in one direction, however inevitable the current direction feels while it lasts. Cloud computing shifted enterprise infrastructure outward because general-purpose computing benefited enormously from scale, specialization, and elasticity, and for fifteen years every workload that mattered confirmed the logic. AI introduces different economics: extremely large training workloads favor centralization even more intensely than anything the cloud era produced, while sensitive, repetitive, latency-bound, and high-volume inference can favor localization for reasons of control, cost, and physics that no provider contract can repeal. The result is not the reversal of cloud computing—the $760 billion capital-expenditure wave of 2026 [20] and the 21 percent growth of public cloud spending [24] refute any reversal thesis on contact—but a restructuring of it, in which the same enterprise consumes hyperscale cloud and operates private intelligence simultaneously, assigning each workload to the venue where its particular combination of sensitivity, volume, and urgency is best served. The future enterprise will rent and own at the same time, and will regard the question “cloud or not?” as a category error, the way a treasurer regards “debt or equity?”—not as a doctrine to be settled but as a mix to be managed.


Pillar 2 — Confidentiality Is Becoming an Architectural Decision

The second lesson is that AI governance cannot be reduced to privacy policies, however well drafted. Contracts and provider safeguards matter enormously; the enterprise commitments published by OpenAI, Anthropic, Microsoft, AWS, and Google [8][9][10][11][12] represent real protection and real progress, and for most workloads they are the right answer. But organizations holding extraordinary concentrations of intellectual property are concluding that contractual privacy is not equivalent to architectural isolation, because a promise and a perimeter are different kinds of things, enforced by different mechanisms, and failing in different ways. The closer AI moves toward a corporation’s most valuable knowledge, the more likely management is to ask whether the model should move toward the data rather than the data continually moving toward the model—the founding maneuver of every case examined in Section 3. The scientist’s unpublished result, the startup’s business plan, the lawyer’s privileged memorandum, and the prospective ICANN applicant’s confidential list of strings [1] all illustrate the same principle, which deserves to be engraved above the door of every AI governance committee: information can possess maximum value precisely before anybody else knows it, and the architecture of that information’s computation is therefore part of its valuation.


Pillar 3 — Inference Economics Will Decide How Much Compute Comes Home

The third lesson is economic, and it disciplines the first two. Private AI infrastructure will not proliferate merely because executives like the idea of owning servers, and any repatriation program justified on sentiment will fail on invoice. It will proliferate where cost, utilization, security, performance, or regulatory requirements make it rational, and the evidence of 2026 shows that threshold being crossed at scale: inference spending surpassing training spending for the first time [39]; agentic workloads consuming tokens at a thousand times chat-era rates [40]; total token demand projected to grow twenty-four-fold by 2030 [38]; and tiered routing architectures cutting blended costs eight-fold relative to frontier-only consumption [39]. A corporation operating a handful of AI queries has no case for GPUs; a corporation operating billions of inference transactions through employees, customers, agents, robots, factories, and instruments faces a straightforward baseload-versus-peak calculation that increasingly favors owning the baseload. The decisive metric will be cost per useful unit of intelligence—not cost per GPU, not cost per token, but the fully loaded price of a correct answer delivered where and when the business needs it—and enterprises that learn to measure that quantity will make placement decisions with a confidence their competitors cannot match.


Pillar 4 — The Enterprise Becomes an AI Operator, Not Merely an AI Customer

The fourth lesson is organizational, and it is the one that will consume the most management attention per dollar. During the first stage of generative AI, companies purchased access to intelligence; during the next stage, a meaningful subset of them will operate it, and operating intelligence is a discipline with its own required competencies: model engineering; GPU infrastructure and capacity planning; AI networking; data governance at training-set quality; inference optimization; AI-specific security; agent orchestration; and continuous model evaluation. These capabilities do not currently exist inside most corporations, and building them is expensive—Latham’s 900-person technology organization with 100 AI specialists [6], and running costs plausibly in the tens of millions of dollars annually [7], indicate the entry price at the high end of professional services. The MIT NANDA research showing 95 percent of enterprise generative-AI pilots delivering no measurable profit-and-loss impact [28], with the failures concentrated in organizations that treated AI as software to be installed rather than a capability to be operated [29], is best read as a portrait of what under-investment in operatorship looks like: adoption without transformation. Companies whose competitive advantage depends heavily on proprietary knowledge will come to regard AI operatorship as being as fundamental as cybersecurity or financial management, and Latham & Watkins is significant precisely because it demonstrates that this transition can reach industries as far from computing as the profession of law [3].


Pillar 5 — The Winning Architecture Is Hybrid, Not Ideological

The fifth lesson is a warning against the tribalism that technology debates reliably generate. The discussion should not be allowed to harden into cloud versus on-premise, open versus closed models, or rent versus own, because each architecture is genuinely superior at something: frontier clouds provide extraordinary capability; open models provide control and customization; private infrastructure provides predictable economics and architectural confidentiality; public clouds provide elasticity; edge models provide latency; and enterprise data provides the differentiation that makes any of it worth doing. The optimal corporation combines all six, and the durable competitive advantage lies not in any single choice but in the meta-capability of knowing which intelligence belongs where—maintained continuously as prices, capabilities, and regulations move. Hybridity here is not a compromise between pure positions; it is the pure position, the only architecture that respects all four forces of Section 2 at once, and the survey data showing 86 percent of CIOs repatriating something while only 8 percent repatriate everything [24][25] shows the entire market converging on exactly this understanding.


Pillar 6 — The Productivity Payoff Is Real, and It Will Be Won or Lost Inside the Enterprise

The sixth pillar—the first of the two this research compels adding—concerns the macroeconomic stakes, because the question of where intelligence runs matters only if intelligence is actually producing value, and on that question the economics profession spent 2024 through 2026 conducting a vigorous argument in public. Erik Brynjolfsson of Stanford, applying his productivity J-curve framework, estimates that United States productivity growth reached roughly 2.7 percent in 2025—nearly double the prior decade’s 1.4 percent average—as the economy transitioned from the investment phase of AI adoption into the harvest phase, with a small cohort of power users automating end-to-end workstreams delivering outsized gains [33]. He has framed the moment in terms that this paper’s enterprise-level evidence supports:

“We are transitioning from an era of AI experimentation to one of structural utility.”

— Erik Brynjolfsson, Director, Stanford Digital Economy Lab, Financial Times, February 2026 [33]

Nobel laureate Daron Acemoglu of MIT stakes out the skeptical pole, forecasting in his “Simple Macroeconomics of AI” a total factor productivity boost he characterizes as modest but far from trivial—on the order of 0.7 percent, with GDP effects near 1.1 to 1.6 percent over a decade [35][36]—on the ground that AI’s task coverage is narrower and its implementation harder than enthusiasts assume [37]. He has also warned, in his 2026 commentary, that the decisive developments to watch are agentic products and the institutional machinery around them rather than raw model capability [46]. What matters for this paper is not adjudicating between 2.7 percent and 0.7 percent; it is noticing what the two positions share. Both locate the binding constraint inside the firm—in Brynjolfsson’s organizational complements and power users, in Acemoglu’s implementation and task-selection frictions—and the MIT NANDA finding that the successful 5 percent of deployments were deeply integrated, workflow-embedded systems [29] says the same thing at case-study resolution. The productivity dividend of AI, whatever its final size, will be won or lost in exactly the layer that Compute Repatriation addresses: the enterprise’s own architecture of data, models, agents, and control. Nations debate the AI payoff; enterprises determine it.


Pillar 7 — The Most Important AI May Become the Least Visible

The seventh and final pillar is epistemic, and it concerns what the world will be able to see. The public’s picture of artificial intelligence is assembled from what is observable: consumer chatbots, published benchmarks, leaderboards, API traffic, and the announcements of a dozen laboratories. Compute Repatriation systematically moves value away from every one of those observation points. The bank’s alpha-generating models never touch a leaderboard; the pharmaceutical company’s discovery models reason over data no benchmark contains; the law firm’s privileged assistants run in cages only employees can enter [7]; the defense enclave is invisible by statute. Even measurement institutions are already struggling at the visible tier—the Stanford AI Index reports that foundation-model transparency declined as leading laboratories stopped disclosing training details, with its transparency index falling from 58 to 40 even as documented AI incidents rose to 362 [31]. As the private tier grows, an increasing share of the world’s most economically consequential AI will be un-benchmarked, un-announced, and un-observed, which carries three implications worth stating plainly: analysts and investors will systematically mismeasure AI-driven advantage, because the advantage is deliberately hidden; policymakers will regulate the visible tier while the invisible tier compounds, which is precisely the oversight asymmetry Section 5.6 warned about; and the discourse about “what AI can do” will lag what AI is actually doing inside perimeters, by widening margins. The history of general-purpose technologies suggests this is normal—electricity’s greatest effects occurred invisibly inside factory redesigns, not at public exhibitions—but it obliges intellectual humility from everyone, this author included, who writes about AI from the outside. The map of the AI economy is about to become permanently incomplete, and knowing that is itself strategic knowledge.


Conclusion: Why “Compute Repatriation” Fits the Coming Enterprise AI Economy

For more than a decade, corporate computing followed an extraordinarily powerful economic logic: move applications, servers, databases, and infrastructure into increasingly centralized clouds. Enterprises exchanged ownership for elasticity, capital expenditure for operating expenditure, and physical control for the convenience of global hyperscale infrastructure, and the exchange was, for the workloads of that era, overwhelmingly correct. Artificial intelligence does not eliminate those advantages. It introduces new considerations that sit beside them, and occasionally above them.

When computing simply stores a document, the infrastructure provider does not need to understand what the document means, and the customer does not need to care whether it could. When an AI model reasons over that document, combines it with proprietary databases, interprets a confidential strategy, proposes a scientific hypothesis, negotiates with another agent, or takes actions inside corporate software, computing becomes inseparable from institutional intelligence, and the question of whose computer is doing it acquires a weight it never had before. That changes the strategic calculation in ways this paper has traced through four forces, six industries, and the entire five-layer economy. A scientist asking AI to evaluate an unpublished discovery is not merely consuming computing. A pharmaceutical company applying a model to proprietary molecular data is not merely renting server capacity. A law firm placing millions of privileged documents inside an AI retrieval system is not merely purchasing software. A startup allowing autonomous agents to analyze its source code, customer pipeline, product roadmap, and financing plan is not merely subscribing to another SaaS platform. Each is deciding where its institutional intelligence is allowed to exist.

That is why the Latham & Watkins example deserves more attention than its immediate impact on the legal profession might suggest. A major law firm purchasing Nvidia infrastructure and fine-tuning open-weight models is evidence that the boundary between an AI company and an ordinary enterprise is beginning to blur [3], and Nvidia’s own enterprise strategy—AI factories designed for corporate and sovereign operators, promoted by a chief executive who tells governments that intelligence is infrastructure to be produced rather than imported [18]—reinforces the direction from the supply side. The hyperscalers, far from resisting, are racing to sell control itself: sovereign clouds, isolated capacity, hybrid control planes, and the whole apparatus by which repatriation will, for most enterprises, actually be delivered [42][44].

The crucial word in Compute Repatriation is therefore not compute. It is repatriation, and repatriation does not require the computer to return physically to corporate headquarters. It means that decision rights return. Security boundaries return. Economic optionality returns. Model choice returns. Sensitive inference returns. Proprietary knowledge remains under a level of corporate control that management, not a vendor, considers appropriate. Sometimes the infrastructure will sit in a corporate building; sometimes in a locked cage of a colocation datacenter, as Latham’s does [7]; sometimes in a dedicated environment inside AWS, Azure, Google Cloud, or Oracle; sometimes on an employee workstation; sometimes at the edge inside a factory, robot, laboratory, vehicle, or defense system. And sometimes the organization will continue sending a task to the world’s most powerful frontier model, because that model remains the best tool available and the task carries nothing that needs protecting. Every one of those placements is a choice, and the existence of the choice is the entire point.

Compute Repatriation should therefore not be mistaken for a prediction that corporations abandon the cloud; the evidence assembled here—$760 billion of hyperscaler capital expenditure [20], public cloud growth of 21 percent [24], Nvidia’s doubling revenue [16]—forecloses that reading. It predicts something more subtle and potentially more important: corporations will stop treating the cloud as the automatic destination for every form of intelligence. The AI era may consequently produce one of the great technological ironies of the late 2020s. After spending more than a decade convincing companies that owning computing infrastructure was unnecessary, the technology industry may now sell those same companies powerful GPU servers, private AI factories, open-weight models, dedicated inference clusters, sovereign clouds, and enterprise agents—because artificial intelligence has made direct computational control strategically valuable again. Cloud computing separated companies from their machines. Artificial intelligence may reconnect some of them.

By 2030, the sophisticated corporation will operate simultaneously across hyperscale cloud, private compute, dedicated accelerators, open models, frontier APIs, proprietary datasets, and edge intelligence. It will be neither completely centralized nor completely decentralized. It will instead continuously decide which workloads deserve elasticity, which demand capability, which require low cost, which require low latency, and which are simply too important to leave its own governance boundary. That is the central argument of this paper and the reason the title fits: Compute Repatriation is the selective return of artificial-intelligence infrastructure, models, inference, and decision-making authority to the enterprise, after a decade in which corporate computing moved steadily outward into the cloud. And the question that follows may define the enterprise architecture of the next decade: will the AI era partially reverse the cloud era? The emerging evidence—from a law firm’s locked server cage to the sovereign clouds of Europe to the routing tables of the world’s banks—suggests that, for the most valuable forms of corporate intelligence, it already has begun.


Footnotes / Endnotes:

[1] ICANN. ICANN 2026 Round Closes with More Than 1,600 New gTLD Applications (August 13, 2026). https://www.icann.org/en/announcements/details/icann-2026-round-closes-with-more-than-1600-new-gtld-applications-13-08-2026-en

[2] Kevin Murphy, Domain Incite. New gTLDs: What We Do and Don’t Know About the 2026 Round (August 2026). https://domainincite.com/31864-new-gtlds-what-we-do-and-dont-know-about-the-2026-round

[3] Legal IT Insider, reporting on the Financial Times. Latham Builds Its Own AI Models with Nvidia GPU Server Investment (September 12, 2026). https://legaltechnology.com/latham-builds-its-own-ai-models-with-nvidia-gpu-server-investment/

[4] Legal Cheek. Latham Buys Its Own AI Servers in BigLaw First (September 2026). https://www.legalcheek.com/2026/09/latham-buys-its-own-ai-servers-in-biglaw-first/

[5] Pulse 2.0. Latham & Watkins Buys Nvidia GPU Servers To Build In-House AI Systems (September 2026). https://pulse2.com/latham-watkins-buys-nvidia-gpu-servers-to-build-in-house-ai-systems/

[6] Complete AI Training. Latham & Watkins Buys Nvidia Chips to Fine-Tune AI Models on Its Own Servers (September 2026). https://completeaitraining.com/news/latham-watkins-buys-nvidia-chips-to-fine-tune-ai-models-on/

[7] InView Independent News. Latham & Watkins Buys Nvidia Servers to Build Its Own AI (September 2026). https://inview.info/news/195278-latham__watkinc_buys_nvidia_servers_to_build_its_own_ai

[8] OpenAI. Enterprise Privacy at OpenAI. https://openai.com/enterprise-privacy/

[9] Anthropic. Anthropic Privacy Center: Commercial Products and Training Defaults. https://privacy.anthropic.com/

[10] Microsoft Learn. Data, Privacy, and Security for Azure OpenAI Service. https://learn.microsoft.com/en-us/legal/cognitive-services/openai/data-privacy

[11] Amazon Web Services. Amazon Bedrock — Security, Privacy, and FAQs. https://aws.amazon.com/bedrock/faqs/

[12] Google Cloud. Generative AI on Vertex AI — Customer Data Governance. https://cloud.google.com/vertex-ai/generative-ai/docs/data-governance

[13] Vectra AI (citing UpGuard, IBM). Shadow AI Explained: Risks, Costs, and Enterprise Governance (2026). https://www.vectra.ai/topics/shadow-ai

[14] JustDoers (citing Cyberhaven 2025 AI Adoption and Risk Report). Shadow AI in the Enterprise (2026). https://www.justdoers.com/blog/shadow-ai-in-the-enterprise-how-employees-are-secretly-using-unapproved-ai-tools-and-the-security-nightmares-it-creates

[15] FPB Legal (citing IBM Cost of a Data Breach 2025). Shadow AI and Company Data: Risks and Compliance (2026). https://www.fpblegal.com/en/shadow-ai-company-data/

[16] NVIDIA Newsroom. NVIDIA Announces Financial Results for Second Quarter Fiscal 2027 (August 26, 2026). https://nvidianews.nvidia.com/news/nvidia-announces-financial-results-for-second-quarter-fiscal-2027

[17] CNBC. Nvidia Earnings Takeaways: Huang Forecasts 70% Fiscal 2028 Revenue Growth (August 26, 2026). https://www.cnbc.com/2026/08/26/nvidia-nvda-earnings-report-q2-2027-live-updates.html

[18] NVIDIA Blog. “Largest Infrastructure Buildout in Human History”: Jensen Huang on AI’s Five-Layer Cake at Davos (January 2026). https://blogs.nvidia.com/blog/davos-wef-blackrock-ceo-larry-fink-jensen-huang/

[19] Time News. Nvidia CEO Jensen Huang Tells G20 Nations AI Is Essential Infrastructure (September 2026). https://time.news/nvidia-ceo-jensen-huang-tells-g20-nations-ai-is-essential-infrastructure/

[20] Statista (Q2 2026 earnings analysis). Big Tech’s AI Spending to Reach $760 Billion in 2026 (July 31, 2026). https://www.statista.com/chart/35046/capital-expenditure-of-meta-alphabet-amazon-and-microsoft/

[21] Nick Patience, The Futurum Group. AI Capex 2026: The $690B Infrastructure Sprint (February 12, 2026). https://futurumgroup.com/insights/ai-capex-2026-the-690b-infrastructure-sprint/

[22] CNBC. Tech AI Spending Approaches $700 Billion in 2026, Cash Taking Big Hit (February 6, 2026). https://www.cnbc.com/2026/02/06/google-microsoft-meta-amazon-ai-cash.html

[23] UncoverAlpha. Amazon, Google, Microsoft, Meta Q2 2026 Earnings Analysis (August 2026). https://www.uncoveralpha.com/p/amazon-google-microsoft-meta-q2-earnings

[24] DataBank (citing Barclays CIO Survey; Gartner). Why 86% of CIOs Are Rethinking Their Cloud Strategy (November 2025). https://www.databank.com/resources/blogs/why-86-of-cios-are-rethinking-their-cloud-strategy/

[25] LOGIX (citing Barclays, IDC, Flexera 2026). Cloud Repatriation and Colocation: Why Workloads Are Moving Back (July 2026). https://logix.com/blog/cloud-repatriation-colocation/

[26] HyScaler (citing Flexera 2025 State of the Cloud). Cloud Repatriation in 2026: Why Enterprises Are Moving Back from the Cloud (May 2026). https://hyscaler.com/insights/cloud-repatriation-the-strategic-shift-in-it/

[27] Sarah Wang and Martin Casado, Andreessen Horowitz. The Cost of Cloud, a Trillion Dollar Paradox (2021). https://a16z.com/the-cost-of-cloud-a-trillion-dollar-paradox/

[28] MIT Project NANDA, reported by Fortune. The GenAI Divide: State of AI in Business 2025 — 95% of Generative AI Pilots Are Failing (August 2025). https://finance.yahoo.com/news/mit-report-95-generative-ai-105412686.html

[29] Virtualization Review (on MIT Project NANDA). MIT Report Finds Most AI Business Investments Fail, Reveals “GenAI Divide” (August 2025). https://virtualizationreview.com/articles/2025/08/19/mit-report-finds-most-ai-business-investments-fail-reveals-genai-divide.aspx

[30] Stanford Institute for Human-Centered AI. The 2026 AI Index Report. https://hai.stanford.edu/ai-index/2026-ai-index-report

[31] Stanford HAI. Inside the AI Index: 12 Takeaways from the 2026 Report. https://hai.stanford.edu/news/inside-the-ai-index-12-takeaways-from-the-2026-report

[32] United Nations University Campus Computing Centre. What the 2026 Stanford AI Index Report Tells Us About the State of AI (July 2026). https://c3.unu.edu/blog/2026-stanford-ai-index-report-takeaways

[33] Erik Brynjolfsson (Stanford), via Fortune / Financial Times op-ed. The AI Productivity Take-Off Is Finally Visible — Stanford’s Brynjolfsson on the Harvest Phase (February 2026). https://finance.yahoo.com/news/one-stanford-original-ai-gurus-205316027.html

[34] Erik Brynjolfsson, TIME. AI Changed Work Forever in 2025 (January 2026). https://time.com/7342494/ai-changed-work-forever/

[35] MIT Technology Review. A Nobel Laureate on the Economics of Artificial Intelligence — Daron Acemoglu (February 2025). https://www.technologyreview.com/2025/02/25/1111207/a-nobel-laureate-on-the-economics-of-artificial-intelligence/

[36] Daron Acemoglu, National Bureau of Economic Research. The Simple Macroeconomics of AI, NBER Working Paper 32487 (2024); Economic Policy 40(121) (2025). https://www.nber.org/papers/w32487

[37] American Enterprise Institute. There’s Nothing Simple About the Macroeconomics of AI (2024). https://www.aei.org/economics/theres-nothing-simple-about-the-macroeconomics-of-ai/

[38] Goldman Sachs Research (Jim Schneider). AI Agents Forecast to Boost Tech Cash Flow as Usage Soars — 24x Token Growth by 2030 (May 2026). https://www.goldmansachs.com/insights/articles/ai-agents-forecast-to-boost-tech-cash-flow-as-usage-soars

[39] Tech Times (citing Gartner; EY). Gartner Marks First Year Inference Spending Beats AI Training (August 2026). https://www.techtimes.com/articles/323879/20260811/gartner-marks-first-year-inference-spending-beats-ai-training-55-cents-every-cloud-dollar.htm

[40] The Modern Data Company (citing Microsoft/Stanford Digital Economy Lab; Satya Nadella). Why Cheaper AI Tokens Are Increasing Enterprise AI Costs (July 2026). https://www.themoderndatacompany.com/blog/why-cheaper-ai-tokens-are-increasing-enterprise-ai-costs

[41] Suplari (citing Gartner; Morgan Stanley; Menlo Ventures). What Does Enterprise AI Actually Cost? A Finance Leader’s Breakdown (June 2026). https://suplari.com/blog/what-does-enterprise-ai-actually-cost

[42] Amazon Press Center. AWS Launches AWS European Sovereign Cloud and Announces Expansion Across Europe (January 2026). https://press.aboutamazon.com/aws/2026/1/aws-launches-aws-european-sovereign-cloud-and-announces-expansion-across-europe

[43] InfoQ. AWS Launches European Sovereign Cloud amid Questions about U.S. Legal Jurisdiction (January 2026). https://www.infoq.com/news/2026/01/aws-european-sovereign-cloud/

[44] Molderez Consult (citing Microsoft; Gartner; Synergy Research). Sovereign Cloud and AI: Where to Host Your Data and Models in 2026 (July 2026). https://molderez-consult.be/blog/en/cloud-souverain-ia-europe.html

[45] Dealroom News. How Big Is the Open-Model Threat to AI Hyperscalers? (September 2026). https://dealroom.co/news/other-1cabevz-how-big-is-the-open-model-threat-to-ai-hyperscalers/

[46] Metaintro (on Daron Acemoglu, MIT Technology Review interview). Nobel Economist Names Three AI Shifts to Watch in 2026 (May 2026). https://www.metaintro.com/blog/nobel-economist-three-ai-things-watch[47] TweakTown (BG2 interview). NVIDIA CEO on Sovereign AI for Countries (September 2025). https://www.tweaktown.com/news/107910/nvidia-ceo-on-sovereign-ai-for-countries-no-one-needs-atomic-bombs-everyone-needs-ai/index.html