Introduction: When a Chip Company Bought a Way to Understand the World
On September 28, 2026, Advanced Micro Devices announced that it had entered into a definitive agreement to acquire World Labs, the spatial-intelligence company co-founded by Stanford professor and computer-vision pioneer Fei-Fei Li, in an all-stock transaction valued at approximately $8.2 billion and expected to close by the end of 2026, subject to regulatory approval.[1][2] The transaction is the second largest in AMD’s history, and the company said that following the closing Li would join as executive vice president and chief scientist, reporting directly to chair and chief executive Lisa Su.[3][4] At first glance, the deal could be interpreted in familiar semiconductor terms: a chip company expanding its artificial-intelligence portfolio as it competes with Nvidia across GPUs, rack-scale systems, software, inference and the rapidly expanding market for what the industry has begun to call physical AI. But the more consequential part of the announcement was not simply that another chip company was moving higher in the artificial-intelligence stack. It was what AMD believed it needed to understand in order to design the compute platforms that come next.
AMD’s own description of the acquisition was unusually candid on this point. World Labs, the company said, develops spatial-intelligence models that generate, reconstruct and simulate interactive 3D environments from text, image and video inputs, together with technology for robotic learning and simulation, and the purpose of bringing that team inside a semiconductor company was to strengthen AMD’s ability to develop hardware, software and systems around the needs of emerging models and applications.[1] Su framed the logic in terms of talent and integration rather than product:
“to have access to the world-class talent Fei-Fei has brought together” and to combine it with AMD’s capabilities in hardware, software and systems.
— Dr. Lisa Su, Chair and CEO, AMD [5]
Li, for her part, described the combination as a way to give World Labs the compute and engineering scale that its expanding ambitions required, and she spoke of accelerating the flywheel between software and hardware development for AMD and for the broader ecosystem.[5] Her statement in the official announcement was similarly direct:
Joining AMD would give the team “the resources and engineering depth to accelerate our research” and help define the infrastructure needed for the next era of AI.
— Dr. Fei-Fei Li, Co-founder and CEO, World Labs; Professor, Stanford University [1]
Industry analyst Patrick Moorhead summarized the strategic rationale from the model developer’s side in a single observation:
“Without having a focused hardware effort, AI is hobbled in efficiency.”
— Patrick Moorhead, Founder, Moor Insights & Strategy [6]
That distinction between a model that needs hardware and a chip company that needs to understand models is the heart of this paper, and it matters for reasons that extend well beyond one transaction. During the large-language-model era, computing platforms were largely optimized around artificial intelligence that ingested tokens, images and increasingly video, then returned digital outputs to a screen. The emerging spatial-intelligence era asks something considerably harder of the entire stack. An intelligent system may have to understand whether a box will fit through a doorway, whether a robotic hand can grasp an irregular object without dropping it, whether a forklift can safely pass another vehicle inside a crowded warehouse, whether a humanoid can reorganize an unfamiliar kitchen, or what will happen after a machine changes its position, speed, force, angle or sequence of actions. Language can describe those situations with great eloquence. Spatial intelligence must model them, predict them and, increasingly, act within them.
World Labs itself describes spatial intelligence as the ability to move from seeing toward doing: perceiving worlds, reasoning about objects and places across space and time, anticipating consequences and interacting reliably with virtual and physical environments.[1] Its first product, Marble, constructs persistent and navigable 3D worlds from text, images, video, panoramas and layouts, and the company raised $1 billion in February 2026, with Autodesk contributing $200 million and serving as an adviser, to accelerate that work.[7] AMD had already invested in World Labs before the acquisition and had struck an inference-optimization and training partnership with the company in 2025, and Li appeared as a guest speaker during AMD’s keynote at CES in January 2026.[4][3] The September announcement therefore did not arrive from nowhere; it was the culmination of an argument AMD had been making for most of the year, namely that understanding and navigating space could become a foundational layer of the next era of AI, with applications in simulation, interactive 3D environments and physical AI.
The economic significance of that argument goes considerably beyond 3D graphics, and it is worth pausing on why.
Large language models benefited from an extraordinary historical accident. Humanity had spent decades digitizing language before frontier AI systems arrived. Web pages, books, software code, encyclopedias, academic literature, forums, photographs, subtitles, documentation and other forms of digital information had accumulated at enormous scale. Whatever the unresolved questions surrounding quality, permission, licensing and copyright, the technical fact was remarkable: a substantial representation of human symbolic knowledge was already stored in machine-readable form, waiting for a model architecture capable of absorbing it. The transformer arrived, the corpus was there, and the two combined to produce the most rapid capability expansion in the history of computing.
There is no equivalent internet containing every physical interaction humans perform. There is no public corpus containing billions of properly synchronized demonstrations of workers assembling aircraft engines, technicians repairing transformers, nurses manipulating medical equipment, electricians entering substations, warehouse employees handling deformable packages, farmers operating machinery in rain, mechanics reaching around an engine block, or people opening millions of differently shaped drawers with different friction, weight and geometry. A language model can encounter millions of descriptions of a coffee cup, and it can discuss the cup with considerable sophistication. A robot must eventually learn how cups actually behave. It must distinguish ceramic from paper, full from empty, hot from cold, stable from slipping, reachable from obstructed. It must infer depth, orientation, force, collision, momentum, deformation and consequence. And when the physical world presents a configuration the machine has never encountered before, prediction must become action under uncertainty, in real time, with physical consequences if the prediction is wrong.
Fei-Fei Li made this argument in her widely circulated November 2025 essay, “From Words to Worlds,” in which she characterized today’s language models, for all their fluency, as systems that have never touched the world they describe:
Today’s most advanced models remain “wordsmiths in the dark; eloquent but inexperienced.”
— Dr. Fei-Fei Li, Professor, Stanford University; Co-founder, World Labs [8]
In the same essay she defined the world models that spatial intelligence requires through three essential capabilities: they must be generative, producing simulated worlds with perceptual, geometric and physical consistency; they must be multimodal, ingesting images, video, depth, text, gestures and actions; and they must be interactive, rolling the state of the world forward when an agent takes an action.[9] Those three properties are the technical specification for the category this paper examines.
This creates a fundamental transition in the economics of artificial intelligence. The scarce input is no longer only the GPU that trains the model or the megawatt powering the datacenter. Increasingly, it is grounded experience: high-quality video, synchronized sensors, robot trajectories, human demonstrations, teleoperation sessions, factory interactions, autonomous-driving encounters, digital twins and simulated variations that teach machines what happens when actions occur in space.
Figure made the issue unusually explicit in August 2026 when it introduced Index, its effort to construct a large physical-world training dataset for humanoid robots. The company argued that the information needed to scale a general-purpose robot does not exist on the internet and therefore must be gathered from physical reality:
The data needed to scale a truly general-purpose robot is not on the internet; “it has to come from the real world.”
— Figure AI, “Introducing Index” [10]
Figure reported that after four months in stealth its app had crossed 264,000 downloads across more than 100 countries, that more than 44,000 weekly active contributors—whom it calls Creators—had uploaded more than 16 million videos, that the system was processing 30 minutes of video every second, and that it had already paid $15 million to contributors while planning more than $1 billion in spending on data and compute over the following twelve months.[10][11] Nine days later the same company announced a partnership with Nscale for up to 100,000 Nvidia Vera Rubin GPUs, an initial compute commitment of $3.5 billion with intent to scale beyond $6 billion, and its founder described the company as now largely bound by the data and compute needed to train its Helix model.[12]
That pairing—a billion dollars for physical data and several billion for the compute to learn from it—is an important economic clue. In the language-model era, the internet was first treated primarily as a corpus. In physical AI, the world itself may increasingly be treated as a training environment. Warehouses become datasets. Factories become datasets. Autonomous fleets become datasets. Homes, construction sites and hospitals become datasets. Robot deployments become datasets. Simulators become machines for manufacturing additional experience at industrial scale. A robot that successfully performs one million tasks does more than produce labor; subject to privacy, contractual and legal constraints, it can potentially generate observations that improve subsequent machines. A fleet therefore has two possible economic outputs: the work performed today and the learning accumulated for tomorrow.
This paper calls that transformation Spatial Intelligence Models because the central development is larger than robotics and more specific than another generation of multimodal artificial intelligence. These models seek to acquire representations of space, persistence, geometry, motion, causality and action. They are systems designed not merely to describe the world but increasingly to predict it, simulate it and operate within it.
The transition has implications throughout the Five-Layer AI Economy that this series of papers has used as its analytical frame. Layer 1, energy, powers Layer 2, chips, which populate Layer 3, datacenters, which train and operate Layer 4, models, which enable Layer 5, applications and agents. Spatial intelligence changes what each layer must do. Layer 2 chips must increasingly accelerate simulation, rendering, perception, world modeling, robot-policy training and edge inference. Layer 4 models must move beyond linguistic and visual recognition toward representations of environments and consequences. Layer 5 applications and agents increasingly acquire bodies: humanoids, autonomous vehicles, drones, industrial machines, warehouse systems and other devices that continuously encounter physical reality. And the crucial feedback loop therefore changes shape, from a vertical hierarchy into a circuit.
Energy → Chips → Datacenters → Spatial Intelligence Models → Physical Agents → Real-World Experience → Better Spatial Intelligence Models
That final return path—from machines operating in the world back into model training—may become one of the most important economic loops of the 2027–2030 artificial-intelligence economy, and the remainder of this paper is an attempt to understand its structure, its participants, its bottlenecks and its consequences.
Why the Title “Spatial Intelligence Models”
I chose Spatial Intelligence Models because it names the capability at the center of the paper rather than one particular technology used to achieve it. “World model” is becoming an important industry term, and it will be used throughout this paper where companies use it themselves, but the broader objective is spatial intelligence: enabling artificial intelligence to understand objects, environments, positions, movement, relationships, time and consequences across both simulated and physical worlds. World Labs uses the term explicitly, and Li has argued for years that it is a distinct and in some respects older form of intelligence than language, one that humans exercise every time they pick up a mug or infer the three-dimensional structure of a molecule.[13] Google DeepMind, Nvidia, Waymo and the robotics developers are approaching similar problems through world models, physical-AI foundation models, simulation engines and vision-language-action systems. The title therefore creates enough room to analyze an emerging technological category rather than tying the paper to a single architecture or a single company, including the one AMD has just agreed to buy.
The title also describes the transition in language understandable beyond the AI research community. Large language models learned patterns in words; spatial intelligence models increasingly learn patterns in worlds. They must understand where things are, how they relate, how they move and what may happen after an action. That shift connects the AMD–World Labs transaction, Nvidia’s physical-AI ecosystem, Google DeepMind’s interactive world models, Waymo’s simulation strategy, Figure’s data-collection campaign, industrial digital twins and the debates among academic roboticists within one conceptual framework. For a Five-Layer AI Economy increasingly extending from datacenters into machines, “Spatial Intelligence Models” identifies the bridge between computational intelligence and physical action.
The paper proceeds in six sections. Section 1 examines why the data inheritance that powered language models has no physical equivalent, and why that scarcity reverses the cost curve of AI training. Section 2 maps the emerging spatial-intelligence stack across World Labs, AMD, Nvidia, Google DeepMind and Waymo, and situates it against the latest corporate results through the second quarter of 2026. Section 3 turns to the question of who owns the experience a robot generates, a question that industrial contracts will increasingly be forced to answer. Section 4 examines factories, warehouses and fleets as continuous learning systems. Section 5 looks forward to 2027–2030 and the possible emergence of markets for physical experience. Section 6 distills what we have learned into seven pillars, before a conclusion returns to why spatial intelligence belongs at the center of the next chapter of the Five-Layer AI Economy.

Section 1: From Text Abundance to Physical-Data Scarcity
Every technological revolution inherits something from the era that preceded it, and the first modern AI boom inherited an information infrastructure that previous generations had unknowingly spent decades constructing. Understanding the nature of that inheritance, and why it has no equivalent in the physical world, is the necessary starting point for understanding why spatial intelligence changes the economics of artificial intelligence rather than merely extending them. This section argues that the scarce input of the language-model era was compute applied to abundant data, that the scarce input of the spatial-intelligence era is grounded experience itself, and that the difference is not one of degree but of kind. It also argues that the most rigorous academic voices in robotics—at Berkeley, MIT and Stanford—have been making versions of this argument since at least 2025, and that the corporate announcements of 2026 are best read as the industry’s response to a constraint the academy identified first.
1.1 The Extraordinary Data Inheritance of the Language-Model Era
The World Wide Web was not designed as a training dataset for transformer models. GitHub was not created to educate coding agents. Digital books were not scanned for future foundation models. Wikipedia was not written as a pretraining corpus. Social networks, discussion boards and instructional videos did not emerge because engineers anticipated that neural networks would someday consume them. Yet collectively these systems created something unprecedented: an enormous digital representation of human communication, accumulated over roughly three decades and distributed across servers that any sufficiently capable crawler could reach.
This mattered because machine learning scales differently when observations already exist. For language, collection costs were often separated historically from training costs. The expensive human work of writing, translating, photographing, documenting, programming and explaining had already occurred, paid for by publishers, universities, employers, advertisers and the unpaid enthusiasm of hundreds of millions of people. Model developers subsequently faced the substantial challenge of assembling, cleaning, filtering, licensing and processing enormous corpora, but much of the underlying informational activity preceded the AI boom by years or decades. The marginal cost of an additional token of training data was, for a long time, close to zero, and the binding constraint on capability became the compute required to learn from data that was already there.
Physical intelligence faces almost the reverse problem. Human civilization contains enormous amounts of physical knowledge, but much of it has never been digitized in a form useful for machine learning. An experienced mechanic knows how much resistance indicates that a bolt is beginning to strip. A warehouse worker recognizes when a loosely packed carton must be lifted differently from a rigid one. A nurse develops an intuitive understanding of how equipment, patients and other people occupy space in a crowded ward. A construction worker anticipates how material will shift while being moved across uneven ground. A driver detects countless weak signals indicating that a pedestrian, cyclist or another motorist might behave unexpectedly. These skills contain information about physics, intentions, material properties, risk, geometry and causality. Yet they historically lived in human nervous systems rather than server racks, and they were transmitted through apprenticeship, imitation and practice rather than through text.
Spatial intelligence therefore confronts a paradox: the world contains nearly unlimited experience, but relatively little of it exists as machine-usable training data.
The most precise quantification of this paradox comes from UC Berkeley’s Ken Goldberg, whose two papers in Science Robotics in August 2025 introduced a phrase that has since become shorthand for the entire problem. Goldberg calculated how long it would take a single human to read all of the text used to train large language models and arrived at roughly 100,000 years; he then observed that robotics possesses nothing remotely comparable in physical-interaction data, and that manipulation is in any case far more complex than language, so the true requirement is likely larger still.[14]
Robotics faces what he calls a “100,000-year data gap” between internet-scale text and the scarcity of physical-interaction data.
— Ken Goldberg, Professor of Industrial Engineering and Operations Research, UC Berkeley [14]
Goldberg was skeptical of the idea that the gap could be closed simply by watching videos of humans, because a two-dimensional recording does not reveal the detailed motions, forces and contacts that a person is actually performing, and recovering three-dimensional structure from flat images remains hard.[15] His prescription was not to abandon learning but to combine it with what he called good old-fashioned engineering, so that robots become useful enough to be deployed and thereby begin generating their own data over time, a pattern he associated with Waymo’s vehicles and with warehouse robots that learn through continued use.[15] That last point deserves emphasis, because it anticipates the central economic argument of this paper: the way out of the data gap is not to find a hidden corpus but to build machines that produce experience as a byproduct of work.
1.2 Why Video Is Necessary but Not Sufficient
Video dramatically expands what artificial intelligence can observe, and the scale of video now being gathered for physical AI is genuinely new. Figure’s Index alone reports ingesting the equivalent of 4.9 years of human activity every day, and per 1,000 hours collected the company counts 373 unique tasks, 1,146 unique manipulated objects and 116 unique environments.[10] But video alone does not automatically provide complete physical understanding, and the strongest voices in robotics have been insistent on this point.
A camera records appearance from a viewpoint. A physical agent may require far more. It can need depth, joint position, force, torque, velocity, contact, tactile information, gripper state, audio, temperature, object identity, robot configuration, human intent, action labels, success or failure signals and precise temporal synchronization across all of these channels. A recording showing a person turning a valve contains useful visual information about the approach, the grip and the rotation. A robotics dataset becomes much more valuable if it also records the exact hand trajectory, the applied force, the tool orientation, the resistance encountered, the sensor state and the eventual outcome, because those are the quantities a controller must actually reproduce.
Rodney Brooks, the MIT roboticist who co-founded iRobot and Rethink Robotics, made this argument at length in a September 2025 essay that became one of the most debated documents in the field. Brooks surveyed fifty years of humanoid development and argued that human dexterity depends overwhelmingly on touch—on the dense mechanoreceptors in the fingertips and on proprioceptive feedback across the body—and that a training pipeline built primarily on camera footage is therefore learning from the wrong signal.[16]
“Collecting just visual data is not collecting the right data.”
— Rodney Brooks, Panasonic Professor of Robotics (emeritus), MIT [17]
Brooks went further, arguing that the field has not yet found the right engineered front-ends to extract the relevant structure from the raw sensory stream that the world presents, and that until it does, end-to-end learning for manipulation will not repeat the success it achieved in speech recognition and language.[16] One need not accept his most pessimistic conclusions to recognize the force of the underlying observation: the objective of a physical-AI dataset is not merely to see what happened but to learn the relationship between state, action and consequence, and that relationship is only partially visible to a camera.
Daniela Rus, director of MIT’s Computer Science and Artificial Intelligence Laboratory, has framed the same constraint from the perspective of safety rather than sensing. Data in the physical world, she has observed, is expensive, time-consuming and often incomplete, and the tolerance for error is categorically different from that of a chatbot:
In the physical world, “we cannot have mistakes and hallucinations,” because safety and reliability are harder to guarantee when AI directly affects motion.
— Daniela Rus, Director, MIT CSAIL [18]
This difference helps explain why physical-AI developers increasingly combine real-world video with teleoperation, multimodal sensor collection and simulation rather than relying on any one source. Figure’s own description of its Helix training regime lists three inputs—human demonstrations captured through teleoperation, real-world robot performance logs and internet-scale video of humans performing everyday tasks—with Index designed to scale the third category rather than replace the first two.[19] Recent academic work from Berkeley and NYU, including a December 2025 paper co-authored by Yann LeCun, Trevor Darrell and Jitendra Malik, has explored whether generated human videos can be converted into physically plausible robot trajectories, precisely because raw video lacks the physical annotations that control requires.[20] The direction of travel is clear: video is the broadest and cheapest source of physical observation, but it must be enriched, grounded and cross-checked against sensors, simulation and outcomes before it becomes experience in the sense that matters.
1.3 The Cost Curve Reverses
Language data historically became cheaper to reproduce as digitization spread, and once a document existed in digital form its marginal cost of copying was effectively zero. High-quality physical experience can remain expensive because collecting it may require hardware, people, environments and time, none of which can be copied.
A robot demonstration may require a real robot, which costs tens or hundreds of thousands of dollars and wears with use. A teleoperation session requires an operator, who must be trained, paid and scheduled. A warehouse dataset requires access to a warehouse, which is someone’s operating business. An autonomous-driving encounter requires vehicles moving through reality, burning fuel or electricity and exposing the operator to liability. Industrial manipulation data may require production equipment and permission from a factory operator who has every reason to protect proprietary processes. Medical robotics data may involve highly regulated environments where consent, privacy and safety review add months to every collection effort. Agricultural robotics must encounter weather, soil, terrain and seasonal variations that cannot simply be downloaded and that arrive on nature’s schedule rather than the developer’s.
Even when sensors become inexpensive, relevant variation is difficult to manufacture. Imagine teaching a home robot to unload a dishwasher. There is no single dishwasher environment. The machine may encounter different layouts, racks, plates, glasses, bowls, utensils, lighting conditions, countertop heights and obstacles. Objects may be wet, fragile, transparent, reflective or partially hidden. A child may interrupt. A drawer may stick. A glass may already contain liquid. Another object may fall while the robot is moving. The combinatorial space becomes enormous, and it is enormous in a way that a single collection campaign cannot exhaust. Physical AI is therefore not simply a problem of accumulating more petabytes. It is a problem of accumulating the right experiences across enough variations, and that is a fundamentally more expensive proposition than crawling the web.
Figure’s payment figures illustrate the new cost structure in miniature. Fifteen million dollars paid for sixteen million videos implies less than a dollar per clip on average, but payment is keyed to recording time rather than clip count, the denominator includes footage discarded in review, and the company has not yet published how much Index improves its models.[21][11] What the numbers do establish is that physical data now has a price, that the price is being paid at scale, and that the company paying it regards a billion dollars of further data and compute spending as the necessary next step.
1.4 From Dataset Size to Experience Coverage
The useful metric for spatial intelligence may eventually become less about raw dataset size and more about experience coverage. A ten-million-hour dataset containing repetitive highway driving may be less useful for some objectives than a much smaller collection containing rare intersections, unusual weather, emergency vehicles, construction zones, atypical pedestrian behavior and edge cases. Likewise, a warehouse robot may benefit more from examples of difficult manipulations than from millions of identical successful movements.
This could change how AI companies discuss data scale. Instead of asking how many tokens a model has seen, they may increasingly ask how many environments, how many tasks, how many object classes, how many embodiments, how many failure modes, how many variations and how much real-world coverage a training set contains. Figure’s decision to publish task, object and environment counts per thousand hours, rather than simply hours, is an early sign that the industry’s own vocabulary is shifting in this direction.[10] The Open X-Embodiment collaboration among Google DeepMind, Stanford and other research groups, which pooled more than one million real robot trajectories across many robot forms, represents the academic version of the same insight: diversity of embodiment and task matters as much as volume.[21]
| Language-Model Era Metric | Spatial-Intelligence Era Metric |
| Tokens in the pretraining corpus | Environments, tasks and object classes covered |
| Parameters and FLOPs | Embodiments and sensor modalities represented |
| Benchmark accuracy on text tasks | Success rate under distribution shift in the physical world |
| Deduplicated web pages | Distinct failure modes observed and recovered from |
| Corpus size in terabytes | Real-world coverage of rare, dangerous or costly scenarios |
Table 1. The vocabulary of data scale shifts from volume to coverage as intelligence moves from words to worlds.
This is the vocabulary of spatial intelligence, and it is also the vocabulary of a market in which not all observations are worth the same amount, a theme to which Section 5 returns.
1.5 Real Experience and Synthetic Experience Become Complements
Physical data scarcity does not mean every experience must occur in reality. In fact, scarcity increases the value of simulation, because one expensive physical observation can become the seed for thousands or millions of controlled variations.
World Labs demonstrated this logic through its real-to-sim-to-real work following its July 21, 2026 acquisition of SceniX, a robotics-simulation startup founded by Columbia University professors Yunzhu Li and Changxi Zheng.[22] The combined team described an engine that reconstructs physical robots, sensors, environments, objects and interactions as simulations that preserve task-relevant observations and dynamics, then varies configuration, appearance, robot state and camera viewpoint so that policies can be trained and evaluated across many circumstances without repeatedly performing every experiment on physical hardware.[23] The company’s claim for the approach was striking in its ambition:
The system allows robots to “learn complex manipulation tasks with zero real-world training data” and to predict in simulation which policies will succeed in reality.
— World Labs, “Building Worlds That Train Robots” [23]
The academic lineage of this idea runs through MIT’s Improbable AI Lab, where Marcel Torne, Pulkit Agrawal and colleagues published a real-to-sim-to-real approach for robust manipulation in 2024 that reconstructed real scenes for policy training and then transferred the resulting policies back to hardware.[24] What changed between that paper and World Labs’ 2026 engine is not the concept but the fidelity and the scale at which generative world models can now produce the simulated side of the loop.
This creates a potentially transformative multiplier that can be written as a single chain:
One physical experience → reconstructed environment → thousands of synthetic variations → improved policy → new physical deployment → additional experience
The boundary between real and synthetic data therefore becomes less useful than the distinction between grounded and ungrounded experience. Simulation becomes valuable when it preserves enough of the relevant structure of reality to improve behavior outside the simulator, and worthless or actively harmful when it does not. Determining which is which—measuring, in effect, the exchange rate between a simulated hour and a real one—becomes one of the defining engineering problems of spatial intelligence, and, as the next section shows, one of the defining commercial contests as well.

Section 2: World Labs, AMD, Nvidia, Google and the Emerging Spatial-Intelligence Stack
If Section 1 established the constraint, this section maps the industrial response to it. Between the autumn of 2025 and the autumn of 2026, the companies best positioned to profit from artificial intelligence made a series of moves that only make sense if one assumes they had concluded that the next generation of models would be trained not primarily on text but on representations of space, motion and consequence. A Stanford professor’s startup raised a billion dollars and then sold itself to a chipmaker for eight times that amount. A Turing Award winner left Meta and raised the largest seed round in European history to build world models rather than language models. Google connected its interactive world model to two decades of Street View imagery. Waymo rebuilt its simulator on top of that world model. Nvidia declared that the “big bang of physical AI” had begun and signed up the four largest industrial-robot manufacturers on earth. And the largest humanoid startup committed billions of dollars to compute explicitly to learn from the physical data it was paying the public to record. Read individually, each is a business story. Read together, they describe the construction of a new technology stack, and this section attempts to describe that stack layer by layer, grounded in the most recent financial disclosures available through the second quarter of 2026.
2.1 World Labs: From Pixels to Persistent Worlds
World Labs provides a useful starting point because its ambition is explicitly spatial and because it has now become part of the semiconductor industry. Its first product, Marble, generates spatially coherent, persistent environments that users can move through, edit and expand, from inputs that include text, images, video, panoramas and 3D layouts, and it represents those environments as Gaussian splats or meshes that can be rendered interactively on phones, laptops and headsets.[13] The important word in that description is not simply generate. It is persistent.
Persistence separates a picture from a world. If an AI creates an image of a kitchen, it produces pixels that are correct from one viewpoint. If an AI creates a spatial environment, the kitchen should retain geometry when the observer changes position. Objects should maintain relationships. Moving behind an island should reveal an internally consistent scene rather than a freshly hallucinated one. An agent should be able to navigate through the environment and encounter a coherent representation of space that responds to its actions. Li has argued that this is precisely the capability language cannot supply, because language is a lossy, low-bandwidth channel for describing a rich three- and four-dimensional world, and she has described spatial intelligence in an earlier interview as a prerequisite for anything deserving the name of general intelligence:
Solving spatial intelligence is “a fundamental and critical step towards full-scale intelligence.”
— Dr. Fei-Fei Li, Professor, Stanford University [25]
World Labs describes its longer-term objective as models capable of perceiving, generating, reasoning about and interacting with both virtual and physical worlds, and its acquisition of SceniX in July 2026 was framed by the company as the moment the research program extended from generating worlds to using them as a closed loop with real-world robot learning.[22][26] That progression—perceive, generate, reason, interact—describes the transition from passive visual intelligence toward active spatial intelligence, and it is the progression AMD has now purchased.
2.2 Why AMD’s Interest Matters: Reading the Deal Against the Q2-2026 Results
AMD’s agreement to acquire World Labs matters because it connects model evolution directly to semiconductor architecture, and it arrives at a moment when AMD’s own results show how completely its business has been reorganized around AI compute.
In its second-quarter 2026 results, reported on August 4, AMD posted record revenue of $11.5 billion, up 50 percent year over year, with Data Center segment revenue of $6.7 billion, up 107 percent, representing 58 percent of company revenue and carrying a 31 percent operating margin.[27][28] Lisa Su told investors that the company was entering the second half with strong momentum as EPYC demand accelerated, Instinct deployments scaled and its Helios rack-scale platform began to ramp:
AMD enters the second half as “Helios begins to ramp,” with Data Center revenue having more than doubled year over year.
— Dr. Lisa Su, Chair and CEO, AMD (Q2-2026 results) [27]
The accompanying earnings presentation referenced a strategic partnership with Anthropic involving the deployment of up to two gigawatts of AMD Instinct GPUs in Helios racks, alongside collaborations with Microsoft and Cisco, and management projected server-CPU revenue growth of more than 80 percent year over year in the second half of 2026 and more than 70 percent in 2027.[29][28] These are the numbers of a company whose center of gravity has shifted decisively toward the datacenter, and the World Labs acquisition should be read as an attempt to ensure that the next shift—toward physical AI—does not catch its hardware roadmap unprepared.
Earlier AI infrastructure was often discussed as though models came first and hardware simply became fast enough to execute them. The relationship is increasingly bidirectional. New model architectures expose bottlenecks. Those bottlenecks influence chips. New chips make different architectures economical. Those architectures create different workloads. Those workloads influence networking, memory, storage and system design. The cycle repeats, and the company that sees the next workload earliest has an advantage in designing for it.
Spatial intelligence adds a distinctly different workload profile. Physical AI may involve enormous video pipelines, 3D reconstruction, multimodal sensor fusion, real-time simulation, physics, rendering, reinforcement learning, world-model inference and low-latency deployment on edge devices with strict power budgets. A model developing a robot policy may train inside a datacenter. The policy eventually executes in a warehouse. The data returns from the warehouse. A simulation system recreates failures. The model retrains. Updated behavior is deployed back to machines. The compute architecture must therefore stretch from Layer 2 datacenter accelerators toward edge processors inside Layer 5 machines, and AMD’s own statement that World Labs’ research would provide deeper insight into emerging AI workloads and help guide its technology roadmaps across hardware, software and systems is an acknowledgment that it intends to design for that entire span.[1] Li’s remark on the Latent Space podcast that world models and spatial data are the most compelling way to absorb modern GPU clusters, compared with language alone, is the same argument stated from the demand side.[13]
AMD’s transaction can thus be interpreted as something more consequential than adding an AI laboratory to a chip company. It gives a semiconductor firm direct visibility into how one important class of future models may actually consume computation, and it does so at a moment when Nvidia is pursuing the identical strategic goal through a broader and more mature platform.
2.3 Nvidia: Building the Operating Environment for Physical AI
Nvidia is approaching the same transition through a physical-AI stack that now spans the entire loop from simulation to deployment, and its financial results provide the clearest available measurement of how large that business has become.
In its second quarter of fiscal 2027, ended July 26, 2026, Nvidia reported revenue of $96.2 billion, up 106 percent from a year earlier, with Data Center revenue of $89.0 billion, up 117 percent, and Edge Computing revenue of $7.2 billion, a category the company now uses to capture devices for agentic and physical AI, including robotics and automotive.[30][31] The company guided to $108 billion for the following quarter and, on the earnings call, chief financial officer Colette Kress indicated a target of roughly 70 percent revenue growth for fiscal 2028.[32] On the same call, management described an expanded partnership under which Amazon Web Services would deploy an additional two million GPUs and adopt Nvidia’s full physical-AI stack—Omniverse, Cosmos, Isaac and Jetson—to power its fleet of warehouse robots, and noted that on-premises automotive revenue had reached $8 billion on a trailing-twelve-month basis.[33] Nvidia’s fiscal-2026 proxy statement had already disclosed $6 billion of physical-AI revenue for that year, alongside the release of Cosmos for physical AI and Alpamayo for autonomous vehicles.[34] Jensen Huang’s framing of the overall buildout, delivered with the first-quarter results, sets the scale:
The buildout of AI factories is “the largest infrastructure expansion in human history.”
— Jensen Huang, Founder and CEO, NVIDIA [31]
Within that expansion, the physical-AI components are architecturally specific. Omniverse provides libraries and services for constructing and operating digital twins and physically based simulations. Isaac Sim provides robotics simulation, testing and synthetic-data generation, accepting inputs from real-world captures, CAD models and robot descriptions that developers can transform into simulated environments in which robot and sensor models operate and from which synthetic datasets can be generated by varying position, appearance or lighting. Isaac Lab supports large-scale robot learning, with version 3.0 introduced in early access at GTC 2026. Cosmos supplies world foundation models, with Cosmos 3 described by Nvidia Research as an omnimodal world-model family for physical AI. GR00T targets humanoid robotics, with GR00T N1.7 released and N2 previewed. Jetson hardware carries inference toward machines at the edge.[35][36][37]
At GTC in March 2026, Nvidia also released what it called the Physical AI Data Factory Blueprint, an open reference architecture for gathering, shaping and assessing both real-world and simulated data, and its vice president of Omniverse and simulation technology stated the company’s thesis about the constraint it was designed to relieve:
Physical AI is the next frontier of the AI revolution, where “success depends on the ability to generate massive amounts of data.”
— Rev Lebaredian, Vice President of Omniverse and Simulation Technology, NVIDIA [38]
One observer summarized Nvidia’s strategy as an attempt to swap robotics’ data problem for a compute problem—to convert a shortage of experience into a demand for GPUs by manufacturing experience in simulation.[39] The characterization is apt, and it reveals something important about the future AI economy. The model alone is not the product. The environment in which the model learns becomes part of the infrastructure. Traditional AI factories convert electricity and silicon into tokens. Physical-AI factories may convert electricity and silicon into simulated experience. That experience trains machines. Machines then generate additional real experience. Nvidia’s ambition is to own the platform on which every stage of that loop runs.
2.4 The Digital Twin Becomes a Training Ground
For decades, industrial digital twins primarily served engineering functions: design, optimization, monitoring and planning. Spatial intelligence transforms them into something else. They can become schools for machines.
The scale of the installed base that could be enrolled in those schools is substantial. According to the International Federation of Robotics, the global operational stock of industrial robots reached 4.66 million units at the end of 2024, up 9 percent in a year, with 542,076 new installations during 2024 and China accounting for more than two million units of the operating stock.[40][41] At GTC 2026, Nvidia announced that ABB Robotics, FANUC, KUKA and Yaskawa—which together account for a global install base exceeding two million of those robots—were integrating Omniverse libraries and Isaac simulation frameworks into their virtual commissioning systems to develop and validate complex robot applications and entire production lines through physically accurate digital twins, and that they were embedding Jetson modules into their controllers for real-time inference at the edge.[35] Nvidia’s Mega Omniverse Blueprint similarly allows enterprises to design, test and optimize robot fleets and AI agents inside a facility digital twin before a single robot is deployed on the floor.[36]
This changes the economic role of a factory model. A digital twin previously helped humans understand the plant. Increasingly it can help machines learn the plant. And once machines learn inside these environments, industrial facilities themselves become part of the model-development stack. This is why Layers 3, 4 and 5 of the Five-Layer AI Economy increasingly overlap. The datacenter trains the intelligence. The simulation environment provides experience. The physical facility supplies grounding. The robot converts intelligence into action. Then the robot produces more observations, which flow back to the beginning.
2.5 Google DeepMind: Worlds as Unlimited Curricula
Google DeepMind’s Genie research illustrates another branch of the same transition. Genie 3 is a general-purpose world model capable of generating interactive environments in real time, and Google describes world models as systems that use their understanding of physical environments to simulate how worlds evolve and how actions affect them. Its broader ambition is that agents can learn inside an effectively unlimited curriculum of generated environments.[42]
The phrase unlimited curriculum deserves attention. Human learning is constrained by time. Robot learning in reality is constrained by hardware wear, safety, personnel, facility availability and cost. Simulation changes the arithmetic because virtual agents can experience events faster, more often and with less physical consequence. A simulated robot can fail repeatedly. A physical industrial robot cannot casually damage equipment thousands of times while learning. A simulated autonomous vehicle can encounter bizarre road scenarios repeatedly. A real vehicle cannot manufacture dangerous incidents merely to collect training examples. A sufficiently capable world model therefore functions as a scenario generator, and the economic value lies partly in producing experiences that reality provides too rarely, too dangerously or too expensively.
In May 2026, Google connected Genie to Street View, anchoring generated environments directly in imagery of real places drawn from a corpus the company has been gathering for roughly two decades. DeepMind researcher Jack Parker-Holder described the purpose in terms that placed robotics and agents ahead of entertainment, and offered the example of a robot being deployed in London that rarely sees direct sunlight, for which Genie could simulate the scarce occasions when sun glints off the housing so that the phenomenon does not surprise the machine when it finally occurs.[43]
Grounded, interactive simulation serves agents and robots as much as people, and “that’s always been the thesis of Genie.”
— Jack Parker-Holder, Research Scientist, Google DeepMind [43]
Google’s own product leaders cautioned that the Street View integration remained experimental and could not yet produce a faithful reconstruction of a specific street.[42] But the strategic implication was widely noted: a dataset gathered for one purpose—helping people find their way—had acquired a second life as spatial-AI training infrastructure that no competitor could easily replicate.[44] Two months later, on July 30, 2026, DeepMind released Gemini Robotics ER 2, an embodied-reasoning model that adds real-time video understanding, task-progress tracking, lower-latency orchestration and multi-robot collaboration, and that the company reports outperforming frontier general-purpose models on physical benchmarks such as human-proximity estimation.[45] Generating worlds and reasoning within them are, at Google as at World Labs, two halves of the same program.
This suggests a much broader future possibility. Companies possessing large repositories of geospatial imagery, mapping data, video, industrial telemetry or autonomous-driving observations may discover that these assets can serve as foundations for spatial models. Maps become environments. Video archives become demonstrations. Vehicle fleets become sensors. Warehouses become laboratories. The value of a company’s historical observations may therefore depend not simply on what the data says about customers today, but on what it can teach machines about reality tomorrow.
2.6 The Contrarian Bet: Yann LeCun and AMI Labs
The intellectual argument for world models also acquired a well-funded champion outside the incumbents. In March 2026, Advanced Machine Intelligence Labs, founded by Turing Award winner Yann LeCun after his departure from Meta, closed a $1.03 billion seed round at a $3.5 billion pre-money valuation—the largest seed round in European history—to build world models based on his Joint Embedding Predictive Architecture, which learns abstract representations of real-world sensor data rather than predicting outputs token by token, and which targets industrial, robotic and healthcare applications where the limitations of language models are most consequential.[46][47] LeCun has argued for years that large language models in their current form will not lead to genuinely intelligent machines, and the size of the round suggests that investors are prepared to fund that thesis at scale. His chief executive was candid about the hype the term was likely to attract:
“In six months, every company will call itself a world model.”
— Alexandre LeBrun, CEO, AMI Labs [46]
The prediction is worth taking seriously as a warning. The term will be stretched. But the convergence is real: World Labs, DeepMind, Nvidia, Waymo, AMI Labs and the robotics foundation-model companies such as Physical Intelligence—founded by Berkeley’s Sergey Levine and Stanford’s Chelsea Finn among others and valued at $5.6 billion after a $600 million raise in November 2025—are all building systems whose central purpose is to represent and predict the physical world rather than to continue text.[48]
2.7 Autonomous Vehicles Show the Economics Early
Autonomous driving provides perhaps the clearest existing demonstration of the loop between real and simulated experience, because it is the one domain in which that loop has been operating at commercial scale for several years.
Waymo announced its World Model in February 2026, describing a frontier generative model for large-scale, hyper-realistic autonomous-driving simulation built on top of Genie 3 and post-trained for the driving domain, capable of jointly generating camera imagery and lidar point clouds and of continuing a real recorded event into a counterfactual simulated one.[49] The company said its Driver had accumulated nearly 200 million fully autonomous real-world miles while simultaneously experiencing billions of miles in virtual worlds, and it crossed the 200-million-mile threshold later the same month.[49][50] Waymo explicitly justified the investment by the need to expose its system to rare, safety-critical scenarios—tornadoes, flooded streets, even an elephant in the road—that a fleet cannot encounter often enough in reality, while acknowledging that it had not published independent benchmarks of the resulting safety gains.[51]
This relationship is important because it is the template. Real miles establish grounding. Virtual miles create scale. Fleet deployment exposes edge cases. Simulation reproduces and varies them. Improved models return to the fleet. The autonomous-driving industry has therefore already built an early version of what robotics may adopt much more broadly, and Section 4 examines how the same logic is now spreading from vehicles to warehouses.
2.8 The Emerging Spatial Intelligence Stack
The resulting technology stack can be summarized as a sequence of layers, each of which is now occupied by identifiable products and companies.
| Layer | Function | Representative Systems (2025–2026) |
| A — Observation | Cameras, lidar, radar, tactile sensors, microphones, robot state, human demonstrations and other measurements | Figure Index; Uber AV Labs sensor fleets; Waymo fleet; teleoperation programs |
| B — Reconstruction | Transforming observations into persistent representations of environments | World Labs Marble; SceniX real-to-sim; Street View grounding in Genie |
| C — Simulation | Generating controllable environments and variations | Nvidia Omniverse and Isaac Sim; Newton physics engine; Isaac Lab |
| D — World Modeling | Predicting how environments evolve under action | Genie 3; Waymo World Model; Nvidia Cosmos 3; AMI Labs JEPA models |
| E — Policy Learning | Learning which actions accomplish objectives | Figure Helix; Physical Intelligence π-series; Nvidia GR00T; Gemini Robotics |
| F — Embodiment | Deploying intelligence into robots, vehicles and other machines | Humanoids; Waymo Driver; Amazon warehouse robots; ABB, FANUC, KUKA, Yaskawa arms with Jetson controllers |
| G — Experience Return | Sending real-world outcomes back into training and evaluation | Fleet logs; failure replay; Nvidia Physical AI Data Factory Blueprint |
Table 2. The seven-layer spatial-intelligence stack. Spatial intelligence models sit at Layers B through D, but their economic significance arises from the entire loop.
Spatial intelligence models sit near the center of this system, but their economic significance arises from the entire loop, and the loop closes only if the experience generated at Layer G can lawfully and technically return to Layer A. Whether it can is not a technical question. It is a question of ownership, and it is the subject of the next section.

Section 3: Who Owns the Experience Generated by a Robot?
The previous two sections argued that grounded experience is becoming the scarce input of artificial intelligence and that an industrial stack is being assembled to collect, simulate and learn from it. This section turns to a question that the language-model era encountered only belatedly, in courtrooms and licensing negotiations, but that the spatial-intelligence era will confront from the outset: when a machine learns from its work, who owns what it learned? The question sounds abstract. It is not. It will be answered, clause by clause, in the deployment contracts that manufacturers, logistics operators, hospitals and homeowners sign with robot providers over the next several years, and the answers will shape the market structure of physical AI as decisively as semiconductor supply shaped the generative-AI industry. This section sets out the problem, explains why robot data is more sensitive than the phrase suggests, and argues that privileged access to reality is becoming a competitive moat in its own right.
3.1 The Ownership Problem Nobody Needed to Solve Before
When industrial robots were mostly deterministic machines following predefined instructions, the ownership of their “experience” mattered relatively little. A welding robot repeatedly performed a programmed sequence. Its operational logs might matter for maintenance and quality control. But few companies considered every robotic movement a strategic training asset, because the robot did not learn and its logs could not make another robot better.
Learning robots change the question entirely. Suppose a humanoid works inside an automobile plant. The robot is manufactured by Company A. Its intelligence comes from a model developed by Company B. The factory belongs to Company C. A systems integrator from Company D configured the deployment. Factory employees occasionally teleoperate the robot when it encounters a task it cannot complete. Sensors record machinery, tools, production processes and worker movements throughout every shift. Over a year the robot performs one million tasks, some of them novel, some of them failures, some of them recoveries that no engineer anticipated.
Who owns the resulting experience? Company A, because it owns the hardware architecture through which the experience was gathered? Company B, because its model generated the actions and its model would be improved by the data? Company C, because the experience occurred inside its facility, on its processes, using its materials? The integrator, because it configured the system that made the experience possible? The workers, because their movements and interventions were captured and may be what actually taught the machine? Or some combination determined contractually, in advance, by parties who may not yet fully understand the value of what they are allocating?
| Party | Basis of Claim | What They Might Demand |
| Robot manufacturer | Owns the hardware and sensors through which experience is captured | Rights to use logs for product improvement across all customers |
| Model developer | Its policy generated the actions; its model gains from the data | Permission to train general models on deployment data |
| Facility owner | Experience occurred on its premises, processes and materials | On-premises retention, exclusion of sensitive zones, compensation or reciprocal model access |
| Systems integrator | Configured the deployment that made the data possible | Rights to derived environment models for reuse |
| Workers and teleoperators | Their demonstrations and interventions supplied the corrective signal | Consent, opt-out, compensation or recognition |
| Regulators and the public | Safety, privacy and labor law | Auditability, data-protection compliance, safety reporting |
Table 3. Competing claims on robot-generated experience. None of these claims was economically significant when robots did not learn.
This question is likely to become increasingly important because robot experience can potentially improve models deployed elsewhere, including in the facilities of the facility owner’s direct competitors. The moment learning becomes transferable, experience becomes an asset, and assets need owners.
3.2 Physical Data Contains More Than Motion
The phrase “robot data” can sound innocuous, as though it referred only to joint angles and gripper positions. It is not innocuous. A robot operating inside a factory may observe production volumes, facility layouts, machine configurations, process sequences, worker behavior, product designs, inventory levels, quality-control procedures, security practices, equipment failures, supplier materials, production bottlenecks and proprietary techniques. Training data collected during deployment can therefore contain commercially sensitive information far beyond the movements necessary to teach a robot to grasp a part, and a model trained on that data may encode aspects of a customer’s competitive advantage in ways that are difficult to audit or remove.
The same issue appears in homes, where the stakes are personal rather than commercial. A domestic robot could potentially observe layouts, possessions, routines, family members, documents, screens, conversations and daily habits. Figure’s Index program, which pays members of the public to record themselves performing household tasks, and its earlier Project Go-Big partnership with Brookfield, which gave the company access to residential and commercial properties for first-person video collection, are precisely efforts to gather this kind of intimate domestic data at scale, and both raise questions about consent and downstream use that the company will have to answer as the dataset grows.[19] In hospitals the sensitivity becomes greater still, because clinical environments combine protected health information with life-safety consequences. Uber, for its part, has said that the sensor kits on its data-collection vehicles are exterior-facing rather than pointed inside the cabin, an early example of a company designing the boundary of collection with these concerns in mind.[52]
The economics of spatial intelligence therefore cannot be separated from data governance. A model developer that can learn from every deployment will improve faster than one that cannot, but a model developer that learns from deployments without clear permission will eventually encounter the legal, reputational and regulatory consequences that the language-model industry is still working through with respect to text.
3.3 The Industrial Contract May Become the New Data License
During the LLM era, disputes focused heavily on whether text, images or other digital works could be used to train models, and those disputes were adjudicated after the fact because the data had been gathered before anyone thought to negotiate. Physical AI creates a different contractual frontier, and it offers the industry a chance to negotiate before the data exists.
The important agreement may increasingly be the deployment contract. A manufacturer purchasing robots might ask whether operational recordings can leave the facility; whether they can be used to improve the general model; whether the model provider can retain derived representations such as reconstructed digital twins even if raw recordings are deleted; whether competitors can benefit from learning generated in the plant; whether the manufacturer can receive compensation for particularly valuable training experiences; whether sensitive areas can be excluded from capture; whether data can remain on-premises; whether training can occur through federated methods that move gradients rather than recordings; whether simulations derived from the factory can be reused elsewhere; and whether workers can opt out of certain recordings.
These are not peripheral legal clauses. They can affect model quality and competitive advantage directly. A robotics provider obtaining permission to learn from thousands of customer deployments may improve faster than a competitor whose customers prohibit centralized training, and a customer who grants broad permission may be subsidizing the capability of a system that will later be sold to its rivals. Both sides have reason to think carefully, and the sophistication of these negotiations will likely rise as the value of the underlying experience becomes clearer.
3.4 Deployment Creates a Data Network Effect
This leads to one of the central economic propositions of spatial intelligence: the more machines a company deploys, the more physical experience it can potentially collect; the more experience it collects, the more capable its models can become; the more capable its models become, the more attractive its machines may become to the next customer. That is a feedback loop, and it has a familiar shape.
The internet economy produced network effects through users. Each additional user made a platform more valuable to every other user. The robotics economy may produce experience effects through deployments. Each additional deployment makes the model more capable for every other deployment. A company with 100 robots sees less of the world than a company with one million—assuming the latter can lawfully and technically use the resulting observations. Amazon’s statement in September 2026 that it has manufactured more than one million robots in the United States, deployed across more than 300 facilities, alongside its decision to open a fourth robot-manufacturing plant in Greenwood, Indiana, illustrates what a fleet of that scale looks like in practice, and Nvidia’s disclosure that AWS is adopting its full physical-AI stack for those warehouse robots suggests that Amazon intends to learn from it.[53][54][33] This makes fleet size strategically important for reasons beyond unit sales, and it explains why companies that have not yet sold many robots are nonetheless spending aggressively to acquire experience through other means.
3.5 Humans May Become Data Producers for Machines
Figure’s Index project provides a striking example of one of those other means. Instead of waiting for humanoids to accumulate every required experience themselves, Figure has built a system through which humans contribute videos intended to expand robot-training coverage, and in August 2026 it reported millions of contributions and direct payments to participants.[10] Uber has built the vehicular equivalent: its AV Labs division is deploying up to 500 Hyundai Ioniq 5 vehicles equipped with fourteen cameras, eight lidars, nine radars and onboard computers, driven by humans, to capture the unusual situations its drivers encounter every day and to sell that data to autonomous-vehicle partners including Nvidia and Wayve.[55][56] Uber’s head of AV Labs was explicit about what its partners lacked:
To fine-tune the software, “what you need is [more] data.”
— Danny Guo, Vice President of Engineering and Science, Uber AV Labs [55]
Uber’s chief technology officer has described data as the bottleneck in autonomous-vehicle development, and the company has said it already operates collection fleets capturing more than 100,000 hours of footage across the United States and Europe.[56] Earlier in 2026 it launched Uber Autonomous Solutions, a division intended to handle training data, mapping, fleet financing and regulatory services for its roughly two dozen autonomy partners, positioning the ride-hailing platform as infrastructure for other companies’ robots rather than as a robot developer itself.[57]
This reverses the traditional relationship between automation and labor. Humans do not merely perform work that robots may later automate. Humans can also produce demonstrations used to teach robots how work is performed, and they can be paid for doing so. A new category of labor may therefore expand: teleoperators, demonstrators, scenario collectors, robot trainers, simulation builders, failure annotators, quality reviewers and physical-data contributors. The labor market surrounding artificial intelligence may include not only people working with AI but people manufacturing experiences for AI, and the terms on which they do so—per minute of footage, per demonstration, per intervention—are being set now, largely without public scrutiny.
3.6 Worker Knowledge Becomes Machine Knowledge
This raises deeper economic questions that the academic literature on automation has been debating for years. A veteran factory employee may possess twenty years of tacit skill. If the employee teleoperates a robot while performing difficult tasks and those sessions are used to improve a general model, some portion of that tacit knowledge becomes reproducible machine capability. What was previously embodied human expertise becomes encoded behavior that can be copied at near-zero marginal cost to every other robot running the same model.
Daron Acemoglu of MIT and Pascual Restrepo have long argued that the direction of technological change is a choice rather than a law of nature, that automation which merely substitutes for labor without raising productivity substantially—what they call so-so automation—reduces the labor share without delivering commensurate gains, and that markets left to themselves may over-invest in that kind of automation relative to technologies that create new tasks for people.[58] Acemoglu has argued that government policy, funding and leadership will be critical in steering AI toward broadly shared prosperity rather than relying on the market to set the direction of change.[59] Erik Brynjolfsson of Stanford, responding in the same forum, emphasized the alternative path:
“When technology complements labor, wages tend to rise,” creating more broadly shared prosperity.
— Erik Brynjolfsson, Director, Stanford Digital Economy Lab [59]
The teleoperation-to-training pipeline sits squarely on the fault line between these two visions. In the near term it is complementary: the human is essential, the robot cannot work without the human’s intervention, and the human is paid for supplying it. In the longer term it is designed to be substitutive: the purpose of recording the intervention is to make the next intervention unnecessary. Companies will have to confront questions about consent, compensation and workplace governance, and different jurisdictions and industries will answer them differently. The point is not that every robotic observation must create a new property right. The point is that organizations can no longer assume physical experience has negligible informational value. When experience improves reusable models, experience becomes economically significant, and the people who supply it have a legitimate interest in the terms.
3.7 The Competitive Moat Moves Into the World
The first AI race emphasized model parameters, research talent, GPUs and datacenters. The spatial-intelligence race may add another strategic asset: privileged access to reality.
A logistics company sees logistics. A manufacturer sees factories. A ride-hailing network sees roads, and Uber has concluded that seeing roads is worth more as a product sold to autonomy developers than as an input to its own abandoned self-driving program.[57] A mapping company sees geography, and Google has connected twenty years of that geography to a world model.[44] A warehouse operator sees manipulation. A robotics company sees machine interaction. A consumer-device company sees homes. A defense contractor sees specialized environments. A healthcare system sees clinical workflows.
The winners may not always be the companies with the largest general datasets. They may be the organizations with access to the most valuable environments for a particular task, and with the contractual permission to learn from them. That permission, as much as any model architecture, may prove to be the decisive asset of the coming decade, and it is why the next section treats factories, warehouses and fleets not as customers of physical AI but as participants in its production.

Section 4: Factories, Warehouses and Fleets as Continuous Learning Systems
For more than a century, industrial economics treated the factory primarily as a location where capital equipment and labor transformed inputs into products, and the warehouse as a location where those products waited to be moved. Spatial intelligence introduces a second output from both places, and from the vehicle fleets that connect them: experience. This section examines how facilities and fleets are becoming continuous learning systems, why warehouses in particular occupy a privileged position in that transformation, why the million-robot threshold matters, why failure data may be worth more than success data, and why teleoperation—so often dismissed as a temporary crutch—deserves to be understood as durable infrastructure. The section builds on the observation, now shared across academic robotics and the largest industrial companies, that the way out of the physical-data gap is to deploy machines that generate experience as a byproduct of useful work.
4.1 The Factory Stops Being Only a Place of Production
Every movement of an autonomous forklift can reveal something about navigation. Every robot manipulation can produce information about grasping. Every intervention by a human operator can identify a model weakness. Every production anomaly can create an edge case. Every successful recovery can become a useful demonstration. The facility becomes both an operating environment and a learning environment, and the two roles reinforce each other.
This can fundamentally change how companies value deployment. A robot installed today may not only reduce the cost of a task. It can improve the system that performs tomorrow’s tasks, in this facility and in others. The capital budgeting of automation has traditionally compared the cost of a machine against the labor it displaces over its depreciation life. Spatial intelligence adds a term to that equation—the value of the experience the machine will generate—that is difficult to estimate but that the most sophisticated operators have clearly begun to price in. When Nvidia reports that AWS is adopting Omniverse, Cosmos, Isaac and Jetson to power its warehouse robots, it is describing an operator that intends to treat its fulfillment centers as a training environment as well as a logistics network.[33]
Daniela Rus and her MIT colleague Russ Tedrake made the academic case for this view in a 2025 debate at CSAIL, arguing that data-driven approaches are critical to unlocking robots’ ability to function reliably in the real world and that adaptation in messy, human-centered environments comes only from experience. Rus’s own laboratory equips volunteers with sensors while they chop vegetables, pour liquids and assemble meals, then trains models that reproduce those tasks on robots and recover when ingredients slip or tools misalign.[60]
“Robots need experience to adapt, and that comes from data.”
— Daniela Rus, Director, MIT CSAIL [60]
If experience is the raw material, then every deployed machine is a mine, and the facility that hosts it is a mining district.
4.2 Warehouses Are Particularly Powerful Learning Environments
Warehouses occupy an unusually favorable position in physical AI because they combine structure with variation. They are structured enough to permit automation: aisles are regular, shelving is standardized, tasks repeat. But they contain enough variation to challenge intelligent systems: packages differ in size, weight, rigidity and packaging; people move unpredictably; inventory changes daily; objects deform; pallets arrive damaged; aisles become obstructed; lighting varies across the day; schedules change; equipment fails. A warehouse therefore offers something valuable that neither a laboratory nor an open street provides in the same proportion: repeated tasks under continuously changing conditions, which is precisely the regime in which a learning system improves fastest.
Nvidia’s work with industrial partners illustrates how warehouses can be reconstructed digitally and used to test robotic fleets before deployment. Its Omniverse-based Mega blueprint allows enterprises to model facilities, machines and fleet behavior inside a physically accurate simulation, validate the entire deployment, and only then commit to the physical floor.[36] This creates a closed loop that can be stated compactly:
Warehouse → Digital Twin → Simulation → Robot Policy → Warehouse → New Observations → Updated Digital Twin
Each pass through the loop tightens the correspondence between the simulated facility and the real one, and each tightening makes the next round of simulated training more transferable.
4.3 The Million-Robot Threshold Changes AI Economics
In September 2026, Amazon announced plans to invest more than $100 million in a 585,000-square-foot advanced-manufacturing facility in Greenwood, Indiana, its fourth robot-manufacturing site, following an August announcement of a similar hub in Austin, Texas. The company said it has manufactured more than one million robots in the United States, that more than one million robots now operate in its fulfillment centers handling stowing, picking, sorting and intra-facility transport, and that those machines are deployed across more than 300 facilities worldwide.[53][54]
Regardless of the specific robotics architectures involved—many of Amazon’s machines are mobile drive units rather than learning manipulators—scale at that magnitude illustrates why deployment footprint matters. One million machines potentially represent one million physical endpoints. If each machine produces information about operational conditions, the fleet becomes a distributed observational network of a size no laboratory could assemble, and the decision to double domestic robot-manufacturing capacity suggests that the operator expects the network to keep growing.
This suggests an important distinction between selling robots and operating a learning fleet. The first produces hardware revenue at the moment of sale. The second can produce continuously improving intelligence for as long as the fleet runs, and that intelligence can be deployed to new machines at negligible marginal cost. The International Federation of Robotics counts 4.66 million industrial robots in operation globally, and the four manufacturers responsible for more than two million of them have now adopted a common simulation and edge-inference platform.[40][35] The installed base that could, in principle, be converted from deterministic machines into learning endpoints is therefore measured in the millions, and the conversion has begun.
4.4 Autonomous Fleets Become Rolling Laboratories
Transportation extends the same logic across geography. Waymo’s autonomous vehicles collect real-world driving experience across ten U.S. metropolitan areas while its world model generates additional scenarios, and the company’s 200 million fully autonomous miles constitute the largest grounded dataset of its kind.[49][50] Uber has taken a different route to the same destination: rather than building its own driver, it is deploying up to 500 sensor-equipped, human-driven vehicles in the specific cities and driving conditions its partners need—crowded entertainment districts, difficult pickup and drop-off points—and its AV Labs leadership estimates that 500 cars operating for six months to a year can produce enough diverse, searchable data to prepare an autonomy developer for deployment in a new market.[55]
This suggests that mobility platforms possess an asset beyond riders and drivers: coverage. A fleet continuously encounters construction, traffic, weather, pedestrians, road geometry and unusual events, and that geographic diversity can improve spatial models in ways that no single test track can. Mobility companies may therefore increasingly compete not just on transportation economics but on observational reach, and Uber’s decision to sell that reach to more than two dozen autonomy partners rather than to hoard it is itself a bet on how the market for physical experience will be structured.[57]
4.5 Simulation Creates an Experience Multiplier
Real-world fleets cannot experience every important scenario frequently enough. Some events are simply too rare. Others are too dangerous. Others are geographically concentrated. Simulation fills the gap, and the Waymo World Model is the most fully documented example of how.
Consider an autonomous vehicle encountering a mattress falling from a truck on a rainy freeway. The real incident might occur once in the life of a fleet. A world model can reconstruct it and generate variants: different speeds, different lighting, different road curvature, different vehicle positions, different weather, different object trajectories, multiple obstacles. Waymo describes exactly this capability—beginning a simulation from a real recorded event and then continuing it with generated camera and lidar data along counterfactual paths—and notes that purely reconstructive methods such as Gaussian splatting break down when the simulated route diverges too far from the original recording, whereas a fully learned world model maintains realism and consistency.[49] One rare observation becomes an entire curriculum.
This is why simulation should not be understood merely as a cost-saving substitute for reality. It can become an experience multiplier, and the size of the multiplier—how many useful simulated variations one real observation can seed—is one of the most important and least publicly measured quantities in physical AI.
4.6 Failure Data May Be More Valuable Than Success Data
Physical-AI systems could also invert the traditional corporate instinct to hide failures. For machine learning, failures are often unusually informative. A successful grasp confirms what the system already knows. A failed grasp reveals a boundary. A robot becoming confused in an unfamiliar environment identifies missing coverage. A human intervention indicates uncertainty. A near collision exposes a weakness. A failed simulation provides a counterexample that sharpens the model’s understanding of where its predictions diverge from reality.
Nvidia’s Physical AI Data Factory Blueprint, with its emphasis on assessing as well as gathering real and simulated data, and World Labs’ claim that its real-to-sim engine can predict which policies will fail in reality before they are deployed, are both attempts to industrialize the extraction of value from failure.[36][23] Companies may therefore construct enormous internal systems dedicated to identifying, clustering and replaying failures, and the economics of spatial intelligence will favor organizations that learn quickly from what goes wrong rather than those that merely accumulate records of what went right.
4.7 Teleoperation as a Bridge Technology
Fully autonomous robots will not instantly master every environment, and teleoperation therefore deserves more attention than its temporary appearance suggests. A remote human can intervene when a robot encounters an unfamiliar task. Operationally, this preserves service and keeps the robot economically useful before it is fully capable. But the intervention also generates a demonstration. The robot failed. A human showed it what to do. The session becomes training data. The next model may solve the problem autonomously.
Ken Goldberg has pointed to precisely this incremental pattern—robots that work well enough to be deployed and thereby generate their own data over time—as the realistic path across the data gap, citing Waymo’s vehicles and warehouse robots that learn through continued use.[15] Sergey Levine of Berkeley and Physical Intelligence has described the same phenomenon as on-the-job learning, in which deployed policies improve from the experience of doing real work rather than from laboratory collection alone.[61] Teleoperation thus serves simultaneously as a fallback mechanism, a safety system, a labor model, a data-collection system and a curriculum generator. It is not simply a transitional crutch. It may remain part of the infrastructure through which spatial models continuously learn, and the people who perform it are, whether or not their contracts say so, among the most important teachers the machines will ever have.
4.8 Physical AI Creates a New Flywheel
The resulting industrial loop can be stated simply:
More deployments → more environments → more experience → better models → better machines → more deployments
The important competitive question becomes whether that flywheel is open or closed. Open ecosystems may distribute improvements broadly: Physical Intelligence’s collaborative approach to data and Nvidia’s decision to release GR00T and Cosmos as open models are bets that a shared foundation grows the market faster than a proprietary one.[48][35] Closed systems may concentrate learning inside a small number of companies whose fleets are large enough to be self-sufficient. Customers may insist that their data remain isolated. Governments may impose restrictions in sensitive sectors. Employees may seek protections. Industrial companies may negotiate compensation or reciprocal model access. These governance choices, examined in Section 3, can shape the market structure of physical AI as profoundly as semiconductor supply shapes today’s generative-AI industry, and they will determine whether the 2027–2030 economy described in the next section is one of many learning fleets or a few.

Section 5: 2027–2030: Markets for Physical Experience
The preceding sections have described a constraint, a stack, an ownership problem and a set of learning environments. This section asks what economic structures emerge when all four mature together, and it does so with the caution appropriate to forecasting. The claims here are not predictions that particular companies will succeed but arguments about the shape of the market that the current trajectory implies: that grounded experience will be bought and sold as a distinct category of input, that intermediaries will emerge to broker it, that its price will vary enormously with rarity and quality, that simulation providers will become producers of synthetic experience, that the reality gap will become a competitive metric, that compute demand will expand rather than contract, that Layer 5 will begin to shape Layer 2, and that the geography of AI infrastructure will become more distributed. Each of these is already visible in embryonic form in the announcements of 2026, and each has implications for the Five-Layer AI Economy that the paper’s final sections draw together.
5.1 Experience Becomes Something Companies Purchase
Today’s AI economy already has functioning markets for GPUs, cloud compute, models, APIs, datasets, annotation, datacenter capacity and electricity. Spatial intelligence could create another market category: licensed physical experience.
A robotics developer might purchase thousands of hours of high-quality demonstrations for a particular manufacturing task. An autonomous-vehicle developer might license difficult driving scenarios from a fleet operator that has encountered them—which is precisely the transaction Uber AV Labs now offers to Nvidia, Wayve and others.[55] An agricultural robotics company might purchase sensor recordings from farms across climates. A warehouse-software company might license anonymized robotic trajectories. A model developer might purchase simulated environments representing rare industrial conditions. Figure’s Index, which pays individuals per minute of recorded activity, is already a functioning retail market for physical experience at the smallest scale, and its $1 billion data-and-compute budget suggests that the company expects the wholesale market to be large.[10]
The market would not necessarily trade raw video. Higher-value products might include curated trajectories, reconstructed environments, synchronized sensor streams, annotated failures, digital twins or synthetic variations derived from licensed physical observations. The value chain runs from observation through enrichment to environment, and margin will likely accrue to whoever performs the enrichment.
5.2 The Experience Exchange
By 2027–2030, specialized intermediaries could emerge between companies possessing environments and companies developing models. Consider an industrial-data provider with agreements covering hundreds of factories. Instead of selling ordinary datasets, it might offer robot demonstrations, simulation-ready facility models, task-specific manipulation traces, rare failure scenarios, validated sensor packages, benchmark environments and domain-specific evaluation suites.
A buyer may not need ownership of the underlying factory recordings. It may license the right to train particular models under controlled conditions, perhaps on the provider’s infrastructure, with the raw data never leaving a secured environment. This resembles data licensing but is economically closer to leasing access to experience, and it maps naturally onto the contractual questions raised in Section 3: a facility owner that has negotiated on-premises retention and exclusion of sensitive zones is exactly the kind of counterparty that could sell controlled training access without surrendering its data. Uber’s Autonomous Solutions division, which bundles training data, mapping, fleet financing and regulatory services for autonomy developers, is an early prototype of such an intermediary in the mobility sector.[57]
5.3 Physical Experience Will Have Different Prices
Not all observations are equally valuable, and a functioning market will price them accordingly. Common experiences may become inexpensive. Rare experiences may command premiums. A basic warehouse pick may become abundant as fleets scale. A delicate aerospace-maintenance procedure may remain scarce because the environments in which it can be recorded are few and tightly controlled. Driving on an ordinary highway may be abundant. A near-accident under unusual weather conditions may be extremely valuable precisely because no operator can safely produce it on demand.
| Pricing Factor | Why It Raises or Lowers the Value of an Observation |
| Rarity | Long-tail events cannot be manufactured in reality; each real instance seeds many simulated variants |
| Task difficulty | Dexterous, contact-rich or multi-step tasks are harder to learn and scarcer in existing data |
| Environment diversity | Coverage across layouts, lighting, weather and clutter improves generalization |
| Sensor quality and synchronization | Force, tactile, depth and precisely aligned streams carry information cameras alone cannot |
| Annotation quality | Success/failure labels, intervention markers and task segmentation multiply usefulness |
| Commercial relevance | Experience in high-value industrial tasks commands more than generic household activity |
| Safety sensitivity | Near-misses and hazardous scenarios are both more valuable and more legally constrained |
| Legal permissions | Clean provenance and consent raise value; ambiguous rights impose discounts |
| Geographic uniqueness | Coverage of a city or region a developer intends to enter is worth more than coverage of one it already serves |
| Transferability across embodiments | Data that improves many robot forms is worth more than data tied to one |
Table 4. Factors likely to determine the price of physical experience in a 2027–2030 market. One trillion low-value observations are not necessarily worth more than one million that cover the important parts of the physical state space.
This would create a different economic structure from web-scale text collection, in which the marginal token was nearly free and value came overwhelmingly from aggregation. In a market for experience, value will come from coverage of the state space, and the most valuable sellers will be those who can reach the parts of it that others cannot.
5.4 Simulation Companies Become Experience Producers
A second market develops around synthetic experience. If simulation fidelity becomes sufficiently high, developers can manufacture scenarios on demand. A customer could request ten million warehouse interactions, half a million forklift near-misses, a hundred thousand household kitchens, fifty thousand construction-site weather combinations or millions of randomized manipulation tasks, and receive them as generated environments rather than as recordings.
The simulation provider becomes analogous to an energy producer, except its output is artificial experience measured in scenario-hours rather than megawatt-hours. This could become a major role for Nvidia’s Omniverse and Cosmos, for World Labs inside AMD, for Google DeepMind’s Genie and for specialized simulation startups. World Labs’ real-to-sim-to-real work already illustrates the economic logic: one physical task can be transformed into many controllable simulated worlds, reducing the need to conduct every experiment on expensive hardware, and an April 2026 academic study using Marble to generate high-fidelity simulations with diverse “digital cousins” from real-world panoramas reported improved policy robustness and strong correlation between simulated and real performance.[23][62] Google’s connection of Genie to Street View has been described as turning the world model into a training-data factory at near-zero marginal cost once the underlying imagery exists.[44] The IMF’s own scenario planning for AI has noted the possibility that, with rapid diffusion of robotics, robots produce more robots; the analogous observation in the data economy is that simulators produce more experience, and experience produces better simulators.[63]
5.5 Model Developers Will Compete on the Reality Gap
Synthetic experience has a fundamental weakness: the simulator can be wrong. A virtual object may behave slightly differently from the real material. Friction may be inaccurate. Lighting may be too clean. Humans may behave unrealistically. A deformable package may not respond correctly. Tiny modeling errors can accumulate until a robot trained perfectly in simulation performs poorly in reality. This is the classic sim-to-real problem, and it has occupied academic robotics for a decade. Spatial intelligence does not eliminate it. It makes closing the reality gap more valuable, because the larger the share of training that occurs in simulation, the more a given gap costs.
Companies that continuously compare simulation against large quantities of real-world experience can recalibrate their models, and the ability to do so becomes a competitive metric. Real data grounds simulation. Simulation scales real data. The strongest systems may therefore depend on both, and the companies best positioned are those that own or can access both sides of the exchange rate: Waymo with 200 million real miles and a Genie-based simulator; Nvidia with Cosmos and Isaac on one side and Jetson controllers inside two million industrial arms on the other; World Labs with a generative engine now attached to a chipmaker that intends to put its models in machines.[49][35][1] Brooks’s warning that vision-only simulation omits the tactile signal on which dexterity depends is, in this framing, an argument that the reality gap for manipulation will remain wide until simulators and sensors capture touch, and that whoever narrows it first will hold a durable advantage.[16]
5.6 Compute Demand Expands Rather Than Disappears
It might appear that better simulation reduces physical costs, and it does. But it can dramatically increase computational demand. A single robot can perform only one physical trajectory at a time. Millions of simulated agents can potentially operate simultaneously. Each simulated environment consumes compute for graphics, physics, world modeling and policy evaluation. Training models on those experiences requires additional accelerators. Evaluation requires still more.
This helps explain why Figure simultaneously describes itself as constrained by data and by compute, and why it followed its data announcement with a compute announcement of a different order of magnitude. On September 3, 2026, the company announced a partnership with Nscale to deploy the Nvidia Vera Rubin platform across up to 100,000 GPUs beginning in the second half of 2027 in Barstow, Texas, with an initial compute commitment of $3.5 billion and intent to scale beyond $6 billion; Nscale took an equity stake in Figure as part of the arrangement.[12][64] Figure’s founder connected the infrastructure requirement directly to training future generations of physical intelligence:
Figure is entering a phase where it is “largely bound by data and compute needed to train Helix.”
— Brett Adcock, Founder and CEO, Figure AI [12]
Jensen Huang described the arrangement as the activation of a robotics flywheel—training on Vera Rubin through Nscale’s cloud, validation in Isaac Sim, deployment on Nvidia GPUs inside Figure’s robots—and placed it in the context of a new industry:
“Humanoid robots extend physical AI into the world designed for people,” opening a major new industry.
— Jensen Huang, Founder and CEO, NVIDIA [65]
One financial analysis noted that Figure, which had raised roughly $1.9 billion in its lifetime, had committed nearly twice that amount to rent computers, and that such commitments represent demand for accelerators in 2027 and 2028 that hyperscaler capital-expenditure figures do not capture.[64] Physical AI therefore does not replace the datacenter economy. It extends it into the world, and it adds a new class of customer—the robot developer—to the list of those competing for accelerator supply.
5.7 Layer 2 Starts Designing for Layer 5
This is where the AMD–World Labs transaction becomes especially important within the Five-Layer AI Economy. Layer 2 companies historically sold computation upward, to datacenters that ran whatever models their customers chose. Spatial intelligence creates stronger feedback from Layer 5. Robots reveal what models need. Models reveal what compute is inefficient. Simulation reveals memory and bandwidth requirements. Edge deployment reveals latency and power constraints. Those requirements move backward toward semiconductor design, and Layer 5 begins influencing Layer 2.
This can produce a new innovation loop:
Robot requirement → model architecture → compute requirement → chip architecture → lower-cost deployment → more robots
AMD’s stated intention that World Labs’ research would help guide its technology roadmaps across hardware, software and systems is a description of exactly this loop being deliberately shortened.[1] Nvidia’s decision to embed Jetson modules in the controllers of the four largest industrial-robot manufacturers is the same loop operating from the other end: the chip company places its silicon at the point of action so that the requirements of action flow directly back to it.[35] In both cases, the semiconductor firm is no longer waiting to learn what the next workload will be. It is buying or building a seat at the table where that workload is defined.
5.8 Spatial Intelligence Creates New Infrastructure Geography
Physical AI may also alter where AI computation happens. Large training runs remain concentrated in datacenters. Simulation can operate in AI factories. But deployment occurs at the edge. A robot cannot always wait for a distant datacenter to decide whether to stop before hitting someone. Vehicles cannot depend entirely on wide-area connectivity. Industrial equipment may operate behind restricted networks for security reasons that no cloud provider can waive.
Daniela Rus has been the most consistent academic voice on this point, arguing that the future of AI will be hybrid—local first, cloud when necessary—and that small, efficient models running on phones, cars, robots and factory equipment represent a different relationship between people and AI as well as a more sustainable one:
After automation, “intelligence will live in your pocket, not a data center.”
— Daniela Rus, Director, MIT CSAIL [66]
Spatial intelligence therefore increases demand for a distributed computing architecture: central training plus regional inference plus on-premises systems plus edge intelligence. Nvidia’s creation of an Edge Computing reporting category—$7.2 billion in the most recent quarter, covering robotics and automotive among other devices—is an early financial acknowledgment of that geography, and Rus’s own research program on compact state-space models for physical AI is the academic counterpart.[30][18] This creates new opportunities for GPUs, CPUs, embedded processors, networking equipment and specialized accelerators, and it makes the Five-Layer AI Economy more geographically distributed. Intelligence leaves the datacenter without becoming independent of it.
5.9 Spatial Intelligence Could Create an Industrial Learning Curve
Traditional manufacturing learning curves reduce costs as cumulative production increases; the more units a factory builds, the cheaper each becomes. Physical AI could introduce a second learning curve layered on the first. As cumulative deployments increase, model capability may improve, and improved capability reduces the cost of the next deployment in ways that have nothing to do with manufacturing.
This means scale can simultaneously reduce hardware cost through manufacturing, software cost through reuse, operational cost through autonomy, error rates through learning, training costs through simulation and deployment time through generalized models. If that happens, robotics could eventually experience a compounding effect similar to earlier software markets, but grounded in machines, and the physical economy would begin acquiring some of the learning dynamics of the digital economy. The World Bank’s 2026 World Development Report offers a sobering counterpoint to any assumption that these dynamics will be evenly distributed: it finds that jobs in high-income countries are more than three times as exposed to automation by generative AI as those in low- and middle-income countries, where 4.5 percent of existing jobs are at risk against 14.2 percent in high-income economies, and it argues that developing countries can benefit from AI without building large models or datacenters of their own.[67]
“AI has thrown developing economies a lifeline, and they should seize it.”
— Indermit Gill, Senior Vice President and Chief Economist, World Bank Group [67]
Whether physical AI follows the same pattern—concentrating displacement where robots are affordable and labor is expensive—is an open question, but the IMF’s scenario analysis explicitly flags requirements for physical human presence, energy and grid capacity, and institutional coordination as bottlenecks that could limit economy-wide productivity gains even as the underlying capabilities of robots become very powerful.[63]
5.10 What 2030 Might Look Like
By 2030, the distinction between an AI company and an industrial company could become increasingly difficult to maintain. A major manufacturer might operate proprietary world models trained on its facilities. A logistics company might maintain models of millions of warehouses and routes. An automotive company might continuously train driving and manufacturing systems from the same fleet. A construction company might accumulate enormous libraries of machine trajectories. Robot providers might operate fleets that learn collectively. Simulation companies might generate billions of artificial task variations each day. Hyperscalers might sell physical-AI training environments alongside GPUs, as Nvidia’s AWS partnership already suggests. Chipmakers might optimize architectures using direct knowledge of world-model workloads, as AMD’s acquisition is designed to permit.
The common asset underneath these systems will be the ability to transform experience into reusable intelligence. That is the economic promise of spatial intelligence, and the next section attempts to distill what the evidence of 2025 and 2026 has taught us about it.

Section 6: What Have We Learned? Seven Pillars
The evidence assembled in this paper spans academic robotics, corporate earnings, acquisition announcements, product launches and the reports of international institutions, and it was gathered across a period—roughly the autumn of 2025 through September 2026—in which the industry’s center of gravity visibly shifted. This section distills that evidence into seven pillars. Each pillar is a claim about how the economics of artificial intelligence change when models learn from worlds rather than words, and each is stated with enough elaboration to be argued with rather than merely asserted. Taken together they describe an economy in which experience, simulation, deployment, ownership, embodiment, edge compute and a circular five-layer structure replace the linear token-to-output logic of the language-model era.
Pillar 1 — Spatial Intelligence Changes the Scarce Input of Artificial Intelligence
The large-language-model revolution taught the industry to think of compute as scarcity against a background of extraordinarily abundant digital information. Physical AI changes that equation. Compute remains scarce, as Nvidia’s guidance of roughly 70 percent growth in fiscal 2028 and Figure’s multi-billion-dollar compute commitment both attest.[32][12] Energy remains scarce. Advanced chips remain scarce. But high-quality grounded experience can also become scarce, and in some domains it is the binding constraint.
The world contains effectively unlimited events, yet relatively few have been captured with the permissions, sensors, labels, trajectories and contextual information required to train machines. Goldberg’s 100,000-year data gap quantifies the shortfall relative to language; Brooks’s insistence on touch explains why the shortfall is not merely one of quantity but of modality; Rus’s observation that physical AI cannot tolerate hallucination explains why the shortfall cannot be papered over with lower-quality data.[14][16][18] That distinction explains why companies are investing in teleoperation, simulation, human demonstrations, autonomous fleets and large-scale physical-data collection. The strategic resource is not simply more data. It is experience that teaches an artificial system how the world changes when actions occur.
Pillar 2 — Simulation Becomes Infrastructure
Simulation should no longer be treated merely as an engineering tool operating somewhere beside the AI stack. For physical AI, simulation becomes infrastructure in the same sense that datacenters are infrastructure. Datacenters manufacture computation. World models and simulation systems manufacture experience. One real interaction can generate many simulated variations. Rare failures can be replayed repeatedly. Dangerous situations can be explored safely. Robots can accumulate enormous virtual curricula before entering a workplace.
Nvidia’s Isaac and Omniverse systems and its Physical AI Data Factory Blueprint, World Labs’ real-to-sim-to-real engine, Google’s Genie with Street View grounding and Waymo’s Genie-based world model all point toward this convergence, and the adoption of a common simulation platform by the four largest industrial-robot manufacturers marks the moment at which simulation became a shared industrial utility rather than a proprietary research tool.[35][23][42][49] For the Five-Layer AI Economy, simulation therefore deserves to be understood as part of the production system through which Layer 4 intelligence becomes reliable enough for Layer 5 action—and, because it consumes accelerators at scale, as a new source of demand for Layers 1 through 3.
Pillar 3 — Every Physical Deployment Can Become a Learning Node
Robots are not only endpoints for intelligence. They can become sources of intelligence. A deployed machine encounters circumstances unavailable inside a laboratory. It experiences variation. It discovers failures. Humans intervene. Sensors record outcomes. Those observations can potentially return to model development, and the operators with the largest fleets—Amazon with more than a million robots, Waymo with 200 million autonomous miles—possess broader experience coverage than any laboratory could assemble.[53][50]
This creates a potentially powerful deployment-learning flywheel: deployment produces experience, experience produces training, training produces capability, capability produces additional deployment. The economics of physical AI may therefore reward not simply the company with the best model today but the company with the strongest mechanism for learning from tomorrow’s deployments, and Goldberg’s prescription—build machines useful enough to be deployed so that they generate their own data—turns out to describe not only the scientific path across the data gap but the commercial one.[15]
Pillar 4 — Ownership of Physical Experience Becomes a Strategic Contractual Question
As robot-generated observations become more useful for training, customers, manufacturers, workers and model providers will increasingly have to determine who may use them. Factories are not public websites. Homes are not public datasets. Industrial processes may contain trade secrets. Workers create tacit knowledge. Healthcare and consumer environments contain sensitive information that no amount of anonymization fully neutralizes.
The expansion of spatial intelligence therefore creates a governance requirement alongside a technical opportunity. Successful physical-AI ecosystems will need architectures and contracts capable of balancing model improvement against confidentiality, privacy, safety and customer control—on-premises retention, federated training, zone exclusion, opt-out rights and compensation mechanisms among them. The language-model industry negotiated these questions after the fact; the physical-AI industry has the opportunity, and the obligation, to negotiate them before the data exists. The companies that solve this well may obtain something more valuable than an isolated dataset: permission to continue learning.
Pillar 5 — Humans Become Teachers, and the Terms of Teaching Matter
A pillar that the original framing of this paper underweighted, and that the events of 2026 have made impossible to ignore, is the emergence of humans as deliberate producers of training experience for machines. Figure pays more than 44,000 weekly active contributors to record household tasks.[10] Uber pays fleet partners and equips human drivers to gather the situations autonomy developers cannot find on their own.[55] Teleoperators in warehouses and factories supply the corrective demonstrations from which the next model learns. Rus’s laboratory instruments volunteers in a kitchen.[60]
This is, in the near term, complementary work of exactly the kind Brynjolfsson argues raises wages, and in the longer term it is designed to make itself unnecessary, which is the pattern Acemoglu and Restrepo warn can erode the labor share if it is the only pattern the market rewards.[59][58] The paper does not resolve that tension. It records that the terms on which humans teach machines—per minute, per demonstration, per intervention, with or without recognition of the tacit expertise being transferred—are being set now, and that they deserve the attention of policymakers, labor representatives and the companies themselves before the pipelines harden.
Pillar 6 — Intelligence Moves to the Edge Without Leaving the Datacenter
Spatial intelligence changes not only what is computed but where. Training remains centralized; simulation runs in AI factories; but action happens on the machine, under latency and power constraints that no wide-area network can satisfy. Nvidia’s Edge Computing revenue, its Jetson modules inside the controllers of ABB, FANUC, KUKA and Yaskawa, and Rus’s argument for local-first hybrid intelligence all describe the same architectural shift.[30][35][66]
The consequence for the Five-Layer AI Economy is that Layer 5 is no longer a thin application layer sitting atop centralized infrastructure. It acquires its own compute, its own inference and its own sensors, and it becomes a source of requirements that flow back to Layer 2. AMD’s acquisition of a model laboratory and Nvidia’s placement of silicon at the point of action are both attempts to sit where those requirements originate.
Pillar 7 — The Five-Layer AI Economy Is Becoming a Closed Learning Loop
The Five-Layer AI Economy originally provided a useful way to understand the physical hierarchy underneath artificial intelligence: energy powers chips, which populate datacenters, which train and operate models, which enable applications and agents. Spatial intelligence adds an important return path.
| Layer | Original Role (LLM Era) | Added Role (Spatial-Intelligence Era) |
| Layer 1 — Energy | Powers training and inference | Also powers simulation at scale and edge inference in millions of machines |
| Layer 2 — Chips | Accelerate token generation | Accelerate physics, rendering, world models and robot policies; designed with direct knowledge of Layer 5 workloads |
| Layer 3 — Datacenters | Host training and serving | Become experience factories manufacturing simulated worlds; joined by regional and on-premises inference |
| Layer 4 — Models | Learn from text and images | Learn from video, 3D space, trajectories, simulation and real-world outcomes |
| Layer 5 — Agents | Produce digital outputs | Acquire bodies; act in the world; generate experience that returns to Layer 4 |
Table 5. How spatial intelligence changes the role of each layer in the Five-Layer AI Economy.
When Layer 5 becomes physical, agents encounter the world. Those encounters generate observations. Observations improve Layer 4 models. New model workloads reshape Layer 2 chips. New chips alter Layer 3 datacenters. Expanding datacenters create additional Layer 1 electricity demand. The framework therefore becomes circular rather than merely vertical:
Energy → Chips → Datacenters → Models → Physical Agents → Experience → Models → Chips → Datacenters → Energy
That feedback loop may become one of the defining industrial structures of artificial intelligence between 2027 and 2030, and the September 2026 decision by a semiconductor company to acquire a spatial-intelligence laboratory is the clearest evidence yet that the participants in the loop understand its shape.

Conclusion: Why Spatial Intelligence Models Define the Next Expansion of AI
The first great wave of modern generative artificial intelligence was built upon something civilization had already produced in enormous quantities: digital information. Humanity had spent decades writing the web before machines learned to read it. It had stored books before models summarized them. It had accumulated code before agents began programming. It had created photographs before multimodal models interpreted them. The transformer arrived to find its corpus waiting, and the result was the fastest expansion of machine capability in history.
The physical world presents a different challenge, and the evidence of the past twelve months suggests that the industry now understands it. The knowledge required by a robot is distributed across billions of objects, movements, environments and interactions that have never been systematically captured. Much of human physical expertise remains tacit. It exists in hands, in eyes, in muscle memory, in judgment, in timing, in the accumulated experience of people who could not write down what they know even if asked. Goldberg’s arithmetic, Brooks’s neuroscience and Rus’s safety constraint all describe the same gap from different directions, and the corporate response—a billion dollars for household videos, three and a half billion for the GPUs to learn from them, eight billion for a laboratory that generates worlds—describes the scale of the effort now being made to close it.[14][16][18][10][12][1]
The next frontier of artificial intelligence is therefore not merely making models larger. It is giving artificial systems sufficiently rich representations of the worlds in which actions occur. That requires cameras, sensors and video. It requires robot trajectories and human demonstrations. It requires teleoperation. It requires digital twins. It requires physically grounded simulation. It requires world models. It requires datacenters capable of manufacturing enormous quantities of artificial experience. And ultimately it requires machines entering reality, discovering what their models still misunderstand, and returning that knowledge to the systems that trained them.
The September 28, 2026 AMD–World Labs announcement is important precisely because it illustrates how deeply this transition may reach. AMD is a semiconductor company whose Data Center revenue more than doubled in its most recent quarter.[27] World Labs is a spatial-intelligence company whose founder has argued for years that language is a lossy channel for describing a three-dimensional world.[13] Their combination links the architecture of computation with emerging models intended to understand space and physical environments, and it does so on the explicit premise that the chip company must understand the model in order to design the chip.[1]
Nvidia is approaching the same frontier by connecting GPUs, Omniverse, Isaac, Cosmos, GR00T and Jetson into a physical-AI platform that now runs inside the controllers of the world’s largest industrial-robot manufacturers and, per its most recent earnings call, inside Amazon’s warehouse fleet.[35][33] Google DeepMind is building interactive world models grounded in two decades of Street View imagery and embodied-reasoning models that plan in the physical world.[42][45] Waymo is combining 200 million real miles with billions of simulated ones.[49] Yann LeCun has raised a billion dollars on the thesis that world models, not language models, are the path forward.[46] Figure is building physical datasets because, in its own formulation, the required robot-training information is not waiting on the internet.[10] Uber is selling the roads it sees to the companies that want to drive them.[55] Factories and warehouses are becoming environments in which intelligent machines continuously learn.
The common thread is spatial intelligence. That is why Spatial Intelligence Models fits this paper better than a title centered on a particular commercial resource, financial metaphor or individual model architecture. The phrase describes the capability toward which all of these developments converge. A world model is one technical mechanism. A digital twin is an environment. A robot policy is a controller. A sensor stream is an observation. A simulation engine produces experience. But spatial intelligence is the larger objective: enabling an artificial system to understand where things are, how they relate, how they move, what actions are possible, what consequences those actions may produce, and how to operate when reality differs from anything previously encountered.
That distinction becomes especially important as artificial intelligence moves beyond screens. A chatbot can fail in language, and the failure is an inconvenience. A robot can fail in space, and the failure has mass and momentum. When an AI system controls a vehicle, manipulates an industrial machine, works beside employees, enters a home or operates critical equipment, understanding the physical world is no longer an optional extension of intelligence. It becomes a prerequisite for action, and the tolerance for error that the language-model era learned to live with will not transfer.
The strategic competition of the coming years may therefore be measured not only by which organization owns the most GPUs, constructs the largest datacenters or trains the biggest models. It may also depend on which organizations can acquire, simulate, govern and learn from the richest representations of physical reality—and on whether the humans who supply that experience, the facilities that host it and the societies that absorb its consequences are treated as participants in the loop rather than as inputs to it.
The language-model era taught machines about what humanity has written. The era of Spatial Intelligence Models will increasingly teach machines about where humanity lives and works, how objects behave, how environments change, and what happens when intelligence acts upon the world. That is why the title belongs at the center of this paper. And it is why spatial intelligence belongs at the center of the next chapter of the Five-Layer AI Economy.

Footnotes and Endnotes:
[1] AMD Newsroom. “AMD to Acquire World Labs to Advance the Future of AI Compute.” Advanced Micro Devices press release, September 28, 2026. https://newsroom.amd.com/news/amd-acquire-world-labs/
[2] Bloomberg News. “AMD to Buy Fei-Fei Li’s World Labs Startup for $8.2 Billion.” Bloomberg, September 28, 2026. https://www.bloomberg.com/news/articles/2026-09-28/amd-to-buy-fei-fei-li-s-world-labs-ai-startup-for-8-2-billion
[3] CNBC. “AMD acquiring Fei-Fei Li’s World Labs AI firm in deal worth $8.2 billion.” CNBC, September 28, 2026. https://www.cnbc.com/2026/09/28/amd-fei-fei-li-world-labs.html
[4] Techstrong.ai. “AMD to Acquire Fei-Fei Li’s World Labs in $8.2 Billion AI Deal.” Techstrong.ai, September 28, 2026. https://techstrong.ai/articles/amd-to-acquire-fei-fei-lis-world-labs-in-8-2-billion-ai-deal/
[5] Stocktwits via TradingView. “AMD Bolsters AI Vision With $8.2B Acquisition Of Fei-Fei Li’s World Labs, Stock Drops 4%.” TradingView News, September 28, 2026. https://www.tradingview.com/news/stocktwits:c0b6118cc094b:0-amd-bolsters-ai-vision-with-8-2b-acquisition-of-fei-fei-li-s-world-labs-stock-drops-4/
[6] The American Bazaar (quoting Patrick Moorhead, Moor Insights & Strategy). “AMD to acquire Fei-Fei Li’s World Labs for $8.2 billion in AI push.” The American Bazaar, September 29, 2026. https://americanbazaaronline.com/2026/09/29/amd-to-acquire-fei-fei-lis-world-labs-for-8-2-billion-in-ai-push-489018/
[7] Reuters (via Yahoo Finance). “AI pioneer Fei-Fei Li’s World Labs raises $1 billion in funding.” Reuters, February 18, 2026. https://finance.yahoo.com/news/ai-pioneer-fei-fei-lis-202957884.html
[8] Fei-Fei Li (Stanford University; World Labs). “Thread introducing “From Words to Worlds”.” X (formerly Twitter), November 10, 2025. https://x.com/drfeifei/status/1987891813387292725
[9] Fei-Fei Li (Stanford University; World Labs). “From Words to Worlds: Spatial Intelligence is AI’s Next Frontier.” Substack essay, November 10, 2025. https://drfeifei.substack.com/p/from-words-to-worlds-spatial-intelligence
[10] Figure AI. “Introducing Index: Building The World’s Largest and Most Diverse Physical Dataset.” Figure newsroom, August 25, 2026. https://www.figure.ai/news/introducing-index
[11] Justin Ryan, Spatial Insiders. “Figure Launches Index to Pay People for Videos That Train Robots.” Spatial Insiders, August 25, 2026. https://spatialinsiders.com/stories/figure-index-robot-training-videos
[12] Figure AI. “Figure and Nscale Sign Strategic Partnership For Up to 100,000 GPUs on the NVIDIA Vera Rubin Platform.” Figure newsroom, September 3, 2026. https://www.figure.ai/news/figure-and-nscale-sign-strategic-partnership
[13] Latent Space (podcast with Fei-Fei Li and Justin Johnson, World Labs). “After LLMs: Spatial Intelligence and World Models.” Latent Space, February 2026. https://www.latent.space/p/after-llms-spatial-intelligence-and
[14] Ken Goldberg (UC Berkeley), interviewed by Berkeley News. “Are we truly on the verge of the humanoid robot revolution?.” Berkeley News (on Goldberg’s Science Robotics papers), August 27, 2025. https://news.berkeley.edu/2025/08/27/are-we-truly-on-the-verge-of-the-humanoid-robot-revolution/
[15] The Robot Report. “Why humanoid robots aren’t advancing as fast as AI chatbots.” The Robot Report, September 2025. https://www.therobotreport.com/why-humanoid-robots-arent-advancing-as-fast-as-ai-chatbots/
[16] Rodney Brooks (MIT, Panasonic Professor of Robotics emeritus; Robust.AI). “Why Today’s Humanoids Won’t Learn Dexterity.” rodneybrooks.com, September 26, 2025. https://rodneybrooks.com/why-todays-humanoids-wont-learn-dexterity/
[17] David Edwards, Robotics & Automation News (quoting Rodney Brooks). “Pioneering roboticist Rodney Brooks says humanoid robots matching human skills is ‘pure fantasy’.” Robotics & Automation News, October 1, 2025. https://roboticsandautomationnews.com/2025/10/01/pioneering-roboticist-rodney-brooks-says-humanoid-robots-matching-human-skills-is-pure-fantasy/95069/
[18] Daniela Rus (Director, MIT CSAIL), interviewed by Capgemini. “A conversation with Daniela Rus: When AI meets robotics.” Capgemini Research Institute, March 2026. https://www.capgemini.com/insights/research-library/a-conversation-with-daniela-rus/
[19] AlphaSignal. “Figure’s Index App Pays 44,000 People to Train Its Humanoid Robots.” AlphaSignal, August 2026. https://alphasignal.ai/news/figure-s-index-app-pays-44-000-people-to-train-its-humanoid-robots
[20] J. Ni, Z. Wang, W. Lin, A. Bar, Y. LeCun, T. Darrell, J. Malik, R. Herzig (UC Berkeley; NYU). “From Generated Human Videos to Physically Plausible Robot Trajectories.” arXiv 2512.05094, December 2025. https://arxiv.org/pdf/2512.05094
[21] Pebblous. “Figure AI Index — Buying Humanoid Training Video.” Pebblous blog, September 2026. https://blog.pebblous.ai/blog/figure-index-robot-training-video-gig/en/
[22] Seedtable. “SceniX acquired by World Labs (Jul 2026).” Seedtable, July 2026. https://seedtable.com/exits/scenix
[23] World Labs. “Building Worlds That Train Robots (Real-to-Sim-to-Real).” World Labs blog, July 2026. https://www.worldlabs.ai/blog/real-to-sim-to-real
[24] Marcel Torne, Anthony Simeonov, Zechu Li, Tao Chen, Pulkit Agrawal et al. (MIT CSAIL). “Reconciling Reality through Simulation: A Real-to-Sim-to-Real Approach for Robust Manipulation.” arXiv 2403.03949 (Robotics: Science and Systems), 2024. https://arxiv.org/pdf/2403.03949
[25] Eliza Strickland, IEEE Spectrum (interview with Fei-Fei Li). “Fei-Fei Li’s World Labs Wants to Give AI Spatial Intelligence.” IEEE Spectrum, 2024. https://spectrum.ieee.org/fei-fei-li-world-labs
[26] Dealroom. “World Labs acquires SceniX to bring generative AI into physical robotics and embodied intelligence.” Dealroom News, July 2026. https://app.dealroom.co/news/feed/world-labs-acquires-scenix-to-bring-generative-ai-into-physical-robotics-and-embodied-intelligence
[27] Advanced Micro Devices, Inc.. “AMD Reports Second Quarter 2026 Financial Results (Form 8-K exhibit).” U.S. Securities and Exchange Commission, August 4, 2026. https://www.sec.gov/Archives/edgar/data/0000002488/000000248826000121/q22026991.htm
[28] Investing.com. “AMD Q2 2026 slides: data center revenue doubles, AI partnerships expand.” Investing.com, August 4, 2026. https://www.investing.com/news/company-news/amd-q2-2026-slides-data-center-revenue-doubles-ai-partnerships-expand-93CH-4836256
[29] Advanced Micro Devices, Inc.. “AMD Financial Results Second Quarter 2026 (earnings presentation, Form 8-K exhibit).” U.S. Securities and Exchange Commission, August 4, 2026. https://www.sec.gov/Archives/edgar/data/0000002488/000000248826000121/amdq22026earningsslidesf.htm
[30] NVIDIA Newsroom. “NVIDIA Announces Financial Results for Second Quarter Fiscal 2027.” NVIDIA press release, August 26, 2026. https://nvidianews.nvidia.com/news/nvidia-announces-financial-results-for-second-quarter-fiscal-2027
[31] NVIDIA Corporation. “NVIDIA Announces Financial Results for First Quarter Fiscal 2027 (Form 8-K exhibit).” U.S. Securities and Exchange Commission, May 20, 2026. https://www.sec.gov/Archives/edgar/data/0001045810/000104581026000051/q1fy27pr.htm
[32] Anthony Lopopolo, Quartz. “Nvidia more than doubled revenue to $96 billion and expects 70% growth ahead.” Quartz, August 26, 2026. https://qz.com/nvidia-earnings-revenue-doubled-stock-fell-082626
[33] NVIDIA Investor Relations. “Q2 Fiscal 2027 Earnings Call Transcript.” NVIDIA, August 26, 2026. https://investor.nvidia.com/files/content_files/TRANSCRIPT_-NVIDIA-Corp-NVDA-US-Q2-2027-Earnings-Call-26-August-2026-5_00-PM-ET.pdf
[34] NVIDIA Corporation. “Definitive Proxy Statement (Form DEF 14A), Fiscal 2026 Business Overview.” U.S. Securities and Exchange Commission, May 12, 2026. https://www.sec.gov/Archives/edgar/data/0001045810/000104581026000036/nvda-20260512.htm
[35] NVIDIA Newsroom. “NVIDIA and Global Robotics Leaders Take Physical AI to the Real World.” NVIDIA press release (GTC 2026), March 2026. https://nvidianews.nvidia.com/news/nvidia-and-global-robotics-leaders-take-physical-ai-to-the-real-world
[36] NVIDIA. “Into the Omniverse: NVIDIA GTC Showcases Virtual Worlds Powering the Physical AI Era.” NVIDIA blog, March 2026. https://blogs.nvidia.com/blog/gtc-2026-virtual-worlds-physical-ai/
[37] NVIDIA Research. “Cosmos 3: Omnimodal World Models for Physical AI.” arXiv 2606.02800, June 2026. https://arxiv.org/pdf/2606.02800
[38] Association for Advancing Automation (A3), quoting Rev Lebaredian, NVIDIA. “NVIDIA Declares ‘Big Bang of Physical AI’ at GTC 2026.” Automate.org Industry Insights, March 2026. https://www.automate.org/ai/industry-insights/nvidia-declares-big-bang-of-physical-ai-at-gtc-2026
[39] The Decoder. “GTC 2026: Nvidia wants to swap robotics’ data problem for a compute problem.” The Decoder, March 2026. https://the-decoder.com/gtc-2026-nvidia-wants-to-swap-robotics-data-problem-for-a-compute-problem/
[40] International Federation of Robotics (IFR). “World Robotics 2025 — Industrial Robots: Executive Summary.” IFR Statistical Department / VDMA Services, September 2025. https://ifr.org/img/worldrobotics/Executive_Summary_WR_2025_Industrial_Robots.pdf
[41] International Federation of Robotics (Takayuki Ito, President). “Global Robot Demand in Factories Doubles Over 10 Years.” IFR press release via Presseportal, September 25, 2025. https://www.presseportal.de/pm/115415/6125403
[42] TechCrunch. “Google’s Genie world model can now simulate real streets with Street View.” TechCrunch, May 19, 2026. https://techcrunch.com/2026/05/19/googles-genie-world-model-can-now-simulate-real-streets-with-street-view/
[43] TechCrunch via Yahoo Tech (quoting Jack Parker-Holder, Google DeepMind). “Google’s Genie world model can now simulate real streets with Street View.” Yahoo Tech, May 19, 2026. https://tech.yahoo.com/ai/gemini/articles/google-genie-world-model-now-175139464.html
[44] The Decoder. “Google pairs its Genie world model with Street View to create explorable AI worlds based on real places.” The Decoder, May 2026. https://the-decoder.com/google-pairs-its-genie-world-model-with-street-view-to-create-explorable-ai-worlds-based-on-real-places/
[45] Google DeepMind. “Gemini Robotics ER 2: Our embodied reasoning model.” Google DeepMind model page, July 30, 2026. https://deepmind.google/models/gemini-robotics/embodied-reasoning/
[46] TechCrunch (quoting Alexandre LeBrun, CEO, AMI Labs). “Yann LeCun’s AMI Labs raises $1.03B to build world models.” TechCrunch, March 9, 2026. https://techcrunch.com/2026/03/09/yann-lecuns-ami-labs-raises-1-03-billion-to-build-world-models/
[47] Nick Patience, The Futurum Group. “Yann LeCun’s AMI Raises $1BN Seed Round — Is the World Model Era Finally Here?.” The Futurum Group, March 13, 2026. https://futurumgroup.com/insights/yann-lecuns-ami-raises-1bn-seed-round-is-the-world-model-era-finally-here/
[48] The Robot Report. “Physical Intelligence raises $600M to advance robot foundation models.” The Robot Report, November 2025. https://www.therobotreport.com/physical-intelligence-raises-600m-advance-robot-foundation-models/
[49] Waymo. “The Waymo World Model: A New Frontier For Autonomous Driving Simulation.” Waymo blog, February 2026. https://waymo.com/blog/2026/02/the-waymo-world-model-a-new-frontier-for-autonomous-driving-simulation/
[50] Waymo. “Over 200 million fully autonomous miles wrapped.” X (formerly Twitter), February 23, 2026. https://x.com/Waymo/status/2025979468620128257
[51] Stewart Burnett, Automotive World. “Waymo unveils DeepMind-powered world simulation model.” Automotive World, February 9, 2026. https://www.automotiveworld.com/news/waymo-unveils-deepmind-powered-world-simulation-model/
[52] CBS News. “Uber looks to cash in on self-driving cars — but not by driving them.” CBS News, January 2026. https://www.cbsnews.com/news/uber-self-driving-cars-autonomous-driving/
[53] The Robot Report. “Amazon to invest $100M in new Indiana manufacturing facility.” The Robot Report, September 25, 2026. https://www.therobotreport.com/amazon-to-invest-100m-in-new-indiana-manufacturing-facility/
[54] Supply Chain Digital. “Amazon Invests $100m in Indiana Robotics Manufacturing.” Supply Chain Digital, September 25, 2026. https://supplychaindigital.com/news/amazon-invests-100m-in-indiana-robotics-manufacturing
[55] Axios (quoting Danny Guo, Uber AV Labs). “Uber uses its ride-hailing fleet to train AI.” Axios, September 23, 2026. https://axios.com/2026/09/23/uber-av-ai-training-ride-hailing-fleet
[56] Gadget Review (quoting Praveen Neppalli Naga, Uber CTO). “Uber’s Fleet Of 500 Data-Collection Vehicles To Hit the Road This Year.” Gadget Review, June 2026. https://www.gadgetreview.com/ubers-fleet-of-500-data-collection-vehicles-to-hit-the-road-this-year
[57] Kirsten Korosec, TechCrunch. “Uber wants to be a Swiss Army Knife for robotaxis.” TechCrunch, February 23, 2026. https://techcrunch.com/2026/02/23/uber-autonomous-solutions-av-robotaxi-delivery-robots/
[58] Daron Acemoglu and Pascual Restrepo (MIT; Boston University). “The Wrong Kind of AI? Artificial Intelligence and the Future of Labor Demand.” IZA Discussion Paper No. 12292, 2019. https://docs.iza.org/dp12292.pdf
[59] Boston Review (Daron Acemoglu, MIT; Erik Brynjolfsson, Stanford). “AI and the Specter of Automation (reading list: “AI’s Future Doesn’t Have to Be Dystopian” and “Augmentation, Not Automation”).” Boston Review, 2025. https://www.bostonreview.net/reading-list/ai-and-the-specter-of-automation/
[60] MIT CSAIL (Daniela Rus and Russ Tedrake). “MIT roboticists debate the future of robotics.” MIT CSAIL News, August 2025. https://www.csail.mit.edu/news/mit-roboticists-debate-future-robotics
[61] Association for Advancing Automation (interview with Sergey Levine, UC Berkeley; Physical Intelligence). “On-the-Job Learning: How Physical Intelligence Is Putting Robots to Work.” Automate.org Industry Insights, 2025. https://www.automate.org/ai/industry-insights/on-the-job-learning-how-physical-intelligence-is-putting-robots-to-work
[62] Stanford/World Labs-affiliated authors. “From Seeing to Simulating: Generative High-Fidelity Simulation with Digital Cousins for Generalizable Robot Learning and Evaluation.” arXiv 2604.15805, April 2026. https://arxiv.org/pdf/2604.15805
[63] International Monetary Fund. “Global Economic and Financial Implications of Artificial Intelligence: Lessons from a Scenario Planning Exercise (IMF Note 2026/002).” IMF, April 2026. https://www.imf.org/-/media/files/publications/imf-notes/2026/english/insea2026002.pdf
[64] Yahoo Finance. “Figure AI Commits $3.5 Billion To Nscale For Up To 100,000 NVIDIA GPUs.” Yahoo Finance, September 9, 2026. https://finance.yahoo.com/technology/ai/articles/figure-ai-commits-3-5-124536120.html
[65] Compare the Cloud (quoting Jensen Huang, NVIDIA). “Nscale signs $3.5bn compute partnership with humanoid robotics firm Figure.” Compare the Cloud, September 3, 2026. https://www.comparethecloud.net/news/nscale-signs-35bn-compute-partnership-with-humanoid-robotics-firm-figure
[66] MIT CSAIL (Daniela Rus). “Daniela Rus talks about AI after automation at Every’s upcoming Thesis conference.” MIT CSAIL Substack, September 29, 2026. https://csailmit.substack.com/p/daniela-rus-talks-about-ai-after
[67] World Bank Group (Indermit Gill, Chief Economist). “AI Offers Lifeline to Developing Economies in an Era of Weak Growth — World Development Report 2026: The Promise of Artificial Intelligence.” World Bank press release, August 4, 2026. https://www.worldbank.org/en/news/press-release/2026/08/04/ai-offers-lifeline-to-developing-economies-in-an-era-of-weak-growth



