Big Tech and VCs Pour Billions into World Models to Make Physical AI a Realityc

Monday, 27/07/2026 | 05:49 GMT by FM
Disclaimer
  • The race to build AI that understands physical space is drawing record levels of funding and talent.
Big Tech is betting billions on AI world models to power the next wave of physical AI.

As talk of an "inevitable" AI bubble burst wanes, the ecosystem is witnessing a fundamental shift in focus. Investors who spent the last two years chasing chatbots, agents and vibe coding engines are now pouring billions into a different bet entirely: AI that can operate in the physical world. Venture capitalists, chipmakers, and cloud giants have funneled more than $3 billion into a single category of startup in just six months, and the pace shows no sign of slowing.

The reason is simple. There's a growing understanding that if AI is to truly change the world and achieve its full economic potential, it cannot stay confined to digital spaces. The next trillion-dollar opportunity, many believe, lies in machines that can work alongside humans in warehouses, drive trucks across the country, and inspect infrastructure without a pilot. That shift from software to physical deployment is why so many AI power players are now laser-focused on “world models.”

This is the overriding goal of AI world models, suddenly the hottest topic in AI. Rather than just predicting the next word, they're able to generate rich simulations of reality, possessing the ability to understand real-world physics, the concept of time and the consequences of actions – the exact ingredients needed to train machines that can operate safely and reliably out here in the real world, at a fraction of the cost of real-world testing.

AI that understands the world’s mechanics

As world models start making more headlines, it has become apparent that most researchers have a hard time agreeing on exactly what a “world model” is. Vincent Sitzmann, an MIT professor who leads the Scene Representation Group within its Computer Science and Artificial Intelligence Laboratory, told Ars Technica that world models differ from LLMs due to their ability to take an interaction and then “simulate what would happen next in some environment.”

Renowned AI researcher Fei-Fei Li, a co-founder of World Labs, asserts that a world model must have three specific criteria: the ability to generate worlds with perceptual, geometrical and physical consistency, be multimodal by design and output the next states of that world, based on input actions. In a way, they’re a kind of 3D sandbox for AI agents to learn about the physical world by predicting how their surroundings will change as a result of physical actions within it.

Quite simply, they’re the key to unlocking the potential of physical AI implementations, such as autonomous vehicles and robots that can work alongside humans in factories, warehouses, shops and construction sites.

These systems already exist, but scaling them up is incredibly difficult due to the massive cost and difficulty of training the underlying AI models that will provide them with the intelligence to work safely alongside real people.

To create a self-driving truck that can travel all the way from New York City to San Francisco without a human at the wheel or steering it remotely, developers need massive volumes of training data that can teach the underlying model how to react to every conceivable hazard or situation it might encounter along the way. It’s simply not possible to recreate all of these scenarios out here in the real world.

World models can solve this dilemma, giving developers a way to run high volumes of extremely realistic simulations that encompass every conceivable thing that could happen to that truck on its journey, from sand storms to blizzards to rabid dogs running onto the road.

A $3 billion-dollar gold rush

Data compiled by Dealroom reveals that venture capitalists and others have funneled more than $2.3 billion into world model developers in the first six months of 2026. It’s an astonishing amount of cash that illustrates just how valuable the market for physical AI could become.

Some of the amounts raised this year are truly staggering, all the more so because many of the startups in question remain relatively unknown out here in the real world. Decart, for instance, recently raised $300 million from Nvidia, eBay Ventures, Adobe Ventures, Toyota Ventures, Atreides Management and others in a deal that took its valuation to more than $4 billion.

A leading frontier AI research lab, Decart is the creator of Oasis 3, a world model that can render interactive, multi-camera robotic training environments in real time.

Li’s World Labs recently closed its own hefty $1 billion round led by Autodesk, AMD, Nvidia, Andreessen Horowitz and Fidelity in February, taking its valuation to $5.4 billion. The company’s flagship model is Marble, which generates persistent 3D worlds. That same month saw Runway close a $315 million round at a $5.3 billion valuation from backers including General Atlantic, Nvidia, Fidelity, AllianceBernstein, Adobe, AMD and Felicis.

Other big raises this year include Yann Le Cunn’s AMI Labs, which snagged $1.03 billion in March from Cathay Innovation, Greycroft, Hiro Capital, HV Capital and Bezos Expeditions. Its value now stands at $3.5 billion. Then there’s General Intuition, which bagged $320 million at a $2.3 billion valuation in June from investors such as Khosla, General Catalyst and Bezos Expeditions, and Odyssey, which raised $310 million from Amazon, AMD and Google’s venture arm, at a $1.45 billion valuation.

The biggest round of all saw Skild AI closing on a whopping $1.4 billion raise led by SoftBank, Nvidia and Bezos Expeditions in January, propelling its value to an incredible $14 billion.

Infrastructure bets on a sure thing

One of the most notable themes in the trending interest in world models is that the money isn’t just coming from traditional VCs. In fact, the most active investors in world models are the compute infrastructure players that will probably benefit from their growth more than anyone else – chipmakers like Nvidia and AMD, and cloud infrastructure providers like AWS (often via Bezos Expeditions).

There’s a fascinating overlap occurring, too. Consider Decart, which has secured massive financial backing from Nvidia’s VC arm while simultaneously signing on to become one of AWS’s flagship Trainium3 partners. Odyssey is a similar story, with AMD being one of its biggest backers, while running its workloads primarily on AWS and Google Cloud infrastructure.

Confusingly, many of these infrastructure players are enthusiastically building their own world models, too. Take AWS, which is leaning on its Amazon Nova AI service to develop systems that can process a combination of visual, temporal and spatial data to power the robots in its massive warehouses. Meanwhile, Nvidia has built Nvidia Cosmos, an environment for training robots and physical AI to perceive real-world environments and react inside them, and Omniverse, for developing digital twins.

There can be only one reason why these infrastructure players are effectively competing against themselves. Quite simply, they see world models as the fundamental drivers of the next era of industrial automation. By investing heavily in the most promising frontier labs while simultaneously pursuing their own ambitious projects, it's clear that they're hedging their bets.

They can't predict which lab will ultimately win this race – their own teams or one of their external investments – but they're confident someone will. And when that winner emerges, world models stand to generate billions of dollars in revenue, not just for the developer itself, but for the cloud services and silicon that power it.

Following the money, it’s clear that the future of AI won't just belong to AI agents and chatbots grabbing the headlines today, but the physical AI systems that will live and work right by our side.

As talk of an "inevitable" AI bubble burst wanes, the ecosystem is witnessing a fundamental shift in focus. Investors who spent the last two years chasing chatbots, agents and vibe coding engines are now pouring billions into a different bet entirely: AI that can operate in the physical world. Venture capitalists, chipmakers, and cloud giants have funneled more than $3 billion into a single category of startup in just six months, and the pace shows no sign of slowing.

The reason is simple. There's a growing understanding that if AI is to truly change the world and achieve its full economic potential, it cannot stay confined to digital spaces. The next trillion-dollar opportunity, many believe, lies in machines that can work alongside humans in warehouses, drive trucks across the country, and inspect infrastructure without a pilot. That shift from software to physical deployment is why so many AI power players are now laser-focused on “world models.”

This is the overriding goal of AI world models, suddenly the hottest topic in AI. Rather than just predicting the next word, they're able to generate rich simulations of reality, possessing the ability to understand real-world physics, the concept of time and the consequences of actions – the exact ingredients needed to train machines that can operate safely and reliably out here in the real world, at a fraction of the cost of real-world testing.

AI that understands the world’s mechanics

As world models start making more headlines, it has become apparent that most researchers have a hard time agreeing on exactly what a “world model” is. Vincent Sitzmann, an MIT professor who leads the Scene Representation Group within its Computer Science and Artificial Intelligence Laboratory, told Ars Technica that world models differ from LLMs due to their ability to take an interaction and then “simulate what would happen next in some environment.”

Renowned AI researcher Fei-Fei Li, a co-founder of World Labs, asserts that a world model must have three specific criteria: the ability to generate worlds with perceptual, geometrical and physical consistency, be multimodal by design and output the next states of that world, based on input actions. In a way, they’re a kind of 3D sandbox for AI agents to learn about the physical world by predicting how their surroundings will change as a result of physical actions within it.

Quite simply, they’re the key to unlocking the potential of physical AI implementations, such as autonomous vehicles and robots that can work alongside humans in factories, warehouses, shops and construction sites.

These systems already exist, but scaling them up is incredibly difficult due to the massive cost and difficulty of training the underlying AI models that will provide them with the intelligence to work safely alongside real people.

To create a self-driving truck that can travel all the way from New York City to San Francisco without a human at the wheel or steering it remotely, developers need massive volumes of training data that can teach the underlying model how to react to every conceivable hazard or situation it might encounter along the way. It’s simply not possible to recreate all of these scenarios out here in the real world.

World models can solve this dilemma, giving developers a way to run high volumes of extremely realistic simulations that encompass every conceivable thing that could happen to that truck on its journey, from sand storms to blizzards to rabid dogs running onto the road.

A $3 billion-dollar gold rush

Data compiled by Dealroom reveals that venture capitalists and others have funneled more than $2.3 billion into world model developers in the first six months of 2026. It’s an astonishing amount of cash that illustrates just how valuable the market for physical AI could become.

Some of the amounts raised this year are truly staggering, all the more so because many of the startups in question remain relatively unknown out here in the real world. Decart, for instance, recently raised $300 million from Nvidia, eBay Ventures, Adobe Ventures, Toyota Ventures, Atreides Management and others in a deal that took its valuation to more than $4 billion.

A leading frontier AI research lab, Decart is the creator of Oasis 3, a world model that can render interactive, multi-camera robotic training environments in real time.

Li’s World Labs recently closed its own hefty $1 billion round led by Autodesk, AMD, Nvidia, Andreessen Horowitz and Fidelity in February, taking its valuation to $5.4 billion. The company’s flagship model is Marble, which generates persistent 3D worlds. That same month saw Runway close a $315 million round at a $5.3 billion valuation from backers including General Atlantic, Nvidia, Fidelity, AllianceBernstein, Adobe, AMD and Felicis.

Other big raises this year include Yann Le Cunn’s AMI Labs, which snagged $1.03 billion in March from Cathay Innovation, Greycroft, Hiro Capital, HV Capital and Bezos Expeditions. Its value now stands at $3.5 billion. Then there’s General Intuition, which bagged $320 million at a $2.3 billion valuation in June from investors such as Khosla, General Catalyst and Bezos Expeditions, and Odyssey, which raised $310 million from Amazon, AMD and Google’s venture arm, at a $1.45 billion valuation.

The biggest round of all saw Skild AI closing on a whopping $1.4 billion raise led by SoftBank, Nvidia and Bezos Expeditions in January, propelling its value to an incredible $14 billion.

Infrastructure bets on a sure thing

One of the most notable themes in the trending interest in world models is that the money isn’t just coming from traditional VCs. In fact, the most active investors in world models are the compute infrastructure players that will probably benefit from their growth more than anyone else – chipmakers like Nvidia and AMD, and cloud infrastructure providers like AWS (often via Bezos Expeditions).

There’s a fascinating overlap occurring, too. Consider Decart, which has secured massive financial backing from Nvidia’s VC arm while simultaneously signing on to become one of AWS’s flagship Trainium3 partners. Odyssey is a similar story, with AMD being one of its biggest backers, while running its workloads primarily on AWS and Google Cloud infrastructure.

Confusingly, many of these infrastructure players are enthusiastically building their own world models, too. Take AWS, which is leaning on its Amazon Nova AI service to develop systems that can process a combination of visual, temporal and spatial data to power the robots in its massive warehouses. Meanwhile, Nvidia has built Nvidia Cosmos, an environment for training robots and physical AI to perceive real-world environments and react inside them, and Omniverse, for developing digital twins.

There can be only one reason why these infrastructure players are effectively competing against themselves. Quite simply, they see world models as the fundamental drivers of the next era of industrial automation. By investing heavily in the most promising frontier labs while simultaneously pursuing their own ambitious projects, it's clear that they're hedging their bets.

They can't predict which lab will ultimately win this race – their own teams or one of their external investments – but they're confident someone will. And when that winner emerges, world models stand to generate billions of dollars in revenue, not just for the developer itself, but for the cloud services and silicon that power it.

Following the money, it’s clear that the future of AI won't just belong to AI agents and chatbots grabbing the headlines today, but the physical AI systems that will live and work right by our side.

Disclaimer

Thought Leadership

!"#$%&'()*+,-./0123456789:;<=>?@ABCDEFGHIJKLMNOPQRSTUVWXYZ[\]^_`abcdefghijklmnopqrstuvwxyz{|} !"#$%&'()*+,-./0123456789:;<=>?@ABCDEFGHIJKLMNOPQRSTUVWXYZ[\]^_`abcdefghijklmnopqrstuvwxyz{|}