← All research

№ 03 · Aug 2026 · ALLOCATORS · OPERATORS · OBSERVERS

The Invoice: How Physical AI Actually Arrives

Robot brains are quickly becoming free downloads. Getting machines to actually work in the real world is still brutally hard, and in my opinion that is exactly where the value is going: to whoever owns the deployment, the environments, the data and the accountability. This paper is my case for why.

The Dock

In July I stood on a dock on an American river. I had flown five thousand miles to sell an underwater inspection service to one of the country’s largest industrial operators, but instead, they spent the afternoon making the case to me.

They run hundreds of vessels, and every single one gets pulled out of the water every ten years just because the regulator says so. They told me plenty of those dockings find nothing wrong. One unplanned docking cost them three months. At no point did anyone ask me whether the robot works. They told me what it would be worth to them.

That was the moment physical AI stopped being a prediction for me. The demand arrived before the machines did.

I am Rohith, the founder of ScrubMarine. We build underwater robots that inspect and clean the assets most people never think about, ship hulls, barges, docks, offshore wind foundations. This is the third paper I have published, after one arguing that India should become the place where the world’s AI actually runs, and one called Selling Shovels in the AI Gold Rush, about who really gets paid when everyone is digging. To be upfront, I am currently raising for ScrubMarine, so of course I am biased towards this future. I have tried to let the evidence do the talking, but you should read this knowing where I stand.

Physical AI, In Plain Terms

Physical AI is software that perceives the world, makes a decision, and then acts on it through a machine. ChatGPT is AI that thinks, and it lives on your screen. Physical AI is AI that acts, and it lives in the real world.

Robots themselves are not new, and I think this distinction matters. Factories have had robot arms since the 1960s, but those machines are blind and scripted. They sit bolted to the floor inside a cage, repeating one motion a million times, and if you move the part an inch the whole line stops. Physical AI is what happens when the machine can actually see, adapt and decide for itself in a messy environment nobody scripted. That jump, from repeating to understanding, is the whole revolution.

And to be clear, the machines are not what changed. My lead engineer made this point when he read a draft of this paper, and he is right: the physical side of robotics has not moved dramatically in a decade. Underwater robots, drones and robot arms have existed for years. What changed is the gold rush in AI. The billions poured into data and compute to build language models quietly paved the way for machines that can finally understand what they are looking at. The hardware was waiting. The intelligence arrived.

The strange thing, which took me a while to appreciate, is that the thinking part turned out to be the easy bit. In 2026 a frontier model can draft a merger agreement in a minute, yet a robot still cannot reliably fold a shirt it has never seen.

Here is what has changed recently, and I know because my team uses these tools every day. The robot brain is now a free download. NVIDIA gives away a foundation model for humanoid robots. Physical Intelligence gave away its general robot brain, trained on roughly 10,000 hours of real robot data. Google runs robot intelligence on the machine itself. Three of the most capable robot brains on earth now cost nothing.

The rest of the stack collapsed with it, or more precisely, the entry level did. The same engineer puts it best. At the start of his degree, 3D printing a part meant building the printer yourself, tuning it for hours, then watching the first layers go down with a coin-flip chance of success. Now the printers in our office just work, and a beginner can run them. Code is the same story. Half our software gets written with AI coding tools now, and structural parts that would have been a machine shop invoice two years ago come off a printer overnight, strong enough to bear load. A small team in Edinburgh can now field industrial robotics, and so can a team in Chennai, or Lagos. Building the machine has stopped being the hard part.

What has not collapsed is everything around the machine. The best models still memorise rather than understand. Move the objects a few centimetres and the benchmark champions fall apart, because there is no internet of touch. The largest open robot dataset in the world holds about 2.4 million recordings. The words that trained ChatGPT were already typed out and sitting there, trillions of them. Nobody has typed out the physical world, and nobody ever will.

So in my opinion the scarce assets are not brains. They are the environments machines work in, the data collected by doing the work, and the accountability that regulated industries demand. The value goes to whoever owns deployment. The robot is the instrument. The record is the business.

A quick word on shape too, because the world is currently obsessed with robots that look like people. A fish is better at swimming than a human, so why would you build a swimmer shaped like a man? Humanoids make sense in environments designed for humans, corridors, stairs, factory lines. Most of the industrial world is not shaped like a corridor. It is a hull, a pipe, a tank, a turbine tower, a sewer. Form follows deployment.

The Invoice

Lots of people ask me when we will see robots everywhere, which in my head I always translate to the same question: when does physical AI get its ChatGPT moment?

It is worth remembering what the ChatGPT moment actually was. I was using OpenAI’s playground long before ChatGPT launched. I used it to write sales emails in an old job, and later inside my first attempt at a startup, if you can even call it that, which built AI for car dealerships. The technology was already there, quietly doing work. Then one November a friendly chat window appeared on top of it and the world flipped in a week. The tech did not change overnight. The belief did.

I do not think physical AI gets its moment the same way. In my opinion, the ChatGPT moment of physical AI will not be a demo. It will be an invoice.

In July I was at Machina in Paris, the first summit anywhere dedicated to physical AI, and the consensus from the people actually building was blunt. Industry first, homes later, because the cost and nuance of deployment decide the order, not imagination. Nobody at a refinery or a shipping line is waiting for a viral video. They are waiting for a number on a quote and a line in an insurance policy. The revolution arrives industry by industry, the first time a robot workforce is cheaper and insurable.

A viral humanoid moment will still come, a machine doing a full shift on a livestream, and the world will lose its mind when it does. But just like ChatGPT, it will be the polished consumer version of years of quiet progress inside businesses, and it will arrive late. By the time that demo lands, I believe the industrial economy will already have been rebuilt underneath it by machines nobody streams. Anyone waiting for the demo to believe is waiting to become a customer of the people who did not wait.

We have actually already seen a preview of that moment, at small scale. Late last year a $20,000 home humanoid went viral overnight, tens of millions of views, everyone losing their minds about the robot butler finally arriving. Then people learned the demos were being piloted by humans in VR headsets, that a stranger would be looking through its cameras inside your house, and the excitement curdled into memes within about a week. That is what belief built on demos does. It evaporates the moment someone checks the invoice.

Who Goes First

Unit economics decide the order, and here is the order I believe it happens in. The popular argument says robots arrive where labour is scarce. I heard it on stage in Paris applied to shipbuilding, not enough welders, so humanoids fill the gap. I think that is only half right, and the half matters.

In my view, robots win first where the wage was never the cost. Think about what it takes to put a human under a ship. Dive training. Oxygen. Decompression schedules. A support vessel idling on the surface. Insurance priced off the people who did not come back up. The diver’s day rate is the cheapest line on that invoice. My robots do not compete with labour. They compete with life support.

And this is dangerous work by any measure. In the United States, the CDC has put the fatality rate for commercial divers at roughly 40 times the national average for workers, and that is in one of the most regulated markets on earth. A huge amount of the same work happens in parts of the world where the rules are thinner and the deaths barely make a report.

Dry dock tells the same story at asset scale. To inspect a vessel properly today you take it out of service, tow it, lift the entire thing out of the water and put people around it for weeks. The inspection itself is a rounding error. The downtime and the logistics are the real bill, and that bill exists because the whole process was designed around human limits.

Call it physiology arbitrage. Wherever keeping a human alive in the environment costs multiples of the human’s wage, the robot’s business case writes itself. Underwater. Inside tanks and reactors. At height. In poisonous air. That, I believe, is the frontier, and it adopts first.

Next come the places where labour is scarce and cannot move. Shipyards are the honest version of the humanoid case, because you cannot import welders at will. Immigration, certification, unions. Scarce and immobile labour is a real gap for machines to fill.

Last, and in my opinion possibly never, comes pure wage arbitrage. A robot competing with abundant cheap labour loses on price for a very long time. Robot baristas and burger-flipping machines keep appearing and keep quietly disappearing, because a six-figure machine going head to head with someone on minimum wage is a bad trade, and I have given exactly this feedback to founders building consumer robots.

The ordering also exposes what I call the generality trap, which in my view is the humanoid problem. A specialist robot needs to be superhuman at one task to earn its keep. A humanoid needs to be decent at dozens before anyone pays, because doing everything is its whole sales pitch. Nobody pays £25,000 for a robot that folds clothes. They would pay it for a full maid, and a maid is a hundred tasks deep. The specialist reaches breakeven years before the generalist has a product. Today the shortcut to generality is teleoperation, a human piloting the robot remotely, which in a home means a stranger with camera eyes in your kitchen. Industry does not much care about that. Homes will. One more quiet reason industry goes first.

The people building better robots are currently the main buyers of humanoids. TIME put it plainly this summer: the modern humanoid industry mostly sells robots to people trying to build better robots. That is not a market. It is a rehearsal.

What the Water Taught Me

I have spent just over two years building underwater robots, and here are five things real water taught me that the demos will not.

One. Autonomy is a fleet technology, not a solo technology. This contradicts most of the money in the sector. With one robot, autonomy is economically worthless to me, because a crew still drives to the dock, lifts the machine into the water and manages the tether. Autonomy does not remove those humans. It pays only when it multiplies them, one crew supervising twenty machines instead of piloting one. Autonomy is a fleet economics unlock, not a magic ingredient. In my opinion, the industry is racing to remove the human from a picture where the human was never the constraint.

Two. The certifier is your second customer. No buyer has ever asked me which model runs the autonomy. Every buyer asks whether the data will stand up in front of a surveyor. Regulated industries adopt at the speed of class societies, coastguards and insurers, not the speed of the technology. You sell twice. Once to the operator, and once to the people who let the operator say yes.

Three. You cannot buy the dataset. When I first started I assumed we could train on existing subsea data, and I was wrong, because it barely exists. The only way to get the data that matters is to do paid work in real environments and record everything. The moat everyone theorises about arrived in my life as a disappointment first and an asset second.

Four. You cannot unit-test the ocean. Software fails in testing. Hardware fails in the world, and the world does not announce its test cases. Our tether once snagged on a slab of concrete no drawing knew about, and our lead engineer went into the water after it. An overexposed camera ate a full day. Pinhole leaks found the one seal we trusted. We once arrived at a test site and realised nobody had packed the controller. None of that was in the simulation. All of it is now in the product.

Five. Customers buy outcomes, not frontier tech. Our customers do not care about the AI or the cloud platform. They care about extending inspection cycles, catching problems before dry dock, and reports a regulator will accept. The most advanced thing about a good robotics company is how boring its value proposition sounds.

The Flywheel

Everything I have learned collapses into one loop, and the loop is the business model. Industries pay for a service. The service funds the robots. The robots collect data nobody can scrape, because it only exists where the work happens. The data improves the decisions and the machines. Better machines win more work. The service funds the data. The data builds the moat.

Notice what the loop does to the autonomy question. The data we collect with piloted machines today is what makes supervised fleets possible tomorrow. The flywheel does not wait for autonomy. It manufactures it.

And I believe the dataset is deeper than sensor readings. When I was younger I used to tell my family, completely seriously, that one day I would upload my mum’s consciousness into a robot so she could be with me forever. It became a running joke in our house, and when ChatGPT launched I heard about it all over again. Obviously nobody is uploading anyone’s consciousness. But the older I get, the more I think the younger me was pointing at something real, because you cannot upload a person, but you can start to capture their judgment.

One of my engineers once asked me how I knew something was wrong before the data showed it, and I told him the truth. Gut instinct is not magic. It is every previous time I made that call, compressed. The more hours you spend on a task, the more decisions you make, and the more decisions you make, the better the map gets. Psychologists have been saying a version of this for decades, that intuition is just pattern recognition earned in environments that give you feedback.

But it got me thinking. If gut instinct is a decision map, then gut instinct is data. And if it is data, it can be collected, and it can be trained into machines.

So the moat is not just what the robot senses. It is every human judgment made around the machine. When to abort a run because the water feels wrong. Which reading is corrosion and which is noise. The field’s biggest labs are waking up to this, and the largest new robot dataset out of China added something telling in its latest release, recordings of robots failing and then recovering, because failure data is the scarcest kind. We record failure and judgment as a byproduct of paid work. That layer, the judgment layer, cannot be scraped, bought or simulated, and it compounds privately.

It also means something humane. For all of history, the experienced diver trained the young diver. In this decade, I believe the last generation of humans who did the dangerous work will train the machines that end it. Their gut becomes the model’s starting point. Nothing about that experience is discarded. It becomes the most valuable input in the system.

Anywhere the Work Is Dirty, Dangerous or Regulated

I believe this playbook goes far beyond water, and this is the claim I care about most. Look at the loop again. Someone pays for a job. The job funds the machine. The machine records what no one else can, and the record compounds. Nothing about that requires an ocean. It runs anywhere work is physical, repeatable and valuable to know about. Pipelines. Tank farms. Substations. Wind turbines. Rail. Nuclear pools. The spectrum runs from the warehouse floor to the reactor pool, and the playbook is the same at both ends: build the service, earn the environment, keep the data. Hub and spoke. A central brain and fleet operation, with spokes into every industry where the work is dirty, dangerous or regulated.

And it is not always the glamorous, high-cost jobs. Growing up, when I would go to India on holiday, I saw a version of this that has never left me. Human beings there still physically climb into sewers and septic tanks. Close to 1,800 people have died doing that work since 2000 on the recorded numbers alone, the real count is higher, and it lands almost entirely on the most marginalised communities. But times are changing. A startup in Kerala built a robot specifically to end the practice, and that machine now works across more than a dozen Indian states. Robotics has a name for this category, the three Ds, dirty, dangerous and demeaning, and I believe physical AI is the first technology that can retire all three. Some of this revolution will be measured in fleet hours and inspection cycles. The part I care about most will be measured in the last person to ever climb into a sewer.

Who Loses

The losers, in my opinion, are everyone holding exactly one piece of the stack, and it starts with brains without bodies. If the robot brain is a free download, a company whose only asset is a robot brain owns a melting asset. Some will buy their way into deployment. The rest are selling something whose price trends to zero.

Bodies without deployment. The hardware makers shipping robots as boxes, no service, no data loop. They are arms dealers in a war where the territory is data, and they hand the territory to whoever operates the machine. There is a human version of this failure too. Engineers who never deploy build like professors, from the drawing, for a world that behaves. So many things changed for my team the day we walked our customer’s site. The tether snags on the slab of concrete the CAD model never had. You cannot build like a professor. The site is the teacher.

Deployment without data. The legacy service firms, and I compete with them, so I say this with respect and no hesitation. Dive companies and inspection houses own today’s customer relationships and treat the data as exhaust. Every job they complete without keeping the record is a moat they chose not to build, and I believe their customers will notice before they do.

Then the structural one. For seventy years the development playbook for poorer countries was cheap manufacturing labour. Japan climbed it, then Korea, then China, then Vietnam. Physical AI burns that ladder. If a machine does the work below any wage, cheap labour stops being a national strategy. I wrote about the white-collar version of this in Selling Shovels, the outsourcing firms and capability centres that sell cognitive output by the head are compressing as thinking gets cheap. This is the same squeeze arriving one floor down, for physical work. The countries that win the next chapter will not be the ones selling labour to the owners of machines. They will be the ones building deployment capability themselves. The tools are free and the hardware is cheap, and a team in Chennai can run this playbook as well as a team in California. Whether nations seize that is the subject of my next paper.

And one casualty nobody will mourn. Danger itself. The diver in the decompression chamber, the worker in the manhole, the inspector on the rope at height. The point was never to remove people from work. It is to remove people from the places that kill them, and put them behind glass where their judgment is the product.

2036

I watched The Martian as a young teenager, and for years afterwards all I wanted was to study aerospace and end up a flight director at NASA. I took a different road, but the obsession with frontier machines never left me, and it shapes how I see the next ten years. So here is what I honestly think the world looks like in 2036.

A senior operator schedules a hull inspection from one screen, because the predictive model flagged it, not because a calendar did. An autonomous vehicle delivers the mothership to the water, the mothership launches the swarm, and the machines work the hull while their data streams to a hub where a small crew supervises every live job on the planet. Humans step in for exceptions. The asset itself was designed to be worked on by machines, docking points, standardised access, machine-readable markings, the way ports were once redesigned around the shipping container.

The same operating system runs everywhere. Crawlers live permanently on wind turbines and the farm inspects itself between storms. Field machines know every plant individually and treat single stems instead of spraying acres. Power lines, substations and water mains are walked, flown and swum continuously, with leaks found before they surface. Building sites are scanned nightly, what was built checked against what was designed every morning, steel signed off without a person at height. Reactor pools and mine shafts are finally human-free. Sewers are entered by machines or not at all.

Inspection stops being an event and becomes a condition. Today a regulator forces an asset out of the water every ten years to find out what state it is in. Before this decade ends, I believe a regulator somewhere will accept continuous robot data in place of a scheduled docking, and the calendar will start giving way to the condition.

The humans have not disappeared. They do the three things machines cannot own, judgment, exceptions and accountability, because someone still signs their name for the regulator, and it will be a person. The veterans of the dangerous trades supervise the fleets, because forty years of gut is the most valuable training data on earth. And the humanoids, when they finally mature, take the role that was genuinely shaped like a person all along, the exception handler, the machine you send when the situation is too human-shaped for anything else.

And the reason The Martian still matters to me is that space is the final version of this whole argument. No air to breathe. Pressure that kills. No rescue coming. Every human hour costing a fortune in life support. These are the exact constraints I deal with underwater, turned up to eleven. The machines proving themselves in the ocean this decade are, in my opinion, the ancestors of the ones that will maintain whatever we build off this planet. Physiology arbitrage does not stop at the waterline, and it does not stop at the sky.

Where I Might Be Wrong

I would be silly not to look at the opposite side of this paper, and I have thought a lot about how it breaks down. In my view it breaks in one identifiable way. If world models close the gap between simulation and reality while cheap open hardware keeps collapsing the cost of the machines, then building and running robots becomes cheap and easy for everyone too, and the value swings back to whoever ships the most capable cheap robot. Regulators could also freeze, and the certifier I called my second customer becomes everyone’s ceiling. The evidence in mid-2026 does not show either happening. The best models still fall apart when the scene changes slightly, and the regulated world is tightening its rules, not loosening them. But I keep both under watch rather than treating them as settled.

The Money Question

People keep asking me whether all of this is a bubble, and my honest answer is that parts of it are. Robotics startups raised $18.8 billion in the first half of 2026, more than the whole of last year, in a sector that has historically returned money slowly and rarely. In my opinion, teleoperated machines dressed up as autonomous ones should not be valued in the billions, and some of today’s capital is pricing demos as if they were deployments.

But here is the part people miss. Bubbles have a habit of building useful things. Railway mania bankrupted its investors and left Britain a railway network. The dot-com crash torched trillions and left the fibre the internet still runs on. Bubbles are how capitalism funds infrastructure ahead of demand. The capital dies. The capability survives. So the question is never whether it is a bubble. The question is what is left when it pops.

The smart demand-side money has already moved, by the way. Y Combinator’s latest request for startups reads like a physical AI shopping list, operating systems for physical work, data from the real world, robotics category after robotics category.

Three things survive a pop, in my view. Revenue from customers who would pay even without the hype. Assets that do not evaporate when funding does, earned data, certified access to regulated environments, the judgment layer. And unit economics that work at today’s capability, not promised capability. Read the flywheel again and do your own maths.

At Three Altitudes

For allocators. Your job in a bubble is not to sit it out. It is owning what survives. Back the deployment owners, contracted environments, live data flywheels, recurring revenue. Judge on fleet hours and named customers, never demo reels. Hold general-purpose humanoid hardware as options, sized as options.

For operators. Your environment just became an asset class. If machines will work in your plant, port, farm or fleet, the data generated by that work is the compounding prize, and whoever keeps it builds the moat. Negotiate data rights with the seriousness you negotiate price. Deploy before you perfect. The site is the teacher.

For observers. Ignore the demos and watch invoices. Watch customer-published numbers, not vendor-published. Watch fleet hours in real environments. And watch the certifiers, because regulated industries move at the speed of the letter of compliance, and the first regulator to accept continuous machine data over a calendar will have started the avalanche.

Back to the Dock

I think about that afternoon on the river a lot. A century-old industrial giant explaining to a startup barely two years old why the future I was selling was one they had already priced. The machines are coming to the water, the sewer, the reactor, the field. Most of them will never be streamed, and the work they do will be boring to watch and enormous to own.

The demand arrived before the machines did. We are building the machines.

Rohith Devanathan

References