Essay · June 2026
The Data Robots Were Never Given
The Manual We Never Wrote
Every robot demo you have ever seen is a magic trick, and the trick is always the same. The impressive part is the thirty-second clip of a humanoid folding your laundry. The real part happened months earlier, off camera, in a room nobody films because it would put you to sleep. In that room, a person is moving a robot arm by hand, over and over, teaching it to pick up a cup. Then a slightly different cup. Then the same cup in worse light. The polished video is the plated dish. What you never see is the years of prep work in the back.
I want to start in that back room, because it is the whole story and almost everyone walks right past it. The popular take on robotics in 2026 is that we are one clever model away from machines that finally do our chores. I think that is exactly backwards, and the reason is a fact so obvious that it manages to hide in plain sight: we never wrote down how we move.
So here is my claim, and I am going to defend it the whole way down. Robots do not have a hardware problem, and they do not even really have a brains problem. They have a data problem. And that data problem exists for a strange, almost funny reason. Moving was always so easy for us that it never occurred to anyone to record it. Nobody ever sat down to document how they pick up a water bottle, or fold a shirt, or load a dishwasher, because none of it ever felt like it needed explaining. That missing instruction manual, the one we never thought to write, ends up deciding everything that follows: who wins, who gets crushed, and where the real money in this whole wave actually ends up.
The Library We Inherited for Free
Let's take a look at LLMs.
The thing that made large language models possible was not one genius idea. It was a library. For thirty years, all of humanity poured its mind onto the internet, every argument, every recipe, every legal filing, every 2 a.m. shower thought, and we did it without realizing we were assembling the most valuable training set ever built. GPT and Claude did not have to create any of that. They walked into a library that was already full and simply started reading. The intelligence came almost by accident, a lucky side effect of the fact that humans happen to be a species that leaves words everywhere we go.
Now ask the question that cracks robotics wide open. Where is the same library for movement?
There isn't one. And the reason there isn't one is almost poetic. We left a record of our words, but never of our hands. Think about the last time you picked up a coffee mug. You ran a routine so automatic that you could not explain it even if I paid you. Which muscles fired, how hard you gripped, the tiny correction you made when it turned out heavier than you expected. You don't know. You were never even aware that you knew. It is the same reason you cannot teach someone to ride a bike with a paragraph, or explain how you keep your balance on stairs. The skill is absolutely real, and it is sitting inside you right now, but you could never put it into words. Millions of years of evolution polished your sense of movement until it turned so smooth it became invisible, even to you. And what stays invisible never makes it onto the page.
This is the part I keep coming back to, because it quietly flips the most famous puzzle in the field. The puzzle is called Moravec's paradox: the strange observation that the things hard for humans, like chess and calculus, turned out to be easy for machines, while the things easy for humans, like walking and grabbing, turned out to be nightmares. The usual explanation is a bit mystical. Our movement is ancient and deep, people say, so it must be almost impossible to copy. I think that explanation is too fancy and basically wrong. The gap is not magic. It is bookkeeping. Machines conquered chess and calculus first because we had handed them every game and every proof ever recorded. Of grabbing a cup, we handed them nothing. The paradox was never about how hard movement is. It is about which of our skills we ever bothered to save, and we saved the thin, recent, thinking layer while the old, automatic one stayed in the dark.
So the real bottleneck was never some unknowable mystery. It is a blank page. And someone now has to fill it in by hand.
Why You Cannot Download a Dishwasher
This is where the optimist shows up, and the optimist has a genuinely good argument that I want to put at full strength before I lay a finger on it.
The optimist says: relax, this is just a data problem, and data problems are the thing we are best at in the world right now. The whole playbook of the last decade was to take some messy thing, turn it into a stream of tokens, throw a mountain of computing power at it, and watch the magic happen. It worked for text. It worked for images. It worked for audio, and even for the way proteins fold. It is how everything, eventually, became the Transformer. Robotics is just next in line. Gather the data, scale the model, and your laundry folds itself. The only thing standing in the way is getting enough robots out into the world to collect it.
I find this argument genuinely seductive, and I think it hides one quiet flaw that sinks the whole thing. Every single domain where this trick worked shared one feature: the data was already there, in enormous piles, free to copy. The internet already existed. The giant image sets already existed. The Transformer did not produce those mountains of data, it just ate them. So the optimist is really pointing at a long list of past wins that all share the exact ingredient robotics is missing. Saying “it worked everywhere there was free data” tells you nothing about the one place where the data is not free. The argument quietly assumes the very thing we are trying to figure out.
Because here is the cruel asymmetry. Copying a sentence costs nothing. You can duplicate the entire works of Shakespeare a billion times for the price of a rounding error. But a single clean recording of a hand doing a real task, with all its little forces and corrections intact, cannot be copied into existence. It has to be made, in the real world, one second per second, by a real robot or a real person, on real hardware that wears out and breaks. It is the difference between streaming a movie and cooking a meal. The movie streams to a million people at once for almost nothing. The meal has to be cooked one plate at a time, every time, forever. Language scaled because the next sentence was basically free. Robotics will not scale that way, because the next motion has a hard price tag made of time, money, and worn-out machines. That price tag is not a minor detail. It is the entire reason this is hard.
The Race to Write It Down
If the manual has to be written by hand, the obvious question is: written how, and written by whom? A whole industry has sprung up to answer this, and watching where it puts its money tells you where the real fight is.
The first approach is to turn humans into walking recorders. Strap on the gear and harvest the data we generate just by living our lives. The serious versions are not just cameras filming you. The good rigs add tactile gloves, cameras on the wrist, and trackers that capture the exact angle of every joint and the actual force your hand applies when it squeezes. Startups are handing these kits to warehouse workers and selling the recordings straight to robot companies. By one estimate, robot firms are already spending north of a hundred million dollars a year buying real-world data (mostly funded by VC dollars right now), and one vendor claims to have logged more than a hundred thousand hours of it. The manual is being written, and people are getting paid by the hour to write it.
The second approach is to have a human puppeteer the actual robot, like a marionette, so the robot's own movements and forces get recorded directly. Everyone agrees this gives the cleanest data. Everyone also agrees it barely scales, because every single station needs a real robot and a real person running it in real time. These teams are counted in dozens, not thousands.
The third approach is the most elegant on paper and the most painful in practice: just put real robots out into the world, let them work, and feed everything they experience back into the model. This is the famous flywheel everyone dreams about. More robots make more data, more data makes better robots, better robots get deployed in bigger numbers, and round it goes. The catch is the cold start, and it is a brutal one. Nobody pays a hundred thousand dollars for a robot that fails half the time while it learns, but the robot cannot stop failing until it gets the very data that only real deployment can give it. It is the same trap as a new restaurant that needs reviews to attract customers but needs customers to get reviews. The flywheel does not start spinning on its own. Somebody has to shove it, hard, before it catches.
Notice what all three of these have in common. They are slow, physical, and expensive, and the data they produce is nothing like the free public internet. It is paid for, recorded in real time, and fought over tooth and nail. The manual for movement is being written, but it is being written in sweat and money, not scraped for free off a webpage. That single fact changes who is even capable of building it. And that changes everything about who wins.
The Generalist's Bargain
There is a real debate buried inside all of this, and I have genuinely flip-flopped on it, so let me lay it out straight instead of pretending I have it all figured out.
One camp says the winners will be specialists, each robot drilled on its own narrow job with its own narrow data, like a machine on an assembly line that does exactly one thing forever. The other camp says the winners will be generalists, a single brain trained on a huge variety of tasks that can then carry what it learned into new ones, the way one person who speaks five languages picks up the sixth faster than someone starting cold. History leans hard toward the generalists. A generalist scales faster, because you pay for one giant training run and then spread that cost across a hundred different jobs, instead of paying a fresh bill for every new task the way the specialist does.
But the generalist's edge comes with a hidden price, and the price is subtle. The generalist does not actually need less data. It needs a different flavor: wide instead of deep. A specialist wants a thousand hours of one task. A generalist wants a hundred hours each across a hundred different kitchens, because what teaches it to handle a situation it has never seen is variety, not repetition. The most striking result I have seen says exactly this: a leading generalist model learned to clean homes it had never set foot in, and the thing that mattered was not how many demonstrations it got, but how many different homes it had seen along the way. Roughly a hundred was enough. Variety beats volume.
So the generalist does not escape the data problem. It just moves it, from “collect a mountain of one thing” to “collect a little of everything, everywhere.” And here is the sting, the part the generalist crowd rarely says out loud. The whole generalist bet is a bet that the data you need for each new task keeps shrinking, because the model already understands the world and only needs a nudge. But if that is true, then data eventually stops being scarce. And the moment data stops being scarce, it stops being a moat. You cannot believe both that the generalist fully wins and that data stays a permanent advantage, because the only world where data stops mattering is the exact world where the generalist runs the table. The two beliefs eat each other alive. Anyone who tells you they are bullish on both at once just hasn't noticed they are holding two cards that cancel out.
Where do I land? I think the generalist is the right long-term direction, and I think data scarcity is the right near-term reality, and I think those two can live together for one specific reason: we are still in the years where the manual is only half-written. While the pages are still mostly blank, data is the thing everything else is waiting on, and the moat is real. The generalist's bet is that this phase ends. Until it does, and it has not, data is the single most valuable thing in robotics.
The Software Was Never the Moat
Now let me grant the optimists the whole argument. Say someone builds a magnificent general-purpose robot brain that learns new tasks from almost nothing. Here is the question that kept nagging me: would that company actually get rich from it?
I don't think it would, and the reason is a pattern I have written about before. In the last wave, the celebrated app-layer startups looked untouchable right up until you noticed they were renting every ingredient that mattered and owning none of it, and the value quietly drained past them to whoever owned the scarce thing underneath. Robotics has its own version of this, and it is already happening out in the open. The leading robot-brain company is giving its model away for free, on purpose, betting it can become a platform instead of a product. When the smartest people closest to the frontier are handing out the brain for nothing, that is the market screaming that the brain is about to become cheap and interchangeable. It will go the way every powerful model now goes: someone proves a thing is possible, and within a year a competitor copies it for a fraction of the price and gives it away.
So if the brain becomes cheap and is once again just borrowed intelligence, where does the lasting value actually go? The easy instinct is to say it stays in the physical robot, because a robot is made of atoms, not bits, and you cannot copy atoms off the internet. That instinct is right, but it is too sloppy, and being too sloppy is exactly how you land on the wrong answer. Most of the physical robot gets cheap just as fast as the software does. The frame, the basic motors, the wiring, the shell, all of it is already in a price war. The cheapest capable humanoid in 2026 sells for around thirteen thousand dollars, built almost entirely without Western parts, because most of a robot's body is no harder to mass-produce than a drone or a washing machine. “Physical” does not mean “protected.” A folding chair is physical too, and nobody got rich selling folding chairs. The Chinese humanoid is all atoms, and its margins are getting crushed anyway.
The lasting value lives somewhere much narrower: the handful of physical parts that stay expensive because they are genuinely, brutally hard to make. Crack open the cost of a humanoid and roughly a third of it is the precision movement parts, the joints and the gearboxes and the screws inside them. The sharpest example is a thing called a planetary roller screw, the part that turns a motor's spin into the exact straight-line push a robot leg needs to stand and walk. Making one is closer to building a Swiss watch movement than stamping out a car door. One supplier that sent samples to Tesla could produce only about three hundred good ones a month, enough to build roughly ten robots. The blunt truth from inside the industry is that even when the motors get assembled in America, the parts that actually matter still come from a tiny set of suppliers in China, Japan, and Germany, because these are precision instruments and the know-how took decades to build. You can spin up a flashy robot brand in a year. You cannot spin up a roller-screw industry in a year. That gap is the moat.
And look at who actually holds it. Not the famous robot brands everyone tweets about. The chokepoint parts are owned by a small group of quiet specialist suppliers most people have never heard of, the firms selling the same screws and gearboxes to every robot maker at once. So when Tesla and Figure and a dozen Chinese upstarts claw each other to death on price, the company selling all of them their roller screws wins no matter which brand is left standing. This is the picks-and-shovels story all over again, the merchant selling shovels to every miner in the gold rush while the miners mostly go broke. It is the same conclusion I keep arriving at from every direction: difficulty is the only moat that doesn't evaporate. In the chatbot wave, the difficulty was the chips. In robotics, it is the precision parts and the rare-earth magnets buried inside the motors, the things whose recipe simply cannot be downloaded.
The Knife That Cuts Both Ways
I want to be honest about where my own argument runs out, because the supply-chain moat is not forever either, and pretending it is would be the same overconfidence I just spent an essay attacking.
A physical chokepoint only stays a moat for as long as the know-how stays rare. The day that knowledge spreads and the designs settle into a standard, even a hard part turns into a cheap one. We can already watch it happening live. The harmonic reducer, another precision part that was a real bottleneck just a few years ago, is now being cranked out by Chinese challengers at roughly half the price of the Japanese pioneer. It is the same arc as flat-screen TVs, which went from a luxury that cost more than a used car to a commodity stacked on a shelf at the supermarket. The roller screw will probably hold out longer because it is harder, but the direction of travel is obvious. So the honest version of my claim is not “physical parts win forever.” It is “the hardest-to-copy parts win for exactly as long as they stay hard,” and the whole game is a bet on how long that is. If you want one number to watch to see whether I am aging well or badly, watch the price and the number of suppliers for planetary roller screws over the next two years. If a dozen firms hit Tesla-grade quality and the price falls by half, my moat is melting. If it stays a tiny club charging four figures a screw, the moat holds.
There is one more layer I genuinely cannot settle, and being honest means saying so out loud. If the brain gets cheap and the body gets cheap, there might be a third place the money pools that is neither one: the platform sitting in the middle, the standard that every robot has to run through. This is the Android bet. Google gave its phone software away for free, let the phone makers and the apps fight it out on price, and quietly became the layer that all of them depend on. It is a real strategy and it has made staggering amounts of money before. Whether it works for robots comes down to a question I honestly do not know the answer to yet: can anyone build a hook, some data flywheel or some safety-and-trust layer, that the free brain quietly funnels everyone's dependence into, the way Google kept its own services locked up on top of free Android? If they can, then the platform is a fourth chokepoint, and it is made of information instead of atoms. If they can't, then the free brain is just a gift to the world, and the money flows right past it to the atoms, exactly like the rest of this essay says. I lean toward the atoms. But I hold it loosely, because that platform play is the one door through which a piece of software could escape the fate I am predicting for all the others.
The Thesis in One Line
Robots are not waiting on a smarter brain. They are waiting on an instruction manual we never wrote, because moving was the one thing we always knew how to do and never had to explain. And when that manual finally gets written, in sweat and money instead of scraped for free, the money will not pool in the brain that reads it or the body that carries it, but in the few small parts so hard to make that nobody can copy them over a weekend.