A language model can learn to write by reading the internet. A robot cannot learn to load a dishwasher the same way, because the internet holds almost no record of what it feels like to grip a wet plate at the right angle and set it down without cracking it. That gap is why humanoid robots still fumble tasks a four-year-old handles without thinking. In the first half of 2026, investors poured $8.6 billion into humanoid robot startups, nearly double last year’s total, according to Dealroom. But that money is chasing machines that still cannot reliably fold laundry.
What they lack is data of a particular kind: recordings of physical work, made by hand.
Most robots today learn by imitation. A person performs a task, the robot records what happened, and a model learns to connect the situation to the right action. The demonstration has to come from somewhere, though, and the cleanest way to produce it is to have a human operate the robot directly. This is teleoperation. An operator wears a headset or holds a set of controls, moves the robot’s arms through the task, and repeats it dozens or hundreds of times with small variations. Each run becomes a training example. Companies that provide teleoperation services for humanoid robots run rooms full of these rigs, capturing the grip, the timing, and the small recovery when something starts to slip.
Teaching a robot to fold a shirt can take a thousand folds, because the model needs to see the fabric bunch, the sleeve catches, the corner gets missed. The failures matter more than the clean runs. A robot that has only seen a task go right has no idea what to do when it goes wrong.
Raw demonstrations are only half the job. Before the footage is useful, someone has to label it: mark where one action ends and the next begins, tag which object the hand is reaching for, flag the frames where the grip failed. This is slow, exacting work, and most robotics teams do not do it themselves. They send it out. Grand View Research projects the data labeling market will reach about $57.6 billion by 2030, and most of that work is outsourced. Much of the data annotation outsourcing that trains an American humanoid happens on the other side of the world, frame by frame, by people who never see the finished machine.
Simulation helps, but the frictions of the real world, a slick handle or a shadow that fools a camera, still have to be captured for real.
None of this shows up in the demo videos. When a humanoid pours a cup of coffee on stage, the clip lasts a few seconds. Behind it sit weeks of operators repeating the motion, annotators marking it up, and quality checks that throw out the sloppy runs. The robot looks autonomous. The autonomy was assembled by hand.
There is a quiet irony in that. The humanoid promises to take on physical work, so people do not have to. For now, teaching it to do that work is itself a large and growing kind of physical labor. And the operators are not a scaffold that falls away once the models improve. Every new task, new environment, and redesigned gripper starts the collection over.
The hard part of the field is no longer building a machine that can move. It is teaching the machine what to do, and that still runs on human hands, human judgment, and a great deal of repetition. The robots are learning to move. People are showing them how.


