RESIDUAL ROBOTICS

Mass UMI data for the pretraining era of robotics

What we do

We're building the infra, ops, and tech to collect mass UMI (Universal Manipulation Interface) data from skilled laborers in India, in order to sell to robotics foundation model labs.

Current robots struggle to generalize to unseen situations. Simple changes in the environment, like shifting objects around or changing the lighting, can lead to failure. This is because robots are trained on small, narrow data sets, usually collected via teleoperation.

UMI data collected from skilled workers and postprocessed well has comparable quality to teleop, but it's multiple times cheaper, easier on the part of the operator, and easier to aggressively scale because it plugs into the existing economy.

LLMs solved the generalization problem by pretraining on the scale of all the text humanity has ever produced (~100 million years of labor). This gives the model a broad and deep understanding of the world, which allows it to learn usefully from posttraining. Pretraining is hard for robots because we don't have enough diverse data -- not enough from teleop, not enough if you scrape every video on the internet. But the evidence suggests it will work.

Data is the lifeblood of machine learning. Data is the bottleneck of the future. Let's collect enough data to unlock the pretraining era of robotics.


Team