Humanoid robots need tens of millions of hours of real physical interaction data, joint movement, force, touch, and egocentric vision, that cannot be scraped from the internet the way text and image data can, and almost nobody outside China is building the infrastructure to capture it at scale. This company builds the physical capture infrastructure, exclusive sites, multi-sensor rigs combining depth cameras, tactile sensors, and egocentric video, either through wearable capture systems that pay ordinary people to record their daily routines, or through dedicated facilities where operators generate structured task data, then sells the labeled corpus or a data-as-a-service relationship to robotics OEMs and foundation-model labs. The buyer is any humanoid robotics company or physical-AI lab that needs training data volume and diversity it cannot generate in-house.
The wedge is picking one capture modality, most plausibly the distributed wearable-headband approach given its lower capital intensity relative to building large dedicated facilities, and proving data quality and diversity is good enough for a paying robotics customer to train on. Distributed capture also has a structural advantage over centralized facilities: it generates diversity across homes, workplaces, and task types that a single controlled facility cannot replicate.
Once a data pipeline and quality bar are established with one or two anchor customers, the company can scale capture volume geographically and expand into additional sensor modalities and task categories, deepening the moat with every additional hour of data collected.