Robot Data Collection (the data-supply stack)

last updated 2026-06-13
Vision-Language-Action (VLA) ModelsVision-Language-Act…World Models (for robotics & autonomy)World Models (for r…Sim-to-Real, Robot Simulation & Synthetic DataSim-to-Real, Robot …Tactile Sensing & Electronic SkinTactile Sensing & E…Dexterous Manipulation & Robot HandsDexterous Manipulat…Robot Dat…

The substrate under Physical AI (robotics cluster hub): how you actually MANUFACTURE the training data for robot foundation/world models, since there is no internet-of-robot-actions to scrape. This is the picks-and-shovels layer; the investability router is Robot Data Supply Stack.

It’s a stack, not a versus

The central 2026 reframe: it is not teleop vs sim vs video. Each source supplies a different property and you need all of them: teleop = fidelity (native, on-distribution, expensive), simulation = scale (cheap, but contact-physics-poor), human video = diversity (internet-scale priors, but the embodiment gap), world models = counterfactuals (dreamed rollouts, but physically inconsistent), tactile = the missing modality (force/slip/contact, absent from all the others). The investment question (→ Robot Data Supply Stack) is which layer has a defensible wedge.

The nine data sources

  1. Teleoperation capture — leader-follower rigs (ALOHA ~$20-32k; Mobile ALOHA +base). Gold-standard data; cost curve ~$340/hr (2024) → ~$118/hr (2026), but ~300-1,200 demos/task = $50-150k. The 1:many operator-scaling software (PATO) is the prize. See Teleoperation Bridge.
  2. Handheld / wearable capture — UMI (handheld gripper + GoPro, no robot needed), DexUMI (hand exoskeleton for dexterous capture), data gloves (SenseGlove R1). The cheap-data frontier, but the methodology is open-sourced from academia faster than a hardware moat forms.
  3. Egocentric human video — Ego4D, Apple EgoDex (829 hrs, hand pose), EgoScaler (20,854 hrs, hinted log-linear scaling law); captured on Meta Aria Gen 2. Diversity at scale, but human→robot transfer is bounded by the embodiment gap (works for intent/navigation, not yet reliable for manipulation).
  4. Simulation & synthetic — Isaac Sim/Lab, Genesis; vendor layer Lightwheel (SimReady assets). NVIDIA-adjacent. See Sim-to-Real, Robot Simulation & Synthetic Data.
  5. World-model-generated — NVIDIA Cosmos / GR00T-Dreams, 1X world model, DeepMind Genie 3. Dreamed rollouts, gated by physical consistency (grasping/contact unreliable). See World Models (for robotics & autonomy).
  6. Real-world fleet self-reinforcing loop (more usage gives more data, which improves the product and drives more usage) — Figure (watch humans), 1X (embody robots, teleop in homes), Tesla (simulate), Neura (build a gym). Most claimed, fewest real: nobody has a proven manipulation self-reinforcing loop yet (Tesla/Waymo have it for driving).
  7. Tactile / contact capture — the under-built modality; force/slip/shear during teleop or via instrumented grippers/skins. Where the 2026 capital is rushing (PaXini, DAIMON). See Tactile Sensing & Electronic Skin.
  8. Sensors as substrate — tactile (binding) > force-torque > event cameras (niche-binding) > depth/LiDAR/RGB/proprioception (commodity). Tactile is the one binding constraint on data quality.
  9. Data-as-a-service / data engine — Encord (the multimodal platform, the venture-scale layer) vs annotation services (Objectways, the commodity/services trap).

Why it matters

Connections

Physical AI (robotics cluster hub) · Robot Data Supply Stack · Teleoperation Bridge · Robot Foundation Models · World Models (for robotics & autonomy) · Sim-to-Real, Robot Simulation & Synthetic Data · Tactile Sensing & Electronic Skin · Tactile Sensing Silicon · Robot Autonomy Destination

Related concepts

Frontier questions