Skip to content

Physical AI Data and Training

Last reviewed: September 2026 | This area changes quickly and is subject to quarterly review.

This document covers the front end of the Physical AI pipeline — from data entering the physical world to becoming a training asset. For an overview of the full pipeline see Physical AI Overview; for the back end that deploys and operates trained models see Deploy and Operate.

Data from the physical world is high-volume and real-time, making it impractical to send everything to the cloud. The baseline structure is to infer first on site (at the edge) and upload only what is needed to the cloud.

Item AWS Azure Google Cloud OCI
Edge runtime IoT Greengrass Azure IoT Operations / IoT Edge Google Distributed Cloud (Edge) Roving Edge Infrastructure
Edge ML inference Greengrass ML components (SageMaker AI model deployment) IoT Edge modules + Azure AI services Edge TPU / Coral (verify current support status) Build your own on RED compute
Industrial data ingestion IoT SiteWise (OPC UA) IoT Operations (OPC UA) — (partner or self-built) — (self-built)

If the edge runtime is the software layer, the choice of accelerator hardware beneath it determines feasibility and TCO. There are four criteria — the compute performance to run the target model, the memory capacity the model must fit into, the robot’s power and thermal budget, and the lifespan of the software ecosystem.

Family Characteristics Considerations
Robotics-oriented edge modules (for example the NVIDIA Jetson family) Spans low-power compact modules to high-performance ones, and is a frequently considered option for on-device robotics VLA inference Performance and memory differ greatly across generations and modules, with wide price spreads. Confirm first that the target model fits in that module’s memory
General-purpose CPUs with integrated NPUs and small accelerators Sufficient for lightweight vision such as classification and detection, with low power and unit cost Often short on memory and bandwidth for large multimodal and VLA inference
FPGAs and industrial SoCs Favorable for equipment where deterministic latency and long-term supply guarantees matter High development difficulty and significant model porting cost
Cloud provider edge appliances Extend cloud operational tooling and management to the field Intended as field servers and gateways; they do not replace real-time control onboard the robot

Sensor Data Pipeline — Ingestion, Storage, Labeling

Section titled “Sensor Data Pipeline — Ingestion, Storage, Labeling”

Data arriving from the edge cannot be used for training as-is. Physical AI data is multimodal — 1D time series (joint angles, current, temperature), 2D video, 3D point clouds, and equipment metadata mixed together — so it must pass through ingestion, storage, and refinement/labeling before entering the training pipeline.

Stage AWS Azure Google Cloud OCI
Stream and video ingestion IoT Core (edge ingress) / Kinesis Video Streams and Data Streams (downstream landing) Event Hubs + IoT Operations Pub/Sub OCI Streaming
Data lake S3 Data Lake Storage Cloud Storage Object Storage
Parallel file system for training FSx for Lustre Azure Managed Lustre Managed Lustre / Parallelstore File Storage with Lustre
Labeling SageMaker Ground Truth (closed to new customers) Azure ML data labeling — (managed service ended; partner or open source) OCI Data Labeling

Training robots and vehicles only in the real world is costly, risky, and slow. This is why the sim-to-real approach took hold: generate and train on large volumes of scenarios in a digital twin and simulation that virtually replicate the physical environment, then transfer to reality.

Item AWS Azure Google Cloud OCI Cross-vendor
Digital twin IoT TwinMaker Azure Digital Twins — (self-built with Spanner Graph, BigQuery, etc.) — NVIDIA Omniverse
Robot and physics simulation — (RoboMaker end of support; self-built) — (partner or self-built) — (partner or self-built) — NVIDIA Isaac Sim / Isaac Lab

Simulation is a premise rather than an option because of data scarcity. Language models started from effectively unlimited pre-training data in the form of internet text, but data of “a robot actually picking up and moving an object” is orders of magnitude smaller. Physical AI training starts from a design that fills the gap with cheaper, more plentiful data.

Data tier Characteristics Limitation
Internet video, images, text Effectively unlimited and cheap. Provides general knowledge about the world and objects Not directly connected to a robot’s body or joint commands
Human task video (egocentric) Relatively plentiful. Provides hints about motion sequence and intent Framed around a human body, so it does not transfer directly to a robot embodiment
Robot teleop episodes Most accurate, containing actual joint commands A human must directly operate the robot to produce it, making collection the most expensive

Teleoperation is the practice of a person moving a robot directly with a controller or VR rig to produce demonstration data, and imitation learning trains a model to reproduce that data. It is the most reliable method, but human time is the cost, which makes it hard to scale.

Because data scarcity is hard for any single organization to overcome alone, an open ecosystem has formed in which multiple institutions pool data and share evaluation tasks. When choosing a stack, “which open assets can be used as-is on this stack” becomes a practical criterion for judging lock-in.

Category Representative assets What it is used for
Cross-robot datasets Open X-Embodiment (paper) A collection consolidating robot data from many institutions. The baseline for cross-embodiment pre-training
Large manipulation datasets DROID Teleop data collected across diverse environments
Simulation benchmarks LIBERO, CALVIN, RoboCasa, Meta-World Compare policies by task success rate on standard task suites
Human task video Ego4D Egocentric task video. Corresponds to the middle tier of the three data tiers above
Open toolchains LeRobot An open source stack bundling data format, training, and evaluation

The value of simulation is clear, but a simulator’s physics, sensor, and material models differ subtly from reality. Differences in friction coefficients, lighting, sensor noise, and mechanical play accumulate, producing the phenomenon where a policy that succeeded in simulation fails on the real robot. This is the sim-to-real gap.

Three mitigations are common.

  • Domain randomization — Randomly vary physical parameters such as friction, mass, lighting, and texture during training so that reality falls inside that distribution.
  • Small-scale fine-tuning on real data — Finish a simulation-pretrained policy with a small amount of data from the actual robot.
  • Physical validation gate — Use success rates on real equipment as the deployment criterion, separately from simulation success rates.

A common misconception in Physical AI is that “training robot models always requires a large GPU cluster.” In practice, most work is fine-tuning a pre-trained robot foundation model to your own robots and tasks, and this phase is nothing like LLM pre-training in scale. Small-scale PEFT and adapter training often completes as a short job on a single GPU.

Stage Nature of the work Infrastructure pattern Cost strategy
Initial validation Small demonstration dataset, LoRA/PEFT adapter training A single GPU instance Spot/preemptible instances as the default. Interruption costs little to restart
Task specialization Medium demonstration dataset, full fine-tuning Single node, multiple GPUs plus a managed training job A managed training service with automatic checkpoint and resume
Platformization Many robots and tasks, repeated retraining Multiple nodes plus a high-speed interconnect Reserved or committed discounts. Requires automatic node recovery (see Distributed Training)
  • Sending all data to the cloud — A design that ignores latency, bandwidth, and cost fails at real-time control. Dividing inference to the edge comes first.
  • Using EOL products in new designs — Do not adopt services that have ended or curtailed support — RoboMaker, Percept, managed labeling services — based only on older material.
  • Locking into a single vendor’s simulator — Tying the training pipeline to one cloud’s proprietary simulation makes porting and comparison difficult.
  • Estimating TCO from GPU cost alone — Data collection, storage, and labeling, plus simulation occupancy cost, arise separately from training compute.
  • Using simulation success rates as deployment evidence — Deploying without a physical rollout validation gate lets the sim-to-real gap surface in the field.
  • Have you designed the storage, refinement, and labeling layer for sensor data, and confirmed that labeling tooling is replaceable?
  • Have you estimated data collection, storage, labeling, and simulation costs separately from GPU cost?
  • Have you checked the licenses of the open datasets and benchmarks you plan to use, and whether they cover your own embodiment?
  • Did you select the edge accelerator by measuring the target model rather than by TOPS figures, and confirm the support lifespan of its drivers and runtime?
  • Have you chosen an infrastructure stage matched to your training scale (without over-committing during initial validation)?
  • Is the digital twin and simulation stack portable to another cloud (lock-in check)?
  • Are the IoT and robotics services you plan to use currently supported (EOL check)?