Physical AI Data and Training
Last reviewed: September 2026 | This area changes quickly and is subject to quarterly review.
Overview
Section titled “Overview”This document covers the front end of the Physical AI pipeline — from data entering the physical world to becoming a training asset. For an overview of the full pipeline see Physical AI Overview; for the back end that deploys and operates trained models see Deploy and Operate.
Layer 1 — Edge Inference and IoT
Section titled “Layer 1 — Edge Inference and IoT”Data from the physical world is high-volume and real-time, making it impractical to send everything to the cloud. The baseline structure is to infer first on site (at the edge) and upload only what is needed to the cloud.
| Item | AWS | Azure | Google Cloud | OCI |
|---|---|---|---|---|
| Edge runtime | IoT Greengrass | Azure IoT Operations / IoT Edge | Google Distributed Cloud (Edge) | Roving Edge Infrastructure |
| Edge ML inference | Greengrass ML components (SageMaker AI model deployment) | IoT Edge modules + Azure AI services | Edge TPU / Coral (verify current support status) | Build your own on RED compute |
| Industrial data ingestion | IoT SiteWise (OPC UA) | IoT Operations (OPC UA) | — (partner or self-built) | — (self-built) |
Edge Inference Hardware
Section titled “Edge Inference Hardware”If the edge runtime is the software layer, the choice of accelerator hardware beneath it determines feasibility and TCO. There are four criteria — the compute performance to run the target model, the memory capacity the model must fit into, the robot’s power and thermal budget, and the lifespan of the software ecosystem.
| Family | Characteristics | Considerations |
|---|---|---|
| Robotics-oriented edge modules (for example the NVIDIA Jetson family) | Spans low-power compact modules to high-performance ones, and is a frequently considered option for on-device robotics VLA inference | Performance and memory differ greatly across generations and modules, with wide price spreads. Confirm first that the target model fits in that module’s memory |
| General-purpose CPUs with integrated NPUs and small accelerators | Sufficient for lightweight vision such as classification and detection, with low power and unit cost | Often short on memory and bandwidth for large multimodal and VLA inference |
| FPGAs and industrial SoCs | Favorable for equipment where deterministic latency and long-term supply guarantees matter | High development difficulty and significant model porting cost |
| Cloud provider edge appliances | Extend cloud operational tooling and management to the field | Intended as field servers and gateways; they do not replace real-time control onboard the robot |
Sensor Data Pipeline — Ingestion, Storage, Labeling
Section titled “Sensor Data Pipeline — Ingestion, Storage, Labeling”Data arriving from the edge cannot be used for training as-is. Physical AI data is multimodal — 1D time series (joint angles, current, temperature), 2D video, 3D point clouds, and equipment metadata mixed together — so it must pass through ingestion, storage, and refinement/labeling before entering the training pipeline.
| Stage | AWS | Azure | Google Cloud | OCI |
|---|---|---|---|---|
| Stream and video ingestion | IoT Core (edge ingress) / Kinesis Video Streams and Data Streams (downstream landing) | Event Hubs + IoT Operations | Pub/Sub | OCI Streaming |
| Data lake | S3 | Data Lake Storage | Cloud Storage | Object Storage |
| Parallel file system for training | FSx for Lustre | Azure Managed Lustre | Managed Lustre / Parallelstore | File Storage with Lustre |
| Labeling | SageMaker Ground Truth (closed to new customers) | Azure ML data labeling | — (managed service ended; partner or open source) | OCI Data Labeling |
Layer 2 — Digital Twins and Simulation
Section titled “Layer 2 — Digital Twins and Simulation”Training robots and vehicles only in the real world is costly, risky, and slow. This is why the sim-to-real approach took hold: generate and train on large volumes of scenarios in a digital twin and simulation that virtually replicate the physical environment, then transfer to reality.
| Item | AWS | Azure | Google Cloud | OCI | Cross-vendor |
|---|---|---|---|---|---|
| Digital twin | IoT TwinMaker | Azure Digital Twins | — (self-built with Spanner Graph, BigQuery, etc.) | — | NVIDIA Omniverse |
| Robot and physics simulation | — (RoboMaker end of support; self-built) | — (partner or self-built) | — (partner or self-built) | — | NVIDIA Isaac Sim / Isaac Lab |
Why Training Data Is Scarce
Section titled “Why Training Data Is Scarce”Simulation is a premise rather than an option because of data scarcity. Language models started from effectively unlimited pre-training data in the form of internet text, but data of “a robot actually picking up and moving an object” is orders of magnitude smaller. Physical AI training starts from a design that fills the gap with cheaper, more plentiful data.
| Data tier | Characteristics | Limitation |
|---|---|---|
| Internet video, images, text | Effectively unlimited and cheap. Provides general knowledge about the world and objects | Not directly connected to a robot’s body or joint commands |
| Human task video (egocentric) | Relatively plentiful. Provides hints about motion sequence and intent | Framed around a human body, so it does not transfer directly to a robot embodiment |
| Robot teleop episodes | Most accurate, containing actual joint commands | A human must directly operate the robot to produce it, making collection the most expensive |
Teleoperation is the practice of a person moving a robot directly with a controller or VR rig to produce demonstration data, and imitation learning trains a model to reproduce that data. It is the most reliable method, but human time is the cost, which makes it hard to scale.
Open Datasets and Benchmark Ecosystem
Section titled “Open Datasets and Benchmark Ecosystem”Because data scarcity is hard for any single organization to overcome alone, an open ecosystem has formed in which multiple institutions pool data and share evaluation tasks. When choosing a stack, “which open assets can be used as-is on this stack” becomes a practical criterion for judging lock-in.
| Category | Representative assets | What it is used for |
|---|---|---|
| Cross-robot datasets | Open X-Embodiment (paper) | A collection consolidating robot data from many institutions. The baseline for cross-embodiment pre-training |
| Large manipulation datasets | DROID | Teleop data collected across diverse environments |
| Simulation benchmarks | LIBERO, CALVIN, RoboCasa, Meta-World | Compare policies by task success rate on standard task suites |
| Human task video | Ego4D | Egocentric task video. Corresponds to the middle tier of the three data tiers above |
| Open toolchains | LeRobot | An open source stack bundling data format, training, and evaluation |
Sim-to-Real Gap
Section titled “Sim-to-Real Gap”The value of simulation is clear, but a simulator’s physics, sensor, and material models differ subtly from reality. Differences in friction coefficients, lighting, sensor noise, and mechanical play accumulate, producing the phenomenon where a policy that succeeded in simulation fails on the real robot. This is the sim-to-real gap.
Three mitigations are common.
- Domain randomization — Randomly vary physical parameters such as friction, mass, lighting, and texture during training so that reality falls inside that distribution.
- Small-scale fine-tuning on real data — Finish a simulation-pretrained policy with a small amount of data from the actual robot.
- Physical validation gate — Use success rates on real equipment as the deployment criterion, separately from simulation success rates.
Matching Training Infrastructure to Scale
Section titled “Matching Training Infrastructure to Scale”A common misconception in Physical AI is that “training robot models always requires a large GPU cluster.” In practice, most work is fine-tuning a pre-trained robot foundation model to your own robots and tasks, and this phase is nothing like LLM pre-training in scale. Small-scale PEFT and adapter training often completes as a short job on a single GPU.
| Stage | Nature of the work | Infrastructure pattern | Cost strategy |
|---|---|---|---|
| Initial validation | Small demonstration dataset, LoRA/PEFT adapter training | A single GPU instance | Spot/preemptible instances as the default. Interruption costs little to restart |
| Task specialization | Medium demonstration dataset, full fine-tuning | Single node, multiple GPUs plus a managed training job | A managed training service with automatic checkpoint and resume |
| Platformization | Many robots and tasks, repeated retraining | Multiple nodes plus a high-speed interconnect | Reserved or committed discounts. Requires automatic node recovery (see Distributed Training) |
Common Mistakes
Section titled “Common Mistakes”- Sending all data to the cloud — A design that ignores latency, bandwidth, and cost fails at real-time control. Dividing inference to the edge comes first.
- Using EOL products in new designs — Do not adopt services that have ended or curtailed support — RoboMaker, Percept, managed labeling services — based only on older material.
- Locking into a single vendor’s simulator — Tying the training pipeline to one cloud’s proprietary simulation makes porting and comparison difficult.
- Estimating TCO from GPU cost alone — Data collection, storage, and labeling, plus simulation occupancy cost, arise separately from training compute.
- Using simulation success rates as deployment evidence — Deploying without a physical rollout validation gate lets the sim-to-real gap surface in the field.
Checklist
Section titled “Checklist”Data and Cost
Section titled “Data and Cost”- Have you designed the storage, refinement, and labeling layer for sensor data, and confirmed that labeling tooling is replaceable?
- Have you estimated data collection, storage, labeling, and simulation costs separately from GPU cost?
- Have you checked the licenses of the open datasets and benchmarks you plan to use, and whether they cover your own embodiment?
Hardware and Training
Section titled “Hardware and Training”- Did you select the edge accelerator by measuring the target model rather than by TOPS figures, and confirm the support lifespan of its drivers and runtime?
- Have you chosen an infrastructure stage matched to your training scale (without over-committing during initial validation)?
- Is the digital twin and simulation stack portable to another cloud (lock-in check)?
- Are the IoT and robotics services you plan to use currently supported (EOL check)?
Related Documents
Section titled “Related Documents”- Physical AI Overview — Full pipeline, layered structure, and open problems
- Deploy and Operate — Robotics foundation models, safety, architecture, fleet deployment
- Block and File Storage — Parallel file system comparison
- GPU Infrastructure — GPU clusters for cloud training and simulation
- Distributed Training — Multi-node training and checkpoint strategy
Further Reading
Section titled “Further Reading”- AWS IoT Greengrass Developer Guide
- AWS IoT TwinMaker
- AWS IoT SiteWise
- Amazon FSx for Lustre
- Amazon SageMaker Ground Truth documentation
- Azure IoT Operations documentation
- Azure Digital Twins documentation
- Azure Managed Lustre
- Azure Machine Learning data labeling