Physical AI Deploy and Operate
Last reviewed: September 2026 | This area changes quickly and is subject to quarterly review.
Overview
Section titled “Overview”This document covers the back end of the Physical AI pipeline — from deploying a trained model into the physical world to operating it safely. For an overview of the full pipeline see Physical AI Overview; for the data and training layers see Data and Training.
Layer 3 — Robotics Foundation Models
Section titled “Layer 3 — Robotics Foundation Models”Basic Terms — Policy and Embodiment
Section titled “Basic Terms — Policy and Embodiment”The function that decides “what to do in this situation” is called a policy. Taking an observation — camera images, range sensors, joint angles — as input and producing an action — joint commands, motion commands — as output is the smallest unit of robot intelligence. Every model discussed below is a question of how to construct that policy.
Robot bodies differ. A manipulator arm, a quadruped, a humanoid, and an autonomous mobile robot (AMR) all differ in degrees of freedom and control method, and these different bodies are called embodiments. Transferring a policy learned on one body to another — cross-embodiment transfer — is a well-known unsolved problem in the field.
Just as LLMs generalized language, robot foundation models that aim to generalize robot perception, planning, and motion are emerging. The representative approach is VLA (Vision-Language-Action), which takes natural language instructions and connects vision, language, and action.
| Item | Status |
|---|---|
| Representative stack | NVIDIA Isaac GR00T — an open foundation model (VLA) for robots, with Omniverse/Cosmos-based simulation and synthetic data, and Jetson Thor on-device inference |
| Major clouds | Their own general-purpose robot foundation models remain limited — they generally run the NVIDIA stack on GPU infrastructure or offer it through partnerships |
| National policy | Japan adopted robotics foundation model development as a national project under GENIAC (see Japan AI Landscape) |
Connecting the Physical World and Agents
Section titled “Connecting the Physical World and Agents”If a robot foundation model handles perception, planning, and motion, the agent is the autonomous execution layer above it that takes a goal, plans steps on its own, and calls tools, sensors, and actuators to execute them. In Physical AI, unlike digital agents, actions take effect immediately in the physical world, so the connection method and permission boundaries are directly tied to safety.
- Edge agent vs. cloud orchestration — It is common to divide responsibility so that real-time judgment and control loops run autonomously on site (at the edge), while long-horizon planning, multi-robot coordination, and model updates are handled in the cloud. Even if the network is severed, the edge agent must be able to continue operating safely or stop safely.
- Tool and actuator connectivity (MCP and similar) — For an agent to read sensor values and issue higher-level tasks, a standardized connection layer is needed. However, a protocol such as MCP is a high-level task instruction and tool-call layer; real-time actuator control (motors, joints, and so on) is handled by a separate deterministic low-level control layer (fieldbus, robot middleware, and so on) with guaranteed latency and safety. The two must not be conflated. For general concepts of autonomous execution and tool calling, see AI Agents; for agent-tool integration protocols, see AI Agent Integration (MCP).
Why the Layers Are Separated — Control Frequency Mismatch
Section titled “Why the Layers Are Separated — Control Frequency Mismatch”This separation is not a design preference but a consequence of physically different operating frequencies. The low-level control loop holding motors and joints must run deterministically at sub-millisecond periods, while inference by a large multimodal model is far slower with much greater latency variance. The common structure therefore splits into two layers.
- Slow layer (understanding and planning) — Understands the scene and sets the next goal. It can be heavy and run at a slow period. It may even sit in the cloud.
- Fast layer (execution and control) — Converts a given goal into actual joint trajectories and executes them at a fixed period. It must be on site, and its latency must not fluctuate.
Agent-to-Hardware Connection Standards
Section titled “Agent-to-Hardware Connection Standards”After MCP established itself as the standard connecting agents to data and software tools, standards connecting agents to physical equipment have begun to appear as well. In August 2026 Anthropic published a research preview of the Model Hardware Standard (MHS). Instead of building a custom adapter per device, it exposes devices through a common driver so that a single agent can operate multiple devices in parallel. Safety limits are enforced at the driver level, below the agent, so a model cannot talk its way past a hard limit.
Safety Layer — Autonomous Driving and Robotics
Section titled “Safety Layer — Autonomous Driving and Robotics”AI that moves in the physical world is directly tied to human life and equipment, making functional safety central. Domain-specific safety standards and certification regimes apply separately — ISO 26262 for autonomous driving, ISO 13849 and IEC 61508 for industrial machinery and robots — and the principle is to maintain an independent safety layer (safety stop, hardware interlocks, safety PLC) that operates regardless of the AI model’s judgment.
Vendors and suppliers offer commercial stacks implementing this. For example, NVIDIA offers the safety system Halos.
- Autonomous vehicles (AV): the DRIVE platform (AGX, Hyperion) with the Halos safety system (spanning cloud to vehicle, targeting ISO 26262), with simulation via Omniverse and Cosmos.
- Robotics: In June 2026 NVIDIA announced Halos for Robotics (IGX Thor, Holoscan Sensor Bridge, Halos OS, AI Systems Inspection Lab), extending its autonomous driving safety foundation to industrial robots, humanoids, and AMRs.
Multicloud and Edge Architecture Considerations
Section titled “Multicloud and Edge Architecture Considerations”What Goes at the Edge, What Goes in the Cloud
Section titled “What Goes at the Edge, What Goes in the Cloud”Physical AI design starts by deciding where each task belongs — edge or cloud. The criteria are latency sensitivity, data volume (bandwidth), safety requirements, and behavior when the network is severed.
| Task | Primary location | Reason |
|---|---|---|
| Real-time perception and control loop | Edge | Latency-sensitive and must not stop even when the network is severed |
| Safety stop and emergency shutdown | Edge | Cannot tolerate a cloud round trip |
| First-pass sensor data filtering and aggregation | Edge | Uploading all raw data costs too much bandwidth and money |
| Data storage and labeling | Cloud | Aggregate data from many machines and manage it as a training asset |
| Model training and retraining | Cloud | Requires large-scale GPUs and datasets (see GPU Infrastructure) |
| Synthetic data generation and simulation | Cloud | Digital twins and simulators require large-scale compute |
| Multi-robot fleet coordination, long-horizon planning | Cloud | Global coordination beyond an individual edge’s field of view |
| Model version management and deployment (OTA) | Cloud → Edge | Managed centrally and distributed to the field |
The Closed-Loop Operating Cycle
Section titled “The Closed-Loop Operating Cycle”Physical AI is not deployed once and finished; it operates as a cycle in which field data returns to the model. Each stage of the pipeline flow diagram in the overview corresponds to the following operating cycle.
- Edge inference (
Edge Inferencein the diagram) — Perceive, decide, and control in real time on site, selecting only meaningful events and anomalous data. - Telemetry collection and refinement (
Telemetry→Data Lake, Labeling) — Upload the selected data and operation logs to the cloud, store them, and refine and label them for training use. - Cloud retraining and simulation (
Cloud Training, Model Management↔Simulation, Digital Twin) — Improve the model with collected data and validate new scenarios in the digital twin and simulation. - OTA deployment (
Deploy→Edge Inference) — Deploy validated models and policies back to the edge. For general patterns such as signing and rollback against deployment failure or regression, see Hybrid and Edge Computing.
What Determines Whether a Model Passes
Section titled “What Determines Whether a Model Passes”To run the closed loop, you need a criterion for “is this model fit to ship.” In Physical AI, however, a lower training loss does not mean a higher real success rate. A model trained to faithfully reproduce demonstration data cannot recover once it falls into a state absent from those demonstrations.
The criterion must therefore be the success rate of rollouts carried through to completion, not a loss value. In practice, three tiers are recorded separately.
| Evaluation tier | What it measures | Limitation |
|---|---|---|
| Offline metrics | Loss and prediction accuracy on validation data | Weakly correlated with success rate. Use for regression detection only |
| Simulation rollouts | Success rate completing the task in the simulator | Diverges from reality by the size of the sim-to-real gap |
| Physical rollouts | Success rate, intervention count, and recovery time on real equipment | Most trustworthy but most expensive |
Fleet Deployment — Abort and Rollback Are Different Layers
Section titled “Fleet Deployment — Abort and Rollback Are Different Layers”Loading a model onto one robot and deploying to a fleet of thousands to tens of thousands are different problems. At fleet scale, the design hinges on how fast a bad model spreads and whether it can be reversed once spread.
- Staged rollout — Deploy to a small group first and expand, halting expansion if the failure rate exceeds a threshold.
- Abort — The mechanism that halts expansion. In many fleet OTA services, however, abort cancels only targets that have not yet started and deployments already in progress run to completion, so confirm the scope of your service’s abort behavior in the vendor’s documentation.
- Rollback — The mechanism that returns devices that already received the new version to their previous state. It is a separate layer from abort and works only if the device side has a recovery path such as previous-version retention or an A/B partition.
Other Considerations
Section titled “Other Considerations”- Data gravity and latency — Sensor data is high-volume and latency-sensitive, so dividing work between on-site edge inference and cloud training is the baseline design. Decide first what is processed at the edge and what is uploaded.
- Simulator portability — If digital twins and simulation are tied to one cloud’s proprietary service, porting becomes difficult. Prioritizing stacks such as NVIDIA Omniverse and Isaac that run anywhere given a GPU reduces lock-in.
- On-device vs. cloud training split — It is common to split training and synthetic data generation to cloud GPUs and real-time inference to on-device hardware (for example the Jetson family).
- Safety and regulation — Autonomous driving and industrial robots are subject to separate functional safety certification and regulation. Reflect certification requirements early in the architecture.
- Check product lifecycle — This area has many retired (EOL) products (for example Azure Percept, AWS RoboMaker, managed labeling services). Always confirm each service’s current support status before designing.
Common Mistakes
Section titled “Common Mistakes”- Bolting on safety later — Autonomous driving and robots must design safety in from the start (“built-in,” not “bolt-on”).
- Granting physical agents unlimited permissions — Letting an autonomous agent call actuators without constraints turns a malfunction directly into physical harm. Action-space restriction and safety-layer validation are essential.
- Setting abort criteria without designing a rollback path — Spread stops, but devices already deployed to do not come back.
- Placing a large model directly in the real-time control loop — Failing to meet the control period leads to safety problems. Separate the slow planning layer from the fast control layer.
- Using simulation success rates as deployment evidence — Deploying without a physical rollout validation gate lets the sim-to-real gap surface in the field.
Checklist
Section titled “Checklist”Edge and Cloud Placement
Section titled “Edge and Cloud Placement”- Have you distinguished inference to process at the edge from data to upload to the cloud?
- Does the edge operate autonomously and safely (offline) when the network is severed?
- Have you defined how many milliseconds each decision must complete within, and separated the slow planning layer from the fast control layer?
- Is the division of roles between on-device inference and cloud training clear?
Safety and Regulation
Section titled “Safety and Regulation”- If agents call physical actuators, have you restricted the action space and added safety-layer validation?
- For autonomous driving or industrial robots, have you reflected functional safety certification requirements in the design?
Deployment and Operations
Section titled “Deployment and Operations”- Have you defined the model pass criterion as rollout success rate rather than a loss value, with a physical validation gate?
- Have you designed both a staged rollout and a rollback path for fleet deployment (rather than assuming abort is sufficient)?
- Does your target fleet size fit within each vendor’s non-adjustable service quotas?
Related Documents
Section titled “Related Documents”- Physical AI Overview — Full pipeline, layered structure, and open problems
- Data and Training — Edge, data pipeline, simulation, training infrastructure
- AI Agents — Autonomous planning and execution concepts
- AI Agent Integration (MCP) — Agent-to-tool and system integration protocols
- GPU Infrastructure — GPU clusters for cloud training and simulation
- LLMOps — General model evaluation and operations
- Japan AI Landscape — Robotics foundation models as a national project (GENIAC)