GPU Workload Characteristics and Reference Architecture
Last reviewed: September 2026 | This is a fast-moving area subject to quarterly review.
Overview
Section titled “Overview”You need this document when a model or its data exceeds the capacity of a single GPU (or a single server). In that case, you have to combine multiple GPUs and multiple servers into one and split the work across them.
A common misconception here is “faster GPUs and more of them makes it that much faster.” It doesn’t work that way. For multiple GPUs to collaborate, they must constantly exchange calculation results, and if you add GPUs without also increasing inter-node communication bandwidth and storage throughput, only GPU idle time grows and performance does not scale linearly. (Just as adding cooks doesn’t speed things up if the kitchen is cramped and the aisles for carrying ingredients are jammed.)
So GPU infrastructure design is less about “which GPU” and more about how you connect the GPUs and how you move the data. This document first classifies what kind of workload you have (training vs. inference, etc.), then compares four vendors across a connection structure split into three tiers: intra-node → inter-node → storage.
What stays the same across clouds, and what differs per vendor
Section titled “What stays the same across clouds, and what differs per vendor”Sorting out what carries over versus what you must relearn when switching clouds makes the rest easier.
- What carries over (portability baseline) — Training/inference code mostly runs on the common foundation of CUDA and NCCL. (CUDA = NVIDIA’s standard software for running computation on GPUs; NCCL = the library that lets multiple GPUs exchange results.) Code written on top of these two generally ports across clouds.
- What differs per vendor — The high-speed network physically connecting the GPUs, the way servers are placed close together, and the fully managed cluster products differ in name and implementation by vendor, and cannot be swapped one-to-one.
Three GPU Workload Classes
Section titled “Three GPU Workload Classes”Because the bottleneck resource differs per workload, the starting point of infrastructure design is understanding workload characteristics.
| Workload | Dominant bottleneck | Communication needs | Storage needs | Representative infrastructure traits |
|---|---|---|---|---|
| Pre-training | Compute + inter-node communication | Very high (all nodes communicate together) | High (streaming large datasets) | Many nodes, high-speed network required, checkpoint bandwidth critical |
| Fine-tuning | Compute + memory | Medium (often within a few nodes) | Medium | Small-to-medium cluster, often feasible on a single node |
| Inference | Memory bandwidth + latency | Low (only when model-parallel) | Low (weights resident after load) | Latency/throughput balance, autoscaling-centric |
Reference Architecture — Three-Tier Communication Model
Section titled “Reference Architecture — Three-Tier Communication Model”There are roughly three kinds of paths data travels in a GPU cluster. Each path is built with different technology and differs in both speed and role. When comparing vendors, mixing these three tiers leads to wrong comparisons, so always separate them.
graph TB
subgraph Node["Single node (8x GPU)"]
G1["GPU"] -->|"NVLink / NVSwitch<br/>(intra-node)"| G2["GPU"]
end
Node -->|"High-speed fabric<br/>(inter-node RDMA)"| Node2["Other node"]
Node -->|"Parallel filesystem / object storage<br/>(data & checkpoints)"| Storage["Storage tier"]
- Intra-node (within one server) — GPUs inside one server are joined by ultra-fast dedicated links called NVLink/NVSwitch. This is the fastest of the three tiers and is determined by NVIDIA hardware characteristics regardless of vendor.
- Inter-node (server to server) — Servers are connected by a high-speed fabric. Here, a fabric means “a dedicated high-speed network that tightly weaves servers together,” and RDMA (Remote Direct Memory Access) is the technology that “exchanges data directly between server memories without going through the CPU.” This tier is implemented differently by each vendor and governs large-scale training performance.
- Storage (data store) — The path for reading training data and writing/reloading intermediate saves (checkpoints). If this path is slow, GPUs sit idle waiting for data.
Inter-node High-Speed Fabric — Vendor Mapping
Section titled “Inter-node High-Speed Fabric — Vendor Mapping”| Tier | AWS | Azure | Google Cloud | OCI |
|---|---|---|---|---|
| Inter-node fabric | EFA (Elastic Fabric Adapter) | InfiniBand (ND series) | GPUDirect-TCPX / RDMA | RDMA Cluster Network |
| Communication library | NCCL | NCCL | NCCL | NCCL |
| Same concept? | Approximate — name, implementation, and performance characteristics differ | Approximate | Approximate | Approximate |
Placement & Topology — Physical Proximity and NUMA
Section titled “Placement & Topology — Physical Proximity and NUMA”Two things must line up to realize inter-node communication performance. First, the GPU servers must be physically close within the data center (if scattered far apart, the round trip takes longer). Second, within one server, the GPU and the network card (NIC) it uses must sit in the same zone. Here, NUMA (Non-Uniform Memory Access) refers to “a structure where even within one server the CPU/memory is split into zones, so same-zone access is fast and crossing to another zone is slow.”
| Item | AWS | Azure | Google Cloud | OCI |
|---|---|---|---|---|
| Proximity placement | Placement Group (Cluster) | Proximity Placement Group + VMSS | Compact Placement Policy | Cluster Network (built-in proximity provisioning) |
| NUMA/GPU-NIC alignment | Instance topology exposed, NCCL topology awareness | Topology exposed | gVNIC + topology awareness | Bare Metal topology pinning |
- Proximity placement — To bind training nodes at low latency, you must explicitly request proximity placement. Nodes scattered without a placement group incur higher collective-communication latency, reducing training throughput.
- NUMA/GPU-NIC affinity — If a GPU and the NIC it uses reside on different NUMA nodes, data traverses the inter-socket link, causing latency and bandwidth loss. Align them with NCCL topology awareness and process binding.
Managed GPU Clusters
Section titled “Managed GPU Clusters”Instead of assembling nodes, fabric, and scheduler yourself, using a vendor-provided managed GPU cluster gets topology, health checks, and restarts pre-integrated. For large-scale training, this is a first-class option.
| Vendor | Managed cluster product | Characteristics |
|---|---|---|
| AWS | SageMaker HyperPod | Node health checks/auto-replacement, checkpoint-based resume built in |
| Azure | CycleCloud + ND series | HPC/AI cluster orchestration, scheduler integration (auto node-replacement not built in — configured via scheduler/scripts) |
| Google Cloud | AI Hypercomputer / Cluster Director | Integrated infrastructure stack, topology-aware provisioning |
| OCI | Supercluster | RDMA cluster network, ultra-low-latency connectivity for large GPU counts (Bare Metal) |
Orchestrator Choice — Slurm vs Kubernetes
Section titled “Orchestrator Choice — Slurm vs Kubernetes”A fork you hit when picking a managed GPU cluster is which orchestrator distributes the work. There are broadly two paths — Slurm and Kubernetes — and they are not a superior/inferior substitution but options you pick by workload characteristics.
- Slurm — An open-source workload manager (job scheduler) long used in HPC (high-performance computing). A user submits a job (“give me N GPUs for M hours”), and Slurm places it in a priority-ordered queue (partition) and allocates it to nodes as they free up. Because it centers on batch jobs rather than containers, it has less friction for large-scale pre-training or when lifting existing on-prem HPC/Slurm jobs as-is, and you can reuse submission scripts and recipes.
- Kubernetes — The container-orchestration standard, stronger at long-running services (such as inference servers) than at batch jobs. It is advantageous when running training, inference, and serving mixed on one cluster, or when you need namespace isolation and multi-tenancy, and it reuses your existing Kubernetes ecosystem. Note that gang scheduling, which distributed training needs, must be added via separate tools (see GPU Kubernetes and Scheduling).
Cloud vendors often provide both Slurm and Kubernetes paths across their product portfolios, but whether both are available as choices within the same managed cluster varies by product. Some products also support a hybrid approach that bridges the two.
| Orchestrator | AWS | Azure | Google Cloud | OCI |
|---|---|---|---|---|
| Slurm (HPC) | HyperPod + Slurm (managed) | CycleCloud Workspace for Slurm (solution template deployed in the customer tenant) | Cluster Director (managed) · Cluster Toolkit (self-deployed) | HPC Cluster Stack + Slurm (self-deployed on GPU cluster networking) |
| Kubernetes | HyperPod + EKS / EKS | AKS | GKE | OKE |
| Hybrid | — | — | Cluster Director — Slurm on GKE (Preview) | — |
Related Documents
Section titled “Related Documents”- Next: Distributed training parallelism (TP/PP/DP) — Distributed Training Standard Architecture
- Kubernetes, scheduling, quotas — GPU Kubernetes and Scheduling
- Inference serving, failures, capacity, cost — Inference Serving, Reliability, and Cost
- Confidential GPU computing — Data Protection — Confidential Computing
Common Mistakes
Section titled “Common Mistakes”- Choosing the top GPU generation without workload analysis — Applying a large-scale pre-training configuration to a workload well served by fine-tuning/inference, spiking cost
- Multi-node training without a placement group — Nodes physically scattered, increasing collective-communication latency and degrading training throughput
- Substituting fabrics one-to-one across vendors — Assuming EFA, InfiniBand, and RDMA are identical, leading to mispredicted performance
- Under-provisioning storage bandwidth — GPUs idle while checkpoints are written, lowering effective utilization
Checklist
Section titled “Checklist”- Did you classify the workload as pre-training/fine-tuning/inference and identify the dominant bottleneck (compute/memory/communication)?
- Did you design the three tiers (intra-node/inter-node/storage) separately?
- Did you explicitly request proximity placement (placement group/cluster network) for multi-node training?
- Did you verify GPU-NIC NUMA alignment and NCCL topology awareness?
- Did you evaluate the portability/operational trade-off between managed clusters and self-built configuration?
References
Section titled “References”Google Cloud
Section titled “Google Cloud”- Cloud GPUs
- AI Hypercomputer
- Cluster Toolkit — self-managed Slurm cluster
- Cluster Director (product page)
- Cluster Director GA announcement (includes Slurm on GKE Preview)