GPU Kubernetes and Scheduling
Last reviewed: September 2026 | This is a fast-moving area subject to quarterly review.
Overview
Section titled “Overview”GPUs are expensive and scarce. So rather than one team monopolizing them, multiple teams and jobs usually share one GPU cluster. What decides “who gets how many GPUs, and when” is the scheduler, and a bad allocation leaves costly GPUs idle or lets one team monopolize them.
Kubernetes (the standard tool for automatically placing and managing containers) treats GPUs as a special resource. The hard part is that one cluster must accept two jobs of opposite nature — a training job starts only when it gets “all the GPUs it needs at once,” while an inference job runs by “holding a few GPUs for a long time.”
GPU Node Pools and Device Plugins
Section titled “GPU Node Pools and Device Plugins”Kubernetes does not recognize GPUs by default, so you install a vendor device plugin along with drivers/operators to expose GPUs as schedulable resources.
- GPU node pool — Configure a dedicated GPU node pool separate from general-purpose nodes. (See Node pool composition.)
- Device plugin / operator — The NVIDIA GPU Operator, which deploys drivers, device plugin, and DCGM together, is the de facto standard.
- taint/toleration — Taint GPU nodes so non-GPU workloads do not occupy expensive GPU nodes.
GPU Sharing — MIG and Time-Slicing
Section titled “GPU Sharing — MIG and Time-Slicing”Letting multiple jobs share a single GPU can raise utilization for small inference/development workloads.
| Method | Isolation level | Suitable workloads | Limitations |
|---|---|---|---|
| MIG (Multi-Instance GPU) | Hardware partition (memory/compute isolation) | Predictable multi-tenant inference | Supported-GPU/profile constraints, dynamic-change overhead |
| MPS (Multi-Process Service) | Shared process space (partial isolation) | Cooperative multi-process, small inference | No memory-isolation guarantee, fault propagation possible |
| Time-slicing | Time division (no isolation) | Development/experiments, bursty workloads | Interference/OOM risk, no fairness guarantee |
MIG is the strongest via hardware isolation, time-slicing has no isolation, and MPS sits in between with multiple processes sharing one GPU context. All three are common NVIDIA GPU features, not vendor-exclusive.
Quotas, Fairness, and Anti-Hoarding
Section titled “Quotas, Fairness, and Anti-Hoarding”The headache of a shared cluster is when one team grabs GPUs and won’t let go (monopoly, hoarding). Then other jobs can’t get GPUs and keep starving. The mechanisms that prevent this:
- ResourceQuota (total cap) — Set an upper bound on the number of GPUs each team (namespace) can use.
- Priority/preemption — Rank jobs by priority (PriorityClass), and when an urgent job arrives, briefly push aside (preempt) a less urgent one to yield GPUs.
- Reclaiming idle GPUs (anti-hoarding) — A queue-based scheduler uses fair-share and reclaim rules to take back GPUs that are held but not actually used, and give them to other jobs.
- Queue-based allocation — The gang-scheduling tier below manages per-team shares and waiting lines.
Gang Scheduling
Section titled “Gang Scheduling”Distributed training can only start once it secures all the GPUs it needs at once. For example, if a job needs 16 GPUs but grabs only 10 and waits for the other 6, those 10 do nothing and just tie up resources (in the worst case, a deadlock where jobs wait on each other). Gang scheduling prevents this by scheduling “start only if all needed are secured, otherwise don’t start at all (all-or-nothing).” It means grabbing the whole “gang” together.
| Tool | Characteristics |
|---|---|
| Kueue | Kubernetes-native job queuing, quotas/fair-share, hierarchical queues |
| Volcano | Batch scheduler, gang scheduling/queues/preemption integrated, HPC/AI-oriented |
Vendor Managed Kubernetes GPU Support
Section titled “Vendor Managed Kubernetes GPU Support”| Item | AWS (EKS) | Azure (AKS) | Google Cloud (GKE) | OCI (OKE) |
|---|---|---|---|---|
| GPU node pool | Managed node groups | GPU node pools | GPU node pools | GPU node pools |
| Driver installation | GPU Operator / EKS-optimized AMI | GPU Operator / AKS GPU image | GPU Operator / GKE driver auto-install | GPU Operator / OKE image |
| GPU sharing | MIG, MPS, time-slicing | MIG, MPS, time-slicing | MIG, MPS, time-slicing | MIG, MPS, time-slicing |
Observability — DCGM and GPU Metrics
Section titled “Observability — DCGM and GPU Metrics”GPU clusters cannot be diagnosed for bottlenecks with CPU-centric observability alone. Collect GPU utilization, memory, temperature, and network traffic with NVIDIA DCGM (Data Center GPU Manager, the standard tool for collecting GPU status and performance).
- Key metrics — GPU utilization (actual compute utilization, not mere allocation), memory usage, fabric bandwidth, power/temperature
- The utilization trap — “A GPU is allocated” and “a GPU is actually computing” are different. Low effective utilization is a sign of a data-loading or communication bottleneck.
- SLO linkage — Integrate collected GPU metrics into your SLO and Observability systems to continuously manage training throughput and inference latency.
Related Documents
Section titled “Related Documents”- Next: Inference serving, failures, capacity, cost — Inference Serving, Reliability, and Cost
- General Kubernetes operations (upgrades, node management) — Kubernetes Operations
- Cluster communication, placement, managed clusters — GPU Workload Characteristics and Reference Architecture
- Parallelism strategies (TP/PP/DP) — Distributed Training Standard Architecture
Common Mistakes
Section titled “Common Mistakes”- Submitting distributed training without gang scheduling — Acquiring only some GPUs and waiting, tying up resources and causing deadlock
- Using time-slicing for production multi-tenancy — No isolation, so one job’s OOM propagates to others
- Absence of ResourceQuota/preemption policies — One team monopolizes (hoards) GPUs, starving other jobs
- Watching only GPU allocation rate, not utilization — Missing low effective utilization (data/communication bottlenecks), wasting expensive GPUs
Checklist
Section titled “Checklist”- Did you taint the dedicated GPU node pool to block non-GPU workload occupancy?
- Did you consider MIG (hardware isolation) instead of time-slicing for multi-tenant inference?
- Did you set per-namespace ResourceQuota and PriorityClass/preemption policies?
- Did you apply gang scheduling (Kueue/Volcano or a managed built-in) to distributed training?
- Did you collect effective GPU utilization with DCGM and link it to SLOs?