| title | Snapshot |
|---|
⚠️ Experimental Feature: Dynamo Snapshot is currently in preview and may only be functional in some cluster setups. Thesnapshot-agentDaemonSet runs in privileged mode to perform CRIU operations. See Limitations for details.
Dynamo Snapshot is infrastructure for fast-starting GPU applications in Kubernetes using CRIU (Checkpoint/Restore in Userspace) and NVIDIA's cuda-checkpoint utility. The usual flow is:
- start a worker once and checkpoint its initialized state
- store that checkpoint on a namespace-local snapshot volume
- restore later workers from that checkpoint instead of cold-starting again
| Startup Type | Time | What Happens |
|---|---|---|
| Cold Start | ~1 min | Download model, load to GPU, initialize engine |
| Warm Start (restore from checkpoint) | ~10 sec | Restore from a ready checkpoint directory |
⚠️ Restore time depends on storage bandwidth, GPU model, and whether the restore stays on the same node.
- x86_64 (
amd64) GPU nodes - NVIDIA driver 580.xx or newer on the target GPU nodes (590.xx or newer if testing multi-GPU snapshots)
- vLLM or SGLang backend today
- Checkpoint storage.
ReadWriteManyis the safest default for cross-node or concurrent multi-node access, butpodMountmode can also use suitableReadWriteOncestorage for sequential checkpoint/restore workflows. - CRI-O / OpenShift: set
runtime.type=crioon the snapshot chart (andopenshift.enabled=trueon OpenShift). Defaults are for containerd; see the chart README for sockets and Helm flags.
- Build a placeholder image
- Install the snapshot chart
- Create a
DynamoCheckpointand wait for it to become ready - Deploy a
DynamoGraphDeploymentthat restores from the correspondingcheckpointRef
Snapshot-enabled workers must use a placeholder image that wraps the normal runtime image with restore tooling. If you do not already have one, build it and push it to a registry your cluster can pull from:
export RUNTIME_IMAGE=registry.example.com/dynamo/vllm-runtime:1.0.0
export PLACEHOLDER_IMAGE=registry.example.com/dynamo/vllm-placeholder:1.0.0
cd deploy/snapshot
make docker-build-placeholder \
PLACEHOLDER_BASE_IMG="${RUNTIME_IMAGE}" \
PLACEHOLDER_IMG="${PLACEHOLDER_IMAGE}"
make docker-push-placeholder \
PLACEHOLDER_IMG="${PLACEHOLDER_IMAGE}"The placeholder image preserves the normal runtime entrypoint/command contract and adds the criu, cuda-checkpoint, and nsrestore tooling needed for checkpoint and restore.
To build either snapshot image against a custom CRIU fork or ref, pass
CRIU_REPO and CRIU_REF through make. If they are unset, the Dockerfile
defaults are used.
make docker-build-agent \
IMG=registry.example.com/dynamo/snapshot-agent:1.0.0 \
CRIU_REPO="${YOUR_CRIU_REPO}" \
CRIU_REF="branch-or-sha"
make docker-build-placeholder \
PLACEHOLDER_BASE_IMG="${RUNTIME_IMAGE}" \
PLACEHOLDER_IMG="${PLACEHOLDER_IMAGE}" \
CRIU_REPO="${YOUR_CRIU_REPO}" \
CRIU_REF="branch-or-sha"Whether you are installing or upgrading dynamo-platform, the operator only needs checkpointing enabled:
dynamo-operator:
checkpoint:
enabled: trueIf the platform is already installed, verify that the operator config contains the checkpoint block:
OPERATOR_CONFIG=$(kubectl get deploy -n "${PLATFORM_NAMESPACE}" \
-l app.kubernetes.io/name=dynamo-operator,app.kubernetes.io/component=manager \
-o jsonpath='{.items[0].spec.template.spec.volumes[?(@.name=="operator-config")].configMap.name}')
kubectl get configmap "${OPERATOR_CONFIG}" -n "${PLATFORM_NAMESPACE}" \
-o jsonpath='{.data.config\.yaml}' | sed -n '/^checkpoint:/,/^[^[:space:]]/p'Verify that the rendered config includes enabled: true.
For the default namespace-local mode, install the snapshot chart in each workload namespace. The chart creates the PVC and the agent in that namespace:
helm upgrade --install snapshot ./deploy/helm/charts/snapshot \
--namespace ${NAMESPACE} \
--create-namespace \
--set storage.pvc.create=trueIn the default agentMount mode, the snapshot-agent DaemonSet mounts the
checkpoint PVC directly. On a multi-node GPU cluster that means agent pods on
multiple nodes may mount the same PVC, so the PVC generally needs
ReadWriteMany. The chart defaults to that mode. If your cluster does not have
a default storage class, also set storage.pvc.storageClass.
If you are reusing an existing checkpoint PVC, do not set storage.pvc.create=true; install the chart with storage.pvc.create=false and set storage.pvc.name instead.
CRI-O or OpenShift: append for example --set runtime.type=crio and, on OpenShift, --set openshift.enabled=true (see deploy/helm/charts/snapshot/README.md).
For clusters that prefer one privileged snapshot agent instead of one DaemonSet per workload namespace, install the chart once in an infrastructure namespace. In this mode the chart does not create workload PVCs; the Dynamo operator either creates each namespace-local PVC or verifies that it already exists:
helm upgrade --install snapshot ./deploy/helm/charts/snapshot \
--namespace dynamo-system \
--create-namespace \
--set storage.accessMode=podMount \
--set storage.pvc.create=false \
--set rbac.namespaceRestricted=falseTo let the operator create the workload PVC in each namespace that uses
checkpoint/restore, configure the operator with create: true:
dynamo-operator:
checkpoint:
enabled: true
storage:
type: pvc
pvc:
pvcName: snapshot-pvc
basePath: /checkpoints
create: true
size: 1Ti
storageClassName: ""
accessMode: ReadWriteManyThe chart and operator use separate configuration surfaces here: the snapshot
chart PVC name is storage.pvc.name, while the operator config field is
checkpoint.storage.pvc.pvcName.
This is a key difference from agentMount: podMount removes the requirement
that the snapshot-agent DaemonSet mount the checkpoint PVC on every GPU node.
Only the active checkpoint/restore workload pod mounts the PVC, and the agent
reaches it through that pod's mount namespace. ReadWriteMany remains the
safest operator-managed default, especially when multiple checkpoint/restore
pods may access the same PVC concurrently or when restore scheduling can span
nodes. Suitable ReadWriteOnce storage classes can still be used for
sequential podMount checkpoint/restore flows when the backend can attach the
volume to the node running the active workload pod.
podMount depends on the target container remaining alive while the agent
resolves /host/proc/<pid>/root/<basePath>. If the container exits or restarts
during checkpoint/restore setup, if the runtime cannot expose a stable host PID,
or if node security settings prevent host proc traversal, the agent fails or
skips that attempt and Kubernetes/operator reconciliation must try again after a
fresh container is available.
To use an already-present PVC instead, omit create or set it to false. The
operator will fail reconciliation with a clear error if the named PVC does not
exist in the workload namespace.
Verify that the DaemonSet is ready. After a checkpoint or restore workload is reconciled, verify the workload namespace PVC:
kubectl rollout status daemonset/snapshot-agent -n dynamo-system
kubectl get pods -n dynamo-system -l app.kubernetes.io/component=snapshot-agent -o wide
kubectl get pvc snapshot-pvc -n ${NAMESPACE}The checkpoint Job pod template should match the worker container you want to checkpoint. For the snapshot flow, the important parts are the checkpoint identity, a container named main, and the placeholder image; the rest of the pod template should mirror your normal worker config. Extra containers are allowed, but only main is checkpointed.
apiVersion: nvidia.com/v1alpha1
kind: DynamoCheckpoint
metadata:
name: qwen3-06b-bf16
spec:
identity:
model: Qwen/Qwen3-0.6B
backendFramework: vllm
tensorParallelSize: 1
dtype: bfloat16
maxModelLen: 2048
job:
activeDeadlineSeconds: 3600
podTemplateSpec:
spec:
...
containers:
- name: main
image: registry.example.com/dynamo/vllm-placeholder:1.0.0
...GMS + Snapshot support is currently disabled.
For a full working example, see deploy/operator/config/samples/nvidia.com_v1alpha1_dynamocheckpoint.yaml.
Apply it:
kubectl apply -f qwen3-checkpoint.yaml -n ${NAMESPACE}kubectl get dckpt -n ${NAMESPACE} \
-o custom-columns=NAME:.metadata.name,HASH:.status.identityHash,PHASE:.status.phase
kubectl wait \
--for=jsonpath='{.status.phase}'=Ready \
dynamocheckpoint/qwen3-06b-bf16 \
-n ${NAMESPACE} \
--timeout=30mThe useful status fields are:
status.phase: high-level lifecycle (Pending,Creating,Ready,Failed)status.identityHash: deterministic hash ofspec.identitystatus.jobName: checkpoint Job namestatus.createdAt: timestamp recorded when the checkpoint became readystatus.message: progress or failure detail when available
Once the checkpoint is Ready, restore a worker from it explicitly:
apiVersion: nvidia.com/v1alpha1
kind: DynamoGraphDeployment
metadata:
name: vllm-checkpointref-demo
spec:
services:
Frontend:
componentType: frontend
replicas: 1
extraPodSpec:
mainContainer:
image: registry.example.com/dynamo/vllm-runtime:1.0.0
VllmDecodeWorker:
componentType: worker
replicas: 1
checkpoint:
enabled: true
checkpointRef: qwen3-06b-bf16
extraPodSpec:
mainContainer:
image: registry.example.com/dynamo/vllm-placeholder:1.0.0
...
...Apply it:
kubectl apply -f vllm-checkpointref-demo.yaml -n ${NAMESPACE}
kubectl get pods -n ${NAMESPACE} -wThe VllmDecodeWorker pod should restore from the ready checkpoint instead of creating a new one.
checkpointRef is the most explicit path. mode: Auto is the higher-level path: the operator computes the checkpoint identity hash, looks for an equivalent DynamoCheckpoint, and creates one only when no matching checkpoint exists. If a DynamoCheckpoint already exists with the same identity, Auto mode reuses it. If no matching checkpoint exists yet, the first worker cold-starts and the operator creates the checkpoint in the background.
checkpoint:
enabled: true
mode: Auto
identity:
model: Qwen/Qwen3-0.6B
backendFramework: vllm
tensorParallelSize: 1
dtype: bfloat16
maxModelLen: 2048Inside a DynamoGraphDeployment, it looks like this:
apiVersion: nvidia.com/v1alpha1
kind: DynamoGraphDeployment
metadata:
name: vllm-auto-demo
spec:
services:
Frontend:
componentType: frontend
replicas: 1
extraPodSpec:
mainContainer:
image: registry.example.com/dynamo/vllm-runtime:1.0.0
VllmDecodeWorker:
componentType: worker
replicas: 1
checkpoint:
enabled: true
mode: Auto
identity:
model: Qwen/Qwen3-0.6B
backendFramework: vllm
tensorParallelSize: 1
dtype: bfloat16
maxModelLen: 2048
extraPodSpec:
mainContainer:
image: registry.example.com/dynamo/vllm-placeholder:1.0.0
...
...Auto mode only hashes checkpoint.identity. GMS-specific checkpoint behavior is not yet available.
Useful inspection commands:
kubectl get dgd vllm-auto-demo -n ${NAMESPACE} \
-o jsonpath='{.status.checkpoints.VllmDecodeWorker.checkpointName}{"\n"}{.status.checkpoints.VllmDecodeWorker.identityHash}{"\n"}{.status.checkpoints.VllmDecodeWorker.ready}{"\n"}'
kubectl get dckpt -n ${NAMESPACE}If you want to force a new restore after the checkpoint becomes ready, scale the worker:
kubectl patch dgd vllm-auto-demo -n ${NAMESPACE} --type=merge \
-p '{"spec":{"services":{"VllmDecodeWorker":{"replicas":2}}}}'Failover restore is not yet available. The current Snapshot flow does not support GMS + Snapshot, so do not use failover restore as a supported checkpoint/restore path. For current GMS and active/passive failover guidance, see Shadow Engine Failover.
It is possible to checkpoint and restore pods without the Dynamo operator via the lower-level snapshotctl utility. However, the snapshot helm chart must be installed, with a running snapshot-agent DaemonSet in the namespace with the checkpoint PVC mounted.
snapshotctl is intended for lower-level debugging and validation workflows, not as the primary user-facing checkpoint interface. For command details and manifest requirements, see deploy/snapshot/cmd/snapshotctl/README.md.
snapshotctl checkpoint \
--manifest ./worker-pod.yaml \
--container main \
--namespace ${NAMESPACE}The checkpoint manifest must be for a pod and use a placeholder image. --container names the workload container to checkpoint.
If you do not pass --checkpoint-id, snapshotctl generates one and prints it:
status=completed
namespace=...
name=...
checkpoint_job=...
checkpoint_id=manual-snapshot-...
checkpoint_location=/checkpoints/...
snapshotctl restore \
--manifest ./worker-pod.yaml \
--namespace ${NAMESPACE} \
--checkpoint-id manual-snapshot-... \
--containers mainThis creates a new restore pod and returns after the request is submitted. Observe progress through Kubernetes readiness, events, and logs.
snapshotctl restore \
--pod existing-restore-target \
--namespace ${NAMESPACE} \
--checkpoint-id manual-snapshot-... \
--containers mainThis patches restore metadata onto an existing pod that is already snapshot-compatible and returns after the patch is accepted.
Checkpoints are uniquely identified by a 16-character SHA256 hash (64 bits) of configuration that affects runtime state:
| Field | Required | Affects Hash | Example |
|---|---|---|---|
model |
✓ | ✓ | meta-llama/Llama-3-8B |
backendFramework |
✓ | ✓ | vllm |
dynamoVersion |
✓ | 0.9.0, 1.0.0 |
|
tensorParallelSize |
✓ | 1, 2, 4, 8 |
|
pipelineParallelSize |
✓ | 1, 2 |
|
dtype |
✓ | float16, bfloat16, fp8 |
|
maxModelLen |
✓ | 4096, 8192 |
|
extraParameters |
✓ | custom key-value pairs |
Fields that do not change the checkpoint hash include:
- replica count
- node placement (
nodeSelector,affinity,tolerations) - resource requests/limits
- logging or observability configuration
The DynamoCheckpoint (shortname: dckpt) is the operator-managed resource for checkpoint lifecycle.
Use it when you want:
- pre-warmed checkpoints before any
DynamoGraphDeploymentexists - explicit lifecycle control independent from a DGD
- a stable human-readable name that services can reference with
checkpointRef
The operator requires:
spec.identityspec.job.podTemplateSpec
spec.job.backoffLimit is deprecated and ignored. Checkpoint Jobs are always single-attempt.
Check status with:
kubectl get dckpt -n ${NAMESPACE}
kubectl describe dckpt qwen3-06b-bf16 -n ${NAMESPACE}
kubectl get dckpt qwen3-06b-bf16 -n ${NAMESPACE} -o yamlThe status block looks like:
status:
phase: Ready
identityHash: 3bff874d069f0ed5
jobName: checkpoint-job-3bff874d069f0ed5-1
createdAt: "2026-01-29T10:05:00Z"
message: ""- Backend support is limited: checkpoint/restore currently supports vLLM workers only, and that support is still a limited preview.
- Worker coverage is narrow: specialized workers such as multimodal, embedding, and diffusion are not supported.
- Multi-GPU remains preview: vLLM tensor-parallel configurations have limited validation and are not yet a broadly supported path across clusters.
- GMS restore remains experimental: GMS + Snapshot is currently disabled.
- Network state is sensitive: restore is sensitive to live TCP socket state. Loopback bootstrap/control sockets are the most reliable path today.
- Privileged DaemonSet required:
snapshot-agentmust run privileged to execute CRIU andcuda-checkpoint. Workload pods do not need to be privileged.
Snapshot only becomes Ready after snapshot-agent confirms the checkpoint contents. A completed Job is not enough by itself.
kubectl get dckpt <checkpoint-name> -n ${NAMESPACE} \
-o custom-columns=NAME:.metadata.name,PHASE:.status.phase,MESSAGE:.status.message,JOB:.status.jobName
JOB_NAME=$(kubectl get dckpt <checkpoint-name> -n ${NAMESPACE} -o jsonpath='{.status.jobName}')
if [ -n "${JOB_NAME}" ]; then
kubectl logs job/"${JOB_NAME}" -n ${NAMESPACE}
fi
kubectl logs daemonset/snapshot-agent -n ${NAMESPACE} --all-containersIf the worker template is wrong, the most common causes are using the raw runtime image instead of the placeholder image, or leaving out normal mounts and secrets that the worker needs to start.
For the default agentMount install, restore discovers checkpoint storage from
the snapshot-agent DaemonSet in the workload namespace. That DaemonSet must be
ready and must mount the checkpoint PVC.
kubectl rollout status daemonset/snapshot-agent -n ${NAMESPACE}
kubectl get daemonset -n ${NAMESPACE} -l app.kubernetes.io/component=snapshot-agent -o wide
kubectl get pvc -n ${NAMESPACE}For a shared-agent podMount install, the snapshot-agent DaemonSet can run in
the infrastructure namespace instead. Verify the shared-agent pods there, then
verify that the workload namespace has the checkpoint PVC that the operator
created or validated:
kubectl rollout status daemonset/snapshot-agent -n dynamo-system
kubectl get pods -n dynamo-system -l app.kubernetes.io/component=snapshot-agent -o wide
kubectl get pvc snapshot-pvc -n ${NAMESPACE}In podMount mode the agent reaches the checkpoint through the workload pod's
mount namespace rather than by mounting the PVC itself. Check the workload pod's
checkpoint storage annotations and the snapshot-agent logs to see the actual
resolved checkpoint path. snapshotctl uses the chart's storage resolution
path, so for lower-level snapshotctl debugging make sure the snapshot chart
configuration matches the access mode you are testing.
snapshotctl requires a Pod manifest and a target-container list. Multi-container manifests are supported as long as every name passed via --container or --containers exists in the pod spec.
snapshotctl checkpoint --manifest ./worker-pod.yaml --container main --namespace ${NAMESPACE}
snapshotctl restore --manifest ./worker-pod.yaml --containers main --namespace ${NAMESPACE} --checkpoint-id <checkpoint-id>If the manifest already carries snapshot target metadata, it must agree with the CLI flag; snapshotctl rejects mismatches instead of silently picking one.
- Stabilize multi-GPU support
- Additional backend support
- Alternative storage backends