diff --git a/confidential-containers/confidential-containers-deploy.rst b/confidential-containers/confidential-containers-deploy.rst index d08aa1be5..115249281 100644 --- a/confidential-containers/confidential-containers-deploy.rst +++ b/confidential-containers/confidential-containers-deploy.rst @@ -55,7 +55,7 @@ After completing the installation, you can :doc:`Run a Sample Workload ` or :ref:`Host OS Administrator ` to confirm or implement hardware prerequisites. -For validated hardware and software versions, refer to :doc:`Supported Platforms `. -Use the checklists below for an at-a-glance summary, then follow each linked section for verification steps. +Use the following checklists as at-a-glance summary and supplement to the :doc:`Supported Platforms ` page. -**Hardware prerequisites** +**Hardware Configuration Requirements** + +For validated hardware, refer to :doc:`Supported Platforms `. .. list-table:: :header-rows: 1 :widths: 30 70 - * - Prerequisite + * - Configuration Requirement - Details - * - :ref:`Use a supported platform ` - - CPU, GPU, and host OS match :doc:`Supported Platforms ` * - :ref:`Hardware virtualization and ACS enabled ` - Hardware virtualization and ACS enabled in host BIOS * - :ref:`IOMMU enabled ` @@ -47,26 +46,26 @@ Use the checklists below for an at-a-glance summary, then follow each linked sec * - :ref:`No host NVIDIA GPU drivers ` - No NVIDIA GPU drivers installed or loaded on worker hosts. -**Cluster prerequisites** +**Cluster Configuration Requirements** + +For validated software components, refer to :doc:`Supported Platforms `. .. list-table:: :header-rows: 1 :widths: 30 70 - * - Prerequisite + * - Configuration Requirement - Details * - :ref:`A Kubernetes cluster and cluster administrator access ` - - Cluster administrator access to a Kubernetes cluster running a supported version (refer to :ref:`Supported Software Components `) - * - :ref:`containerd 2.2.2 installed ` - - containerd 2.2.2 installed on each GPU worker node + - Cluster administrator access to a Kubernetes cluster running a supported version. * - :ref:`Helm installed ` - Helm installed on your cluster administration system * - :ref:`Kubelet configured ` - Enable ``KubeletPodResourcesGet`` (required before Kubernetes v1.34) and ``RuntimeClassInImageCriApi`` feature gates; set ``runtimeRequestTimeout: 20m`` on GPU worker nodes -***************** -Hardware and BIOS -***************** +******************************************** +Hardware and BIOS Configuration Requirements +******************************************** .. _coco-prereq-supported-platform: @@ -172,9 +171,9 @@ In this architecture, the NVIDIA GPU Operator handles GPU driver installation an Refer to `Removing the Driver `_ in the NVIDIA Driver Installation Guide. -****************** -Kubernetes Cluster -****************** +********************************************* +Kubernetes Cluster Configuration Requirements +********************************************* The following sections describe requirements for worker nodes and for the system you use for cluster administration. @@ -186,26 +185,6 @@ Kubernetes Cluster and Cluster Administrator Access You must have cluster administrator access to a Kubernetes cluster running a supported Kubernetes version. Refer to the :ref:`Supported Software Components ` section in :doc:`Supported Platforms ` for supported Kubernetes and component versions. -.. _coco-prereq-containerd: - -containerd 2.2.2 -================ - -Verify the installed version on each GPU worker node: - -.. code-block:: console - - $ containerd --version - -*Example Output:* - -.. code-block:: output - - containerd containerd.io 2.2.2 ... - -Your actual output may vary, but the reported version must be ``2.2.2``. -If you are running a different version on any worker node, refer to the `containerd Getting Started guide `_ for installation instructions. - .. _coco-prereq-helm: Helm @@ -256,7 +235,7 @@ Apply these settings as follows: #. Open the kubelet configuration file: .. code-block:: console - + $ sudo nano /var/lib/kubelet/config.yaml This is typically located at ``/var/lib/kubelet/config.yaml``, but your configuration file may be in a different location. diff --git a/confidential-containers/release-notes.rst b/confidential-containers/release-notes.rst index 94e430ee9..230da0c2a 100644 --- a/confidential-containers/release-notes.rst +++ b/confidential-containers/release-notes.rst @@ -26,6 +26,40 @@ This document describes the new features and known issues for the NVIDIA Confide ---- +.. _coco-v1.1.0: + +1.1.0 +===== + +This release expands hardware coverage and updates the validated software stack. + +New Features +------------ + +* Added support for the NVIDIA HGX B300 platform with both single-GPU and multi-GPU passthrough. + +* Added support for Ubuntu 26.04 as a host operating system. + +* Added support for the following software components: + + * Kata Containers 3.31.0 + * containerd 2.3.x + + +Docs Changelog +-------------- + +The :ref:`coco-install-kata-chart` procedure was updated for this release. +Changes include: + +* Installs ``kata-deploy`` with a values file instead of inline ``--set`` flags. + +* Includes a new sample values file, :file:`samples/kata-nvidia-gpu-values.yaml`, that configures the ``kata-deploy`` Helm chart for the NVIDIA Confidential Containers reference architecture (NVIDIA GPU shims only, NFD disabled, ``nydus`` snapshotter, and per-shim runtime class node selectors). + +* Adds a readiness verification step using ``kubectl rollout status ds/kata-deploy``. This step relies on the readiness reporting added in Kata Containers 3.31.0 and lets you confirm that ``kata-deploy`` has finished extracting artifacts and restarting containerd on every node before continuing. + +---- + .. _coco-v1.0.0: ***** diff --git a/confidential-containers/samples/kata-nvidia-gpu-values.yaml b/confidential-containers/samples/kata-nvidia-gpu-values.yaml new file mode 100644 index 000000000..cb7ebf608 --- /dev/null +++ b/confidential-containers/samples/kata-nvidia-gpu-values.yaml @@ -0,0 +1,107 @@ +# Example values file to enable NVIDIA GPU shims for the NVIDIA +# Confidential Containers Reference Architecture. + +# Set to true for verbose kata-deploy and Kata runtime logging. +debug: false + +# Disable Node Feature Discovery (NFD) deployment by kata-deploy. +# Both kata-deploy and the GPU Operator deploy NFD by default. This +# reference architecture relies on the NFD instance that the GPU Operator +# deploys and manages, so the kata-deploy NFD is turned off to avoid a +# duplicate, conflicting deployment. +nfd: + enabled: false + +# Install the nydus snapshotter on each node alongside containerd. +# The confidential -snp and -tdx shims below use nydus to pull container +# images directly into the confidential VM (guest pull), which keeps image +# contents inside the trusted execution environment (TEE). +snapshotter: + setup: ["nydus"] + +# Disable the chart's default hypervisor/TEE shims and opt in only to +# the NVIDIA GPU shims supported by this reference architecture. +shims: + disableAll: true + + # Non-confidential NVIDIA GPU passthrough shim used when Confidential + # Computing mode is off on the node. The runtime class is restricted to + # nodes where the GPU Operator's Confidential Computing Manager has + # reported nvidia.com/cc.ready.state=false, so it will not schedule on + # CC-ready nodes. The empty containerd snapshotter falls back to the + # default (overlayfs); guest pull is not used for this non-confidential + # path. + qemu-nvidia-gpu: + enabled: true + supportedArches: + - amd64 + allowedHypervisorAnnotations: [] + containerd: + snapshotter: "" + runtimeClass: + # This label is automatically added by the GPU Operator. + nodeSelector: + nvidia.com/cc.ready.state: "false" + + # Confidential NVIDIA GPU passthrough for AMD SEV-SNP nodes. + # Scheduled where the GPU Operator reports CC mode is on AND NFD + # reports SEV-SNP support. Set agent.httpsProxy / agent.noProxy if + # the guest needs a proxy to reach the registry. + qemu-nvidia-gpu-snp: + enabled: true + supportedArches: + - amd64 + allowedHypervisorAnnotations: [] + containerd: + snapshotter: "nydus" + forceGuestPull: false + crio: + guestPull: true + agent: + httpsProxy: "" + noProxy: "" + runtimeClass: + # These labels are automatically added by the GPU Operator and NFD + # respectively. + nodeSelector: + nvidia.com/cc.ready.state: "true" + amd.feature.node.kubernetes.io/snp: "true" + + # Confidential NVIDIA GPU passthrough for Intel TDX nodes. + # Same selectors and snapshotter behavior as the SNP shim above, + # but pinned to TDX-capable hosts. + qemu-nvidia-gpu-tdx: + enabled: true + supportedArches: + - amd64 + allowedHypervisorAnnotations: [] + containerd: + snapshotter: "nydus" + forceGuestPull: false + crio: + guestPull: true + agent: + httpsProxy: "" + noProxy: "" + runtimeClass: + # These labels are automatically added by the GPU Operator and NFD + # respectively. + nodeSelector: + nvidia.com/cc.ready.state: "true" + intel.feature.node.kubernetes.io/tdx: "true" + +# Default shim when a pod does not request a runtime class. Set to the +# non-confidential shim so pods only run in a confidential VM when +# they explicitly request the -snp or -tdx runtime class. +defaultShim: + amd64: qemu-nvidia-gpu # Can be changed to qemu-nvidia-gpu-snp or qemu-nvidia-gpu-tdx if preferred + +# Create one Kubernetes RuntimeClass per enabled shim above +# (kata-qemu-nvidia-gpu, kata-qemu-nvidia-gpu-snp, kata-qemu-nvidia-gpu-tdx). +# createDefault: false suppresses the generic "kata" RuntimeClass since +# you should always reference a specific NVIDIA shim +# by name in pod specs. +runtimeClasses: + enabled: true + createDefault: false + defaultName: "kata" diff --git a/confidential-containers/supported-platforms.rst b/confidential-containers/supported-platforms.rst index 2afb6538c..aa6cba3ab 100644 --- a/confidential-containers/supported-platforms.rst +++ b/confidential-containers/supported-platforms.rst @@ -59,6 +59,9 @@ NVIDIA GPUs * - NVIDIA B200 - Single-GPU, Multi-GPU + * - NVIDIA HGX B300 + - Single-GPU, Multi-GPU + * - NVIDIA RTX Pro 6000 BSE - Single-GPU @@ -84,11 +87,11 @@ CPU Platforms - Kernel Version * - AMD Genoa / Milan - AMD SEV-SNP - - Ubuntu 25.10 + - Ubuntu 25.10 or 26.04 - 6.17+ * - Intel Emerald Rapids (ER) / Granite Rapids (GR) - Intel TDX - - Ubuntu 25.10 + - Ubuntu 25.10 or 26.04 - 6.17+ For additional information on node configuration, refer to the `Confidential Computing Deployment Guide `_ for information about supported NVIDIA GPUs, such as the NVIDIA Hopper H100. @@ -98,7 +101,7 @@ The following topics in the deployment guide apply to a cloud-native environment * Hardware selection and initial hardware configuration, such as BIOS settings. * Host operating system selection, initial configuration, and validation. -When following the cloud-native sections in the deployment guide linked above, use Ubuntu 25.10 as the host OS with its default kernel version and configuration. +When following the cloud-native sections in the deployment guide linked above, use Ubuntu 25.10 or 26.04 as the host OS with its default kernel version and configuration. For additional resources on machine setup: @@ -126,24 +129,18 @@ Supported Software Components * - `QEMU `__ - 10.1 \+ Patches * - `Containerd `__ - - 2.2.2 + - 2.3.x * - `Kubernetes `__ - 1.32 \+ * - `NVIDIA GPU Operator `__ and its components. - + Refer to the :ref:`GPU Operator Component Matrix ` for the list of components and versions included in each release. - - v26.3.1 and higher + - ${gpu_operator_version} and higher * - `Kata Containers `__ - - 3.29 (installed with ``kata-deploy`` Helm chart) + - ${kata_version} (installed with ``kata-deploy`` Helm chart) * - `Key Broker Service (KBS) protocol `__ - 0.4.0 * - `Kata Lifecycle Manager `__ - 0.1.4 Users may leverage `Red Hat OpenShift Sandboxed Containers `__ to deploy Confidential Containers, however, Confidential GPU features are currently classified as Technology Preview by the downstream provider. - - - - - - diff --git a/gpu-operator/deploy-kata-containers.rst b/gpu-operator/deploy-kata-containers.rst index 3d9801d06..2ad1a28a7 100644 --- a/gpu-operator/deploy-kata-containers.rst +++ b/gpu-operator/deploy-kata-containers.rst @@ -32,7 +32,7 @@ Deploy with Kata Containers About the Operator with Kata Containers *************************************** -`Kata Containers `_ is an open source project that creates lightweight Virtual Machines (VMs) that feel and perform like traditional containers such as a Docker container. +`Kata Containers `_ is an open source project that creates lightweight Virtual Machines (VMs) that feel and perform like traditional containers such as a Docker container. A traditional container packages software for user-space isolation from the host, but the container runs on the host and shares the operating system kernel with the host. Sharing the operating system kernel is a potential vulnerability. @@ -129,9 +129,9 @@ To enable Kata Containers for GPUs on your cluster, you do the following: #. Make sure your cluster meets the prerequisites. #. Label the nodes you want to use for Kata Containers. -#. Install the upstream ``kata-deploy`` Helm chart, which deploys all Kata runtime classes, including NVIDIA-specific runtime classes. +#. Install the upstream ``kata-deploy`` Helm chart, which deploys all Kata runtime classes, including NVIDIA-specific runtime classes. The ``kata-qemu-nvidia-gpu`` runtime class is used with Kata Containers. -#. Install the NVIDIA GPU Operator with Kata sandbox mode enabled. +#. Install the NVIDIA GPU Operator with Kata sandbox mode enabled. After installation, you can run a sample workload that uses the Kata runtime class. @@ -225,7 +225,7 @@ Kubernetes Cluster Refer to the `Kata Containers documentation `_ for more details on the Kata runtime and VFIO cold-plug. -* Increase kubelet image pull timeouts configuration to 20 minutes to avoid timeouts when pulling large images. +* Increase kubelet image pull timeouts configuration to 20 minutes to avoid timeouts when pulling large images. Kubelet can de-allocate your pod if the image pull exceeds the configured timeout before the container transitions to the running state. Increase ``runtimeRequestTimeout`` in your `kubelet configuration `_ to ``20m`` to match the default values for the Kata shim configurations in Kata Containers. @@ -246,8 +246,8 @@ Kubernetes Cluster $ sudo systemctl restart kubelet - If you need a timeout of more than 1200 seconds (20 minutes), you will also need to adjust the Kata Agent's ``image_pull_timeout``, which defaults to 1200s. - This setting also sets the confidential data hub's image pull API timeout in seconds. + If you need a timeout of more than 1200 seconds (20 minutes), you will also need to adjust the Kata Agent's ``image_pull_timeout``, which defaults to 1200s. + This setting also sets the confidential data hub's image pull API timeout in seconds. To do this, add the ``agent.image_pull_timeout`` kernel parameter to your shim configuration, or pass an explicit value in a pod annotation in the ``io.katacontainers.config.hypervisor.kernel_params: "..."`` annotation. .. _label-nodes-kata-containers: @@ -305,13 +305,13 @@ Install the Kata Containers Helm Chart Install Kata Containers using the ``kata-deploy`` Helm chart. The ``kata-deploy`` chart installs all required components from the Kata Containers project including the Kata Containers runtime binary, runtime configuration, UVM kernel, and images that NVIDIA uses for Kata Containers. -The minimum required version is 3.29.0. +The minimum required version is ${kata_version}. #. Set the chart version and registry path: .. code-block:: console - $ export VERSION="3.29.0" + $ export VERSION="${kata_version}" $ export CHART="oci://ghcr.io/kata-containers/kata-deploy-charts/kata-deploy" @@ -322,9 +322,12 @@ The minimum required version is 3.29.0. $ helm install kata-deploy "${CHART}" \ --namespace kata-system --create-namespace \ --set nfd.enabled=false \ - --wait --timeout 10m \ + -f kata-nvidia-gpu-values.yaml \ --version "${VERSION}" + The sample `kata-nvidia-gpu-values.yaml` file is included in the GitHub repository at + https://github.com/NVIDIA/cloud-native-docs/blob/main/confidential-containers/samples/kata-nvidia-gpu-values.yaml. + *Example Output:* .. code-block:: output @@ -336,33 +339,45 @@ The minimum required version is 3.29.0. DESCRIPTION: Install complete TEST SUITE: None - .. note:: - - The ``--wait`` flag in the install command instructs Helm to wait until the release is deployed before returning. - It can take a few minutes to return output. - - There is a `known Helm issue `_ on single node clusters, that may result in the Helm command finishing before all deployed pods are finished initializing. - If you are deploying to a single node cluster, you may need to wait for an additional few minutes after the Helm command completes for the ``kata-deploy`` pod to be in the Running state. - .. note:: Both ``kata-deploy`` and the GPU Operator deploy Node Feature Discovery (NFD) by default. The install command includes ``--set nfd.enabled=false`` to prevent ``kata-deploy`` from deploying NFD. The GPU Operator will deploy and manage NFD in the next step. + .. note:: + + The Helm install command returns as soon as the Kubernetes resources are created. + The ``kata-deploy`` DaemonSet then takes several minutes per node to extract artifacts, restart containerd, and label the node before its pods report ready. + You can use either of the optional verification steps below to confirm readiness before continuing. + -#. Optional: Verify that the ``kata-deploy`` pod is running: +#. Optional: Verify that the ``kata-deploy`` DaemonSet has finished rolling out on every node: .. code-block:: console - $ kubectl get pods -n kata-system | grep kata-deploy + $ kubectl -n kata-system rollout status ds/kata-deploy --timeout=20m + + *Example Output:* + + .. code-block:: output + + Waiting for daemon set "kata-deploy" rollout to finish: 0 of 1 updated pods are available... + daemon set "kata-deploy" successfully rolled out + + +#. Optional: Verify that the ``kata-deploy`` pods are running: + + .. code-block:: console + + $ kubectl get pods -n kata-system *Example Output:* .. code-block:: output - NAME READY STATUS RESTARTS AGE - kata-deploy-b2lzs 1/1 Running 0 6m37s + NAME READY STATUS RESTARTS AGE + kata-deploy-b2lzs 1/1 Running 0 6m37s #. Optional: Verify that the ``kata-qemu-nvidia-gpu`` runtime class is available: @@ -649,8 +664,8 @@ If the sample workload does not run, confirm that you labeled nodes to run virtu .. code-block:: output NAME STATUS ROLES AGE VERSION - kata-worker-1 Ready 10d v1.35.3 - kata-worker-2 Ready 10d v1.35.3 + kata-worker-1 Ready 10d v1.35.3 + kata-worker-2 Ready 10d v1.35.3 kata-worker-3 Ready 10d v1.35.3 You might have configured ``vm-passthrough`` as the default sandbox workload in the ClusterPolicy resource. @@ -669,4 +684,3 @@ Also confirm in the ClusterPolicy that ``sandboxWorkloads`` is configured for Ka enabled: true defaultWorkload: vm-passthrough mode: kata - diff --git a/repo.toml b/repo.toml index 374e1d164..269826a02 100644 --- a/repo.toml +++ b/repo.toml @@ -176,7 +176,7 @@ docs_root = "${root}/gpu-operator" project = "gpu-operator" name = "NVIDIA GPU Operator" version = "26.3" # Update repo_docs.projects.openshift.version to match latest patch version maj.min.patch -source_substitutions = { minor_version = "26.3", version = "v26.3.3", recommended = "580.173.02", dra_version = "0.4.1" } +source_substitutions = { minor_version = "26.3", version = "v26.3.3", recommended = "580.173.02", dra_version = "0.4.1", kata_version = "3.31.0" } copyright_start = 2020 sphinx_exclude_patterns = [ "life-cycle-policy.rst", @@ -213,7 +213,8 @@ output_format = "linkcheck" docs_root = "${root}/confidential-containers" project = "confidential-containers" name = "NVIDIA Confidential Containers Architecture" -version = "1.0.0" +version = "1.1.0" +source_substitutions = { kata_version = "3.31.0", gpu_operator_version = "v26.3.1", gpu_operator_minor_version = "26.3" } copyright_start = 2020 [repo_docs.projects.confidential-containers.builds.linkcheck]