# Add-on: gpu

**URL:** <https://discuss.kubernetes.io/t/add-on-gpu/11286>\
**Category:** microk8s\
**Tags:** docs\
**Created:** [June 4, 2020, 3:48pm UTC](https://discuss.kubernetes.io/t/add-on-gpu/11286 "2020-06-04T15:48:58Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![toto](https://sea2.discourse-cdn.com/flex016/user_avatar/discuss.kubernetes.io/toto/32/5256_2.png) [@toto](https://discuss.kubernetes.io/u/toto)\
**Post date:** [June 4, 2020, 3:48pm UTC](https://discuss.kubernetes.io/t/add-on-gpu/11286/1 "2020-06-04T15:48:58Z")

</div>

This addon enables NVIDIA GPU support on MicroK8s using the [NVIDIA GPU Operator](https://github.com/NVIDIA/gpu-operator) and offers:

- Use of any existing NVIDIA host drivers, or compilation and loading of kernel drivers dynamically at runtime.
- Installation and configuration of the `nvidia-container-runtime` for containerd.
- Configuration of the `nvidia.com/gpu` kubelet device plugin, to support resource capacity and limits on GPU nodes.
- Multi-instance GPU (MIG) configuration via ConfigMap resources.

You can enable this addon with the following command:

```bash
microk8s enable gpu

```

> _NOTE_: Starting with MicroK8s 1.36, the GPU addon **no longer forces the NVIDIA runtime as the default containerd runtime**. The GPU Operator determines the appropriate value based on its own settings. As a result, workloads should explicitly set `runtimeClassName: nvidia` in their pod spec to use the GPU.

> _NOTE_: For MicroK8s 1.25 or older, if you see an an error similar to
> 
> ```auto
> Error: INSTALLATION FAILED: failed to download "nvidia/gpu-operator" at version "v22.9.0"
> 
> ```
> 
> You can instead try:
> 
> ```auto
> microk8s helm repo update nvidia
> microk8s enable gpu
> 
> ```

> _NOTE_: For using GPU Operator `v25.10.0+` with MicroK8s releases older than `1.35`, please set the required `RUNTIME_CONFIG_SOURCE` variable like below:
> 
> ```auto
> microk8s enable gpu --gpu-operator-version v25.10.0 --gpu-operator-set toolkit.env[3].name=RUNTIME_CONFIG_SOURCE --gpu-operator-set toolkit.env[3].value='file=/var/snap/microk8s/current/args/containerd.toml'
> 
> ```
> 
> Necessary configuration changes above are included by default on MicroK8s versions `1.35` or newer.

> _NOTE_: The GPU addon is supported on MicroK8s versions 1.22 or newer. For MicroK8s 1.21, see [GPU addon on MicroK8s 1.21](#microk8s-121-12).

### Verify installation

Verify that all components are deployed and configured correctly with:

```bash
microk8s kubectl logs -n gpu-operator-resources -lapp=nvidia-operator-validator -c nvidia-operator-validator

```

which should return:

```bash
all validations are successful

```

### Deploy a test workload

Once the GPU addon is enabled, workloads can request the GPU using a limit setting, e.g. `nvidia.com/gpu: 1`. For example, you can run a `cuda-vector-add` test pod with:

```bash
microk8s kubectl apply -f - <<EOF
apiVersion: v1
kind: Pod
metadata:
  name: cuda-vector-add
spec:
  restartPolicy: OnFailure
  runtimeClassName: nvidia
  containers:
    - name: cuda-vector-add
      image: "registry.k8s.io/cuda-vector-add:v0.1"
      resources:
        limits:
          nvidia.com/gpu: 1
EOF

```

And then check the pod’s logs to verify that everything is okay:

```bash
microk8s kubectl logs cuda-vector-add

```

where a successful run would produce logs similar to:

```bash
[Vector addition of 50000 elements]
Copy input data from the host memory to the CUDA device
CUDA kernel launch with 196 blocks of 256 threads
Copy output data from the CUDA device to the host memory
Test PASSED
Done

```

You are ready to run GPU workloads on your MicroK8s cluster!

## Addon configuration options

> _NOTE_: These require MicroK8s version 1.28 or newer. Check the installed revision with `snap list microk8s`.

In the `microk8s enable gpu` command, the following command-line arguments may be set:

| Argument | Default | Description |
| --- | --- | --- |
| `--driver $driver` | `auto` | Supported values are `auto` (use host driver if found), `host` (force use the host driver), or `operator` (force use the operator driver). |
| `--version $VERSION` | `v25.10.0` | Version of the GPU operator to install. |
| `--toolkit-version $VERSION` | `` | If not empty, override the version of the `nvidia-container-runtime` that will be installed. |
| `--gpu-operator-set-as-default-runtime / --gpu-operator-no-set-as-default-runtime` | `unset (starting 1.36+)` | When not set, the GPU Operator determines whether to set nvidia as the default containerd runtime. Pass `--gpu-operator-set-as-default-runtime` to force it on, or `--gpu-operator-no-set-as-default-runtime` to force it off. |
| `--set $key=$value` | `` | Set additional configuration options to the GPU operator Helm chart. May be passed multiple times. For a list of options see [values.yaml](https://github.com/NVIDIA/gpu-operator/blob/master/deployments/gpu-operator/values.yaml). |
| `--values $file` | `` | Set additional configuration options to the GPU operator Helm chart using a file. May be passed multiple times. For a list of options see [values.yaml](https://github.com/NVIDIA/gpu-operator/blob/master/deployments/gpu-operator/values.yaml). |

## Use host drivers and runtime

### Use host NVIDIA drivers

The GPU addon works with the existing NVIDIA host drivers (if available), otherwise it will deploy the `nvidia-driver-daemonset` to dynamically build and load the NVIDIA drivers into the kernel.

In order to use host drivers, install the NVIDIA drivers **before** enabling the addon. See [Nvidia driver installation docs](https://ubuntu.com/server/docs/nvidia-drivers-installation#p-97843-the-recommended-way-ubuntu-drivers-tool).

Verify that drivers are loaded by checking `nvidia-smi`:

```bash
nvidia-smi

```

Then enable the addon:

```bash
microk8s enable gpu

```

### Use host nvidia-container-runtime

The GPU addon will automatically install `nvidia-container-runtime`, which is the runtime required to execute GPU workloads on the MicroK8s cluster. This is done by the `nvidia-container-toolkit-daemonset` pod.

If needed, this section documents how you to install the nvidia-container-runtime manually. The steps below should be performed **before** enabling the GPU addon.

Install nvidia-container-runtime following the [upstream instructions](https://nvidia.github.io/nvidia-container-runtime/). At the time of writing, the instructions for Ubuntu hosts look like this:

```bash
curl -s -L https://nvidia.github.io/nvidia-container-runtime/gpgkey | \
  sudo apt-key add -
distribution=$(. /etc/os-release;echo $ID$VERSION_ID)
curl -s -L https://nvidia.github.io/nvidia-container-runtime/$distribution/nvidia-container-runtime.list | \
  sudo tee /etc/apt/sources.list.d/nvidia-container-runtime.list
sudo apt-get update
sudo apt-get install nvidia-container-runtime

```

This will install `nvidia-container-runtime` in `/usr/bin/nvidia-container-runtime`. Next, edit the containerd configuration file so that it knows where to find the runtime binaries for the `nvidia` runtime:

```bash
echo '
        [plugins."io.containerd.grpc.v1.cri".containerd.runtimes.nvidia]
          runtime_type = "io.containerd.runc.v2"

          [plugins."io.containerd.grpc.v1.cri".containerd.runtimes.nvidia.options]
            BinaryName = "/usr/bin/nvidia-container-runtime"
' | sudo tee -a /var/snap/microk8s/current/args/containerd-template.toml

```

Restart MicroK8s to reload the containerd configuration:

```bash
sudo snap restart microk8s

```

Finally, enable the gpu addon and make sure that the toolkit daemonset is not deployed:

```bash
microk8s enable gpu --set toolkit.enabled=false

```

## Configure NVIDIA Multi-Instance GPU

[NVIDIA Multi-Instance GPU (MIG)](https://www.nvidia.com/en-us/technologies/multi-instance-gpu/) expands the performance and value of NVIDIA [H100](https://www.nvidia.com/en-us/data-center/h100/), [A100](https://www.nvidia.com/en-us/data-center/a100/) and [A30](https://www.nvidia.com/en-us/data-center/products/a30-gpu/) Tensor Core GPUs. MIG can partition the GPU into as many as seven instances, each fully isolated with its own high-bandwidth memory, cache, and compute cores. This allows for serving workloads under guaranteed quality of service (QoS) while extending the reach of accelerated computing resources to every user.

After enabling the GPU addon in MicroK8s on a host with an NVIDIA GPU that supports MIG, the GPU operator will automatically deploy the `nvidia-mig-manager` daemonset on the cluster. Configuring the GPU card on the node to enable MIG is done by setting an appropriate label on the Kubernetes node.

### Enable MIG

1. First, ensure that your GPU card has support for MIG. If that is the case, then `nvidia-mig-manager` should be running in the cluster, and the node should have a `nvidia.com/mig.available=true` label. Verify this with:

2. Set the `nvidia.com/mig.config` label on the node with the MIG configuration you want to apply. In our example, we have an NVIDIA A100 40GB card, and we will use the `all-1g.5gb` profile, which segments an NVIDIA A100 card to 7 `1g.5gb` GPU instance profiles:

3. `mig-manager` will report the result by setting the `nvidia.com/mig.config.state` label on the node. Check it with `microk8s kubectl describe node $node | grep nvidia.com`. If the configuration has been successful, the labels should look like this:

4. Finally, use `nvidia-smi` to verify that 7 GPU instances are now available for use:

## Features

### GPU addon features

- Use the existing NVIDIA host drivers, or build the drivers and load to the kernel dynamically at runtime.
- Automatically install and configure the `nvidia-container-runtime` for containerd.
- Configure the `nvidia.com/gpu` kubelet device plugin, to support resource capacity and limits on GPU nodes.
- Multi-instance GPU (MIG) can be configured using ConfigMap resources.

### GPU addon components

The GPU addon will install and configure the following components on the MicroK8s cluster:

- `nvidia-feature-discovery`: Runs feature discovery on all cluster nodes, to detect GPU devices and host capabilities.
- `nvidia-driver-daemonset`: Runs in all GPU nodes of the cluster, builds and loads the NVIDIA drivers into the running kernel.
- `nvidia-container-toolkit-daemonset`: Runs in all GPU nodes of the cluster. Once the NVIDIA drivers are loaded, installs the `nvidia-container-runtime` binaries and configures the `nvidia` runtime on containerd accordingly. Prior to 1.36, it sets the default runtime to `nvidia`, so all pod workloads can use the GPU. Starting MicroK8s release 1.36, the GPU operator determines the default runtime, and workloads should use `runtimeClassName: nvidia`.

> _NOTE_: Starting with GPU Operator `v25.10.0` (MicroK8s 1.35+),  
> [CDI (Container Device Interface)](https://github.com/cncf-tags/container-device-interface/blob/main/SPEC.md)  
> is enabled by default. CDI is transparent for standard workloads — pods with  
> `runtimeClassName: nvidia` are unaffected. However, containers that access  
> GPUs directly via `NVIDIA_VISIBLE_DEVICES` (e.g. monitoring agents) must also  
> explicitly set `runtimeClassName: nvidia` in their pod spec.

- `nvidia-device-plugin-daemonset`: Runs in all GPU nodes of the cluster, and configures the `nvidia.com/gpu` kubelet device plugin. This is used to configure resource capacity and limits for the GPU nodes.
- `nvidia-operator-validator`: Validates that the NVIDIA drivers, container runtime and the kubelet device plugin have been configured correctly. Finally, it executes an example cuda workload.

A complete installation of the GPU operator looks like this (output of `microk8s kubectl get pod -n gpu-operator-resources`):

```auto
NAME READY STATUS RESTARTS AGE IP NODE NOMINATED NODE READINESS GATES
nvidia-container-toolkit-daemonset-mjbk8 1/1 Running 0 110m 10.1.51.198 machine-0 <none> <none>
nvidia-cuda-validator-xj2kx 0/1 Completed 0 109m 10.1.51.204 machine-0 <none> <none>
nvidia-dcgm-nvqnz 1/1 Running 0 110m 10.1.51.199 machine-0 <none> <none>
gpu-feature-discovery-dn6lt 1/1 Running 0 110m 10.1.51.202 machine-0 <none> <none>
nvidia-device-plugin-daemonset-zg76f 1/1 Running 0 110m 10.1.51.201 machine-0 <none> <none>
nvidia-device-plugin-validator-k6hdv 0/1 Completed 0 107m 10.1.51.205 machine-0 <none> <none>
nvidia-dcgm-exporter-9vnc5 1/1 Running 0 110m 10.1.51.203 machine-0 <none> <none>
nvidia-operator-validator-ntvdj 1/1 Running 0 110m 10.1.51.200 machine-0 <none> <none>

```

## MicroK8s 1.21

MicroK8s version 1.21 is out of support since May 2022. The GPU addon included with MicroK8s 1.21 was an early alpha and is no longer functional.

Due to a problem with the way containerd is configured in MicroK8s versions 1.21 and older, the `nvidia-toolkit-daemonset` installed by the GPU operator is incompatible and leaves MicroK8s in a broken state.

It is recommended to update to a supported version of MicroK8s. However, it is possible to install the GPU operator by following the steps described in [this GitHub gist](https://gist.github.com/neoaggelos/a4244cc92f76599b0d8febfeae3a28a8).

---

<div class="post-metadata">

**Author:** ![geosp](https://sea2.discourse-cdn.com/flex016/user_avatar/discuss.kubernetes.io/geosp/32/5814_2.png) [@geosp](https://discuss.kubernetes.io/u/geosp)\
**Post date:** [September 4, 2020, 10:23pm UTC](https://discuss.kubernetes.io/t/add-on-gpu/11286/2 "2020-09-04T22:23:06Z")

</div>

You should mention in the documentation that the runtime needs to be docker not containerd and therefore one should change the kubelet container runtime to docker like this:

Modify /var/snap/microk8s/current/args/kubelet:  
–container-runtime=docker  
–container-runtime-endpoint=${SNAP\_COMMON}/run/docker.sock

---

<div class="post-metadata">

**Author:** ![kjackal](https://sea2.discourse-cdn.com/flex016/user_avatar/discuss.kubernetes.io/kjackal/32/1750_2.png) [@kjackal](https://discuss.kubernetes.io/u/kjackal)\
**Post date:** [September 7, 2020, 12:50pm UTC](https://discuss.kubernetes.io/t/add-on-gpu/11286/3 "2020-09-07T12:50:44Z")

</div>

@geosp what MicroK8s version are you using? It has been some time since docker was removed in favor of containerd.

---

<div class="post-metadata">

**Author:** ![robotrapta](https://sea2.discourse-cdn.com/flex016/user_avatar/discuss.kubernetes.io/robotrapta/32/8571_2.png) [@robotrapta](https://discuss.kubernetes.io/u/robotrapta)\
**Post date:** [September 14, 2021, 10:07pm UTC](https://discuss.kubernetes.io/t/add-on-gpu/11286/4 "2021-09-14T22:07:54Z")

</div>

The “obviously” statement could be improved. It’s somewhat condescending, and glosses over details that many people find difficult.

Also, as of 1.21 I believe it’s really incomplete. Because if you just install the drivers and enable the add-on, it doesn’t work, with various pods in the `gpu-operator-resources` namespace complaining that `no runtime for "nvidia" is configured`.  
This problem is pretty well documented for docker, but in containerd land, I can’t figure out how to solve it.

---

<div class="post-metadata">

**Author:** ![John\_Grabner](https://sea2.discourse-cdn.com/flex016/user_avatar/discuss.kubernetes.io/john_grabner/32/6222_2.png) [@John\_Grabner](https://discuss.kubernetes.io/u/John_Grabner)\
**Post date:** [September 24, 2021, 9:26pm UTC](https://discuss.kubernetes.io/t/add-on-gpu/11286/5 "2021-09-24T21:26:11Z")

</div>

To use a GPU in docker-compose, my deployment.yaml looks like

```auto
  transcribe: 
    image: transcribe
    shm_size: '8gb'
    deploy:             
        resources:          
          reservations:
            devices:
            - capabilities: [gpu]

```

What is the equivalent for microk8s when enabling gpu?

---

<div class="post-metadata">

**Author:** ![kjackal](https://sea2.discourse-cdn.com/flex016/user_avatar/discuss.kubernetes.io/kjackal/32/1750_2.png) [@kjackal](https://discuss.kubernetes.io/u/kjackal)\
**Post date:** [September 25, 2021, 4:19am UTC](https://discuss.kubernetes.io/t/add-on-gpu/11286/6 "2021-09-25T04:19:46Z")

</div>

In the 1.21 release the gpu operator was in an alpha state so it may give you a hard time. Could you please try the 1.22/stable track?

---

<div class="post-metadata">

**Author:** ![kjackal](https://sea2.discourse-cdn.com/flex016/user_avatar/discuss.kubernetes.io/kjackal/32/1750_2.png) [@kjackal](https://discuss.kubernetes.io/u/kjackal)\
**Post date:** [September 25, 2021, 4:21am UTC](https://discuss.kubernetes.io/t/add-on-gpu/11286/7 "2021-09-25T04:21:43Z")

</div>

We use the manifest in [microk8s/cuda-add.yaml at master · ubuntu/microk8s · GitHub](https://github.com/ubuntu/microk8s/blob/master/tests/templates/cuda-add.yaml) to test the GPU.

---

<div class="post-metadata">

**Author:** ![John\_Grabner](https://sea2.discourse-cdn.com/flex016/user_avatar/discuss.kubernetes.io/john_grabner/32/6222_2.png) [@John\_Grabner](https://discuss.kubernetes.io/u/John_Grabner)\
**Post date:** [September 25, 2021, 3:03pm UTC](https://discuss.kubernetes.io/t/add-on-gpu/11286/8 "2021-09-25T15:03:35Z")

</div>

The example given appears to be only for vector add (I’m guessing based on the title of “[k8s.gcr.io/cuda-vector-add:v0.1](http://k8s.gcr.io/cuda-vector-add:v0.1)”).

For docker / docker-compose, I install on the host a driver from [Download Drivers | NVIDIA](https://www.nvidia.com/Download/index.aspx) and select 460. This is all I need to do to give deep learning pytorch / tensorflow in containers the full capability of the gpu. A lot more than vector add. In fact, docker is pitched in pytorch as a way to avoid the need to understand and deal with a bunch of cuda and stuff on your host, just install the driver and let the container deal with everything else.

Are there any examples of pytorch or tensorflow in kubernetes containers using the GPU?

---

<div class="post-metadata">

**Author:** ![robotrapta](https://sea2.discourse-cdn.com/flex016/user_avatar/discuss.kubernetes.io/robotrapta/32/8571_2.png) [@robotrapta](https://discuss.kubernetes.io/u/robotrapta)\
**Post date:** [September 26, 2021, 2:51am UTC](https://discuss.kubernetes.io/t/add-on-gpu/11286/9 "2021-09-26T02:51:54Z")

</div>

It works in 1.22, thanks. It also worked fine in 1.20, which made 1.21 extra confusing/frustrating.

---

<div class="post-metadata">

**Author:** ![robotrapta](https://sea2.discourse-cdn.com/flex016/user_avatar/discuss.kubernetes.io/robotrapta/32/8571_2.png) [@robotrapta](https://discuss.kubernetes.io/u/robotrapta)\
**Post date:** [September 26, 2021, 2:57am UTC](https://discuss.kubernetes.io/t/add-on-gpu/11286/10 "2021-09-26T02:57:15Z")

</div>

If vector add works, pytorch and tensorflow will work just fine. I’ve been doing this a long time, and it’s always been all-or-nothing with the cuda support inside containers. You might end up with a driver that’s not quite as fast as another, but unless you’re measuring carefully you wouldn’t notice.

That said, if you want to see if pytorch works, just run

```auto
python -c "import torch; torch.cuda.is_available()"

```

That almost always works. (Occassionally it can get fooled into thinking there’s a working GPU when there isn’t - like if your pytorch build was compiled without support for your cuda-compute level.)

If you’re using TensorFlow, I’m sorry.

---

<div class="post-metadata">

**Author:** ![John\_Grabner](https://sea2.discourse-cdn.com/flex016/user_avatar/discuss.kubernetes.io/john_grabner/32/6222_2.png) [@John\_Grabner](https://discuss.kubernetes.io/u/John_Grabner)\
**Post date:** [September 28, 2021, 11:59am UTC](https://discuss.kubernetes.io/t/add-on-gpu/11286/11 "2021-09-28T11:59:13Z")

</div>

Wow, It workes wonderfully!  
I misread ‘image: “[k8s.gcr.io/cuda-vector-add:v0.1](http://k8s.gcr.io/cuda-vector-add:v0.1)”’ as some sort of cuda driver, but I now see it’s just a test program to validate cuda.

---

<div class="post-metadata">

**Author:** ![fzhan](https://sea2.discourse-cdn.com/flex016/user_avatar/discuss.kubernetes.io/fzhan/32/9714_2.png) [@fzhan](https://discuss.kubernetes.io/u/fzhan)\
**Post date:** [March 19, 2022, 4:50pm UTC](https://discuss.kubernetes.io/t/add-on-gpu/11286/12 "2022-03-19T16:50:31Z")

</div>

I’m using the 3006 stable for 1.22, but the error still says “no runtime for nvidia is configured”
