Case study · Linux troubleshooting
Recovering GPU passthrough after a kernel upgrade broke DKMS
A routine upgrade of the OpenMediaVault VM left the NVIDIA driver unable to build against the new kernel. The passed-through RTX A2000 stopped working in the guest, GPU-backed Ollama workloads went down and the Docker environment sharing the VM was at risk. This is how the fault was isolated and recovered without reinstalling blind.
At a glance
Key facts
- Platform
- Proxmox host → OpenMediaVault VM
- GPU
- NVIDIA RTX A2000, PCI passthrough
- Affected
- Ollama, Docker NVIDIA runtime
- Root cause
- DKMS module failed to build against the new kernel
- Recovery
- Known-good kernel booted and pinned; NVIDIA 550 driver
- Linux
- Proxmox
- OpenMediaVault
- DKMS
- NVIDIA
- Docker
- Ollama
Context
OpenMediaVault runs as a virtual machine on a Proxmox host. It provides the mdadm RAID1 storage and SMB shares for the home network and is also the main Docker host, running DNS, reverse proxy, monitoring and an Ollama service that uses an NVIDIA RTX A2000 passed through from the host.
Because storage, containers and the GPU all live in the same VM, an upgrade problem there has a wide blast radius. Anything that destabilises the guest affects every service on the network.
The change
A routine package upgrade of the VM pulled in a newer Linux kernel. The NVIDIA driver is installed as a DKMS module, which means it has to be rebuilt against every new kernel at install time. On this upgrade, that rebuild did not succeed.
Symptoms
- After the reboot the GPU was no longer usable inside the guest.
- GPU-accelerated Ollama workloads stopped.
- The Docker NVIDIA runtime could not access the device.
- Storage and the other containers appeared to be running, but their state had not yet been verified.
Investigation
The temptation with driver problems is to reinstall components until something works. The aim here was to establish which layer had failed before changing anything.
- Reviewed the system logs and the DKMS build logs.
- Checked the installed kernel version against the kernel the driver module had last been built for.
- Ran nvidia-smi to confirm whether the driver could see the GPU at all.
- Verified the Docker NVIDIA runtime separately from the host, so a container-level problem would not be mistaken for a host-level one.
- Confirmed that the Proxmox PCI passthrough configuration was unchanged and the device was still presented to the VM.
Root cause
The passthrough was intact. The DKMS build of the NVIDIA module had failed against the newer kernel, so the driver never loaded and nothing above it, from the container runtime to the model, could work. It was a kernel and driver compatibility problem, not a hardware or virtualisation problem.
Recovery
- Booted the VM on the last known-good kernel, where the existing module built and loaded.
- Pinned that kernel so an automatic upgrade could not reintroduce the fault.
- Confirmed the NVIDIA 550 driver loaded and that nvidia-smi reported the RTX A2000.
- Made no further changes until a stable recovery point had been confirmed.
Validation
- Tested the GPU from the host operating system.
- Tested from inside the Ollama container and confirmed models loaded into GPU memory.
- Checked the other Docker services individually.
- Checked the storage volumes and SMB access.
The system was not declared recovered until storage, containers, network access and the GPU workload had each been tested. The people relying on the hosted services were kept informed of the outage, its scope and when service was restored.
What changed afterwards
- Kernel and driver compatibility is checked before an upgrade, not after.
- A known-good kernel is retained rather than letting the package manager remove it.
- Working kernel and driver versions are recorded.
- Critical services are validated stage by stage after any change.
- No unrelated changes are made until a stable recovery point is confirmed.
The broader lesson is that an upgrade on a system with kernel-dependent drivers and passed-through hardware cannot be treated as a single package-management step.
Scope
What this does not claim
- This was a single-host homelab incident, not a production outage with formal change control.
- The upgrade process described is a working practice, not an automated pipeline.
More
Other case studies
Cloud administration
Azure administration lab: CLI provisioning, cost controls and private storage with RBAC
Read
Storage
Storage platform: mdadm RAID1, SMB and a validated disk migration
Read
Networking & services
DNS, HTTPS and service routing with Pi-hole, Unbound, Nginx Proxy Manager and Cloudflare Tunnel
Read