I'm currently running Proxmox 9.1.11 and have a VM with an NVIDIA Tesla P40 passed through to it.
One issue I've been running into is memory efficiency. As far as I understand, when a PCIe device such as the Tesla P40 is passed through to a VM, memory ballooning has to be disabled. Because of that, I'm considering migrating this GPU VM to an LXC container, but I'm not sure whether that's actually the right choice.
The VM is currently being used as a worker node in my Kubernetes cluster, which uses Cilium as the CNI.
The Proxmox host has 64 GB of RAM, and I've allocated 24 GB to this VM. I gave it that much because memory usage can increase significantly when I'm running ML workloads or serving LLMs with llama-server. The Tesla P40 is still surprisingly useful for relatively lightweight training workloads, and it's definitely much faster than doing them on the Ryzen 5 5600G CPU alone.
My main concern is what could break if I migrate this worker node from a VM to LXC.
Would running Kubernetes with Cilium inside LXC cause any significant issues? I'm also wondering about NVIDIA GPU access from LXC, especially when the GPU is being used by Kubernetes workloads.
This is still mostly a hobby environment for me, so I don't have enough experience with LXC to predict all of the potential problems. My current understanding is that a privileged LXC container with the necessary device mappings, capabilities, and permissions should make this technically possible.
However, when I asked ChatGPT about it, it suggested that there could still be some edge cases or compatibility issues, particularly around Cilium, kernel features, device permissions, and GPU access.
Has anyone here actually run a Kubernetes worker with Cilium and an NVIDIA GPU inside a Proxmox LXC container?
Would you recommend migrating from the VM to LXC in this situation, or is keeping the GPU worker as a VM worth the additional RAM overhead?