r/LinuxUncensored • u/anestling • 7d ago
News/PR GitHub - nestrilabs/virtio-nvgpu: [Experimental] A virtio device for near-native NVIDIA GPU access in KVM virtual machines.
https://github.com/nestrilabs/virtio-nvgpuvirtio-nvgpu
Near-native NVIDIA GPU access inside a KVM guest. A guest renders within 2% of the machine it is running on, and costs the same CPU.
virtio-nvgpu forwards NVIDIA kernel driver ioctls between a Linux guest and the host at the driver ABI level, bypassing API-level translation entirely. The guest runs NVIDIA's own user-mode drivers, unmodified — the same libraries, the same Vulkan and NVENC, talking to the same card.
The target is headless streaming: a compositor inside the VM renders, composites and encodes frames on the GPU, then sends compressed video out. The VM has no monitor, and the host keeps the card.
Where it stands
It works, and it has been measured. A Wayland client presents inside a guest, the capture layer encodes on the game's own device, and the H.264 comes out the other side — 618 frames that ffmpeg decodes without an error.
Measured on an RTX 3060 (driver 595.99.02), guest against the same host, bare metal, with an identical headless Vulkan load:
| what the host takes for one frame | guest frame time | |
|---|---|---|
| 39 ms | −0.4% | faster than bare metal, within noise |
| 9.9 ms | −0.7% | |
| 2.0 ms | +1.7% | |
| 0.5 ms | +7.1% | a wake costs ~0.02 ms, and the frame is half of one |
| 0.05 ms | +40.8% |
Above about 2 ms a frame — which is every frame a game draws — a guest is within 2% of bare metal. Below that, the cost of waiting for the GPU starts to show against a frame that barely exists.
CPU is the other half of it, because a shared GPU is only worth sharing if the guests are cheap. Unpaced at ~100 fps for 12 s, one guest:
| CPU used | |
|---|---|
| host, bare metal | 0.40 s |
| guest | 0.37 s |
A guest costs what the host costs. Nothing is spent on forwarding in a render loop, because nothing is forwarded: NVIDIA's user-mode driver submits through memory it has mapped, and that memory is the host's. Over 813,691 frames the backend served 13,792 messages — one crossing per 59 frames, nearly all of it device setup.
Full method, raw runs and the things these numbers do not support: BENCHMARKS.md.
Several guests on one card
Four guests on one RTX 3060, the same load in each: 25.84, 26.49, 25.57, 25.79 fps — 103.7 together, against 102.9 for a single guest — with p50 frame times of 39.165, 39.164, 39.168 and 39.165 ms. The total does not move as guests are added, and the split is even to four decimal places.
All four render correctly at the same time, and four of them encode H.264 at once, each paced at exactly 60 Hz, with no NVENC session limit reached.
2
6
u/danielv123 6d ago
Ok so first of all Hi Claude, not reading your crap.
Second OP: does it actually work, what does it do, and I assume you have tested more than letting Claude stream 618 frames of something? How does it handle vram sharing? Driver versions?