Made this for the ROCK 4D: LLM + vision on the RK3576 NPU, on armbian with mainline kernel instead of the vendor 6.1 BSP
It’s the vendor rknpu driver built out-of-tree on mainline-based linux-7.1.3 (mainline linux with a small NPU patch set, not stock mainline), driven by the RKLLM and RKNN runtimes. On a ROCK 4D I get Llama-3.2-1B around 13 tok/s (Qwen2.5-1.5B is around 9), and MobileNet vision around 150 fps, both on the NPU on the same kernel. There’s an OpenAI-compatible API and a slash-command chat CLI on top of it.
Install on Armbian is one line. It pulls the mainline kernel, then after a reboot the driver, runtimes and tools:
curl -fsSL https://raw.githubusercontent.com/gahingwoo/kiln/main/scripts/kiln-install.sh | bash
Script’s in the repo if you want to read it before running it. There’s a flashable buildroot image too if you’d rather not touch your current setup, it’s in the README.
It’s the closed path though, out-of-tree vendor driver and closed runtimes, so it won’t go upstream and it’s not the open Mesa/Teflon answer. Runs today if that’s all you need.
Also chipping away at the open driver (linux-rk3576-npu) on the side, different route to the same place.
Feedback welcome, still rough in spots.