Kiln - Let LLM + vision model run on the RK3576 NPU, mainline kernel (ROCK 4D)

Made this for the ROCK 4D: LLM + vision on the RK3576 NPU, on armbian with mainline kernel instead of the vendor 6.1 BSP

It’s the vendor rknpu driver built out-of-tree on mainline-based linux-7.1.3 (mainline linux with a small NPU patch set, not stock mainline), driven by the RKLLM and RKNN runtimes. On a ROCK 4D I get Llama-3.2-1B around 13 tok/s (Qwen2.5-1.5B is around 9), and MobileNet vision around 150 fps, both on the NPU on the same kernel. There’s an OpenAI-compatible API and a slash-command chat CLI on top of it.

Install on Armbian is one line. It pulls the mainline kernel, then after a reboot the driver, runtimes and tools:

curl -fsSL https://raw.githubusercontent.com/gahingwoo/kiln/main/scripts/kiln-install.sh | bash

Script’s in the repo if you want to read it before running it. There’s a flashable buildroot image too if you’d rather not touch your current setup, it’s in the README.

It’s the closed path though, out-of-tree vendor driver and closed runtimes, so it won’t go upstream and it’s not the open Mesa/Teflon answer. Runs today if that’s all you need.

logs: Kiln on Radxa ROCK 4D (RK3576), pure mainline linux-7.1.3 + Armbian: NPU LLM (Qwen2.5-1.5B & Llama-3.2-1B, live /model switch, multi-turn, 8-13 tok/s) + MobileNet vision (170 fps). Vendor rknpu 0.9.8 + librkllmrt/librknnrt on a mainline kernel. · GitHub

Also chipping away at the open driver (linux-rk3576-npu) on the side, different route to the same place.

Feedback welcome, still rough in spots.