Hello everyone,
I would like to share an experimental, unofficial port of Qwen3.5-0.8B for the Radxa Dragon Q6A (QCS6490 / HTP v68).
Language-model prefill and decode, including the output head, and the complete vision encoder now run on the NPU through the standard ONNX Runtime QNN Execution Provider, with CPU EP fallback disabled. No project-specific ONNX Runtime build is required.
For clarity, this does not mean zero CPU usage: tokenization, image preprocessing, token embedding lookup, state encoding/transfers, and sampling still run on the CPU. The model sessions use session.disable_cpu_ep_fallback=1, and the vision startup check also verifies QNN-only node placement through ORT profiling.
Repository:
Current results
The release supports text chat and single-image questions, with a recommended 2K context profile and an optional 4K profile. Images are processed at a fixed 1024 × 576 resolution, producing 576 image tokens.
Recent measurements on my Q6A:
| Metric | 2K, recommended | 4K, optional |
|---|---|---|
| Decode | About 7.4 tokens/s | About 7.9 tokens/s |
| Chunk16 prefill | About 82 tokens/s | About 63 tokens/s |
These are workload-specific measurements, not throughput guarantees. The prefill figures cover the chunk-prefill stage only, excluding some input preparation and the sequential suffix described below. A higher decode rate does not necessarily mean a shorter time to the first token.
The tested environment is Ubuntu 24.04.4, ONNX Runtime 1.27.0, and QNN EP 2.3.0.
Main changes made during the port
-
Adapted the hybrid Gated DeltaNet/attention graphs for HTP v68. This involved fixed-shape graphs, stable primitive replacements for unsupported operations such as Softplus, and changes to dynamic matrix products and their quantization.
-
Addressed several numerical issues. I corrected weight-quantization settings, widened the final logit range to avoid clipping, and applied device-measured zero-input offset corrections, including 50 integer bias corrections in the vision encoder. These corrections are specific to the tested backend; I have not established whether the underlying offset behavior originates in the SDK or hardware.
-
Repaired recurrent-state and image-input handling. State inputs and outputs use separate buffers, and the final partial prefill chunk preserves the actual visual embeddings and multimodal rotary positions. A short sequential suffix helps reduce observed chunk-prefill errors, at a latency cost.
-
Reduced memory use and runtime overhead. Decode and the 24 prefill graphs share compiled weights, with state managed by a resident C++ host. Revised grouped-query attention layouts improved prefill throughput, while exact-prefix reuse and an output-preserving sampler optimization removed additional repeated work.
Remaining limitations
This is still a work in progress, and I would not claim FP32-equivalent accuracy. Known issues remain in OCR, visual interpretation, instruction following, and Japanese phrasing. Some answers are incorrect even when generation finishes normally, and 4K is not consistently better than 2K.
The repository includes known failures alongside the improvements, as well as English/Japanese documentation, setup instructions, hash-verified model downloads, and tools for rebuilding from the frozen quantized ONNX graphs.
I hope the implementation and porting notes may be useful to others experimenting with the Q6A’s NPU. I would be very grateful for feedback, reproduction reports, or suggestions, particularly on the remaining numerical issues.
Thank you for taking the time to read this.