[HELP] Running llm models on q6a with NPU

I’m trying to get a 3-class YOLO-style detection model (fire / smoke / people) running on NPU Q6A, and have hit the same wall through a few different paths… i hope someone here has done this successfully and can point out what am missing…..

What I’ve tried:

  • Local QAIRT SDK toolchain (qairt-converter → qairt-quantizer → qnn-context-binary-generator).
    Reproducible crashes/segfaults and missing-symbol errors at the context-binary generation
    step, on the exact same model that converts and quantizes cleanly earlier in the pipeline.
  • Qualcomm AI Hub (cloud compile/quantize/link). Gets further, but the real problem shows up
    here: the model’s box-regression and class-confidence outputs get forced onto a shared
    per-tensor INT8 quantization scale, which destroys precision for whichever output has the
    smaller numeric range. Splitting box/confidence into separate output tensors fixes that part.
  • Splitting further into one tensor per class (3 classes) is where I’m stuck two of the
    three class channels come back as literal zero after quantization, reproducible with two
    different ONNX ops (Split and Slice), so it doesn’t look like an op choice issue.

Is there any easy way for a multi class yolo model or any ai model to run on npu of q6a ???
Any pointers appreciated