Decided to try out ik_llama.cpp as I’d heard there were some optimizations for CPU (particularly for Prompt Processing). The improvement is really dramatic for PP and for Qwen30B-A3B-Instruct-2507 Q4_M, it pretty much doubles the TPS.
How to Build
git clone https://github.com/ikawrakow/ik_llama.cpp
cd ik_llama
# THIS IS IMPORTANT! MUST SET ARCH FLAGS MANUALLY!
cmake -B ./build -DGGML_CUDA=OFF -DGGML_BLAS=OFF -DCMAKE_C_FLAGS="-march=armv9-a+sve2+dotprod+i8mm+fp16+fp16fml+crypto+sha2+sha3+sm4+rcpc+lse+crc+aes+memtag+sb+ssbs+predres+pauth" -DCMAKE_CXX_FLAGS="-march=armv9-a+sve2+dotprod+i8mm+fp16+fp16fml+crypto+sha2+sha3+sm4+rcpc+lse+crc+aes+memtag+sb+ssbs+predres+pauth" -DGGML_NATIVE=off
# Build using 7 threads
cmake --build ./build --config Release -j 7
Benchmarks
ik_llama.cpp
./llama-bench -m ../../../models/Qwen_Qwen3-30B-A3B-Instruct-2507-Q4_K_M.gguf -t 7
| model | size | params | backend | threads | test | t/s |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | ---------------: |
| qwen3moe ?B Q4_K - Medium | 17.35 GiB | 30.53 B | CPU | 7 | pp512 | 50.51 ± 0.27 |
| qwen3moe ?B Q4_K - Medium | 17.35 GiB | 30.53 B | CPU | 7 | tg128 | 15.44 ± 0.01 |
build: d99cf7cb (3836)
llama.cpp:
./llama-bench -m ../../../models/Qwen_Qwen3-30B-A3B-Instruct-2507-Q4_K_M.gguf -t 7
| model | size | params | backend | threads | test | t/s |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | --------------: | -------------------: |
| qwen3moe 30B.A3B Q4_K - Medium | 17.35 GiB | 30.53 B | CPU | 7 | pp512 | 24.33 ± 0.06 |
| qwen3moe 30B.A3B Q4_K - Medium | 17.35 GiB | 30.53 B | CPU | 7 | tg128 | 14.94 ± 0.03 |
build: cb9178f8 (5857)