Llama.cpp - Benchmarks

Decided to try out ik_llama.cpp as I’d heard there were some optimizations for CPU (particularly for Prompt Processing). The improvement is really dramatic for PP and for Qwen30B-A3B-Instruct-2507 Q4_M, it pretty much doubles the TPS.

How to Build

git clone https://github.com/ikawrakow/ik_llama.cpp
cd ik_llama

# THIS IS IMPORTANT! MUST SET ARCH FLAGS MANUALLY!
cmake -B ./build -DGGML_CUDA=OFF -DGGML_BLAS=OFF -DCMAKE_C_FLAGS="-march=armv9-a+sve2+dotprod+i8mm+fp16+fp16fml+crypto+sha2+sha3+sm4+rcpc+lse+crc+aes+memtag+sb+ssbs+predres+pauth" -DCMAKE_CXX_FLAGS="-march=armv9-a+sve2+dotprod+i8mm+fp16+fp16fml+crypto+sha2+sha3+sm4+rcpc+lse+crc+aes+memtag+sb+ssbs+predres+pauth" -DGGML_NATIVE=off

# Build using 7 threads
cmake --build ./build --config Release -j 7

Benchmarks

ik_llama.cpp

./llama-bench -m ../../../models/Qwen_Qwen3-30B-A3B-Instruct-2507-Q4_K_M.gguf -t 7       
| model                          |       size |     params | backend    | threads |          test |              t/s |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | ---------------: |
| qwen3moe ?B Q4_K - Medium      |  17.35 GiB |    30.53 B | CPU        |       7 |         pp512 |     50.51 ± 0.27 |
| qwen3moe ?B Q4_K - Medium      |  17.35 GiB |    30.53 B | CPU        |       7 |         tg128 |     15.44 ± 0.01 |

build: d99cf7cb (3836)

llama.cpp:

./llama-bench -m ../../../models/Qwen_Qwen3-30B-A3B-Instruct-2507-Q4_K_M.gguf -t 7
| model                          |       size |     params | backend    | threads |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | --------------: | -------------------: |
| qwen3moe 30B.A3B Q4_K - Medium |  17.35 GiB |    30.53 B | CPU        |       7 |           pp512 |         24.33 ± 0.06 |
| qwen3moe 30B.A3B Q4_K - Medium |  17.35 GiB |    30.53 B | CPU        |       7 |           tg128 |         14.94 ± 0.03 |

build: cb9178f8 (5857)