Home / Models & Releases / Article
Models & Releases

Data point: Qwen3.8-Flash-Next PP/TG speed on M3 Ultra

R/

r/LocalLLaMA

September 14, 2026 at 12:30 AM

📌 Model File: https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF Q4_K_XL llama.cpp configuration through llama-swap: -c 256000 --jinja --temp 1 --top-p 0.95 --top-k 20 --min-p 0.00 --presence-penalty 0 --repeat-penalty 1.0 Testing by: llama-benchy Results: | model | test | t/s | peak t/s | ttfr (ms) | est_ppt (ms) | e2e_ttft (ms) | |:-------------------|---------------:|--------------:|-------------:|----------------:|----------------:|----------------:| | qwen3.8-flash-next | pp1000 @ d10...

Model File: https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF Q4_K_XL

llama.cpp configuration through llama-swap:

-c 256000 --jinja --temp 1 --top-p 0.95 --top-k 20 --min-p 0.00 --presence-penalty 0 --repeat-penalty 1.0 

Testing by: llama-benchy

Results:

| model | test | t/s | peak t/s | ttfr (ms) | est_ppt (ms) | e2e_ttft (ms) | |:-------------------|---------------:|--------------:|-------------:|----------------:|----------------:|----------------:| | qwen3.8-flash-next | pp1000 @ d1000 | 558.83 ± 2.61 | | 3312.36 ± 38.89 | 3300.16 ± 38.89 | 3312.36 ± 38.89 | | qwen3.8-flash-next | tg500 @ d1000 | 31.05 ± 0.14 | 31.67 ± 0.47 | | | | 
submitted by /u/rm-rf-rm
[link] [comments]

Read the full article at

r/LocalLLaMA

Visit Source ↗