Data point: Qwen3.8-Flash-Next PP/TG speed on M3 Ultra
r/LocalLLaMA
September 14, 2026 at 12:30 AM
📌 Model File: https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF Q4_K_XL llama.cpp configuration through llama-swap: -c 256000 --jinja --temp 1 --top-p 0.95 --top-k 20 --min-p 0.00 --presence-penalty 0 --repeat-penalty 1.0 Testing by: llama-benchy Results: | model | test | t/s | peak t/s | ttfr (ms) | est_ppt (ms) | e2e_ttft (ms) | |:-------------------|---------------:|--------------:|-------------:|----------------:|----------------:|----------------:| | qwen3.8-flash-next | pp1000 @ d10...
Model File: https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF Q4_K_XL
llama.cpp configuration through llama-swap:
-c 256000 --jinja --temp 1 --top-p 0.95 --top-k 20 --min-p 0.00 --presence-penalty 0 --repeat-penalty 1.0 Testing by: llama-benchy
Results:
| model | test | t/s | peak t/s | ttfr (ms) | est_ppt (ms) | e2e_ttft (ms) | |:-------------------|---------------:|--------------:|-------------:|----------------:|----------------:|----------------:| | qwen3.8-flash-next | pp1000 @ d1000 | 558.83 ± 2.61 | | 3312.36 ± 38.89 | 3300.16 ± 38.89 | 3312.36 ± 38.89 | | qwen3.8-flash-next | tg500 @ d1000 | 31.05 ± 0.14 | 31.67 ± 0.47 | | | | [link] [comments]
Read the full article at
r/LocalLLaMA
More in Models & Releases
Models & Releases
NVIDIA Unveils RTX PRO 5500 "Blackwell" Workstation GPU with 84 GB GDDR7 Memory
submitted by /u/Lumpy_Phase_9539 [link] [comments]
What happened to AI browsers?
AI browsers were supposedly going to kill Chrome/Safari, and help people quickly buy stuff, book tickets etc. at reasonable prices while avoiding ads. That was such a great sales pitch, but it has been just radio silence from then on. Ar...
[Bi-Weekly Megathread] Project Showcase
Do you have something you'd like to share with the r/LocalLLaMA community. This is the place for it! Recommendation on presentation: Please share plain english description of what your project does and why people should care about it Ho...