Home / Models & Releases / Article
Models & Releases

llama.cpp ngram on RAM/SSD?

R/

r/LocalLLaMA

September 11, 2026 at 10:39 AM

πŸ“Œ I've been out of the loop for some time. Is there already an official way to offload ngram to RAM or SSD in something like Unsloth Studio? Interested in running Qwen3.8-Flash-Next on 72GB VRAM, but naiive attempts failed because even at Q4 it seems to load God knows what to God knows where. Would appreciate if someone share the recipe, or tell what are the official plans to support this (I can wait, knowing that it is upcoming). Thank you in advance. submitted by /u/NickNau [link] [...

I've been out of the loop for some time.

Is there already an official way to offload ngram to RAM or SSD in something like Unsloth Studio?

Interested in running Qwen3.8-Flash-Next on 72GB VRAM, but naiive attempts failed because even at Q4 it seems to load God knows what to God knows where.

Would appreciate if someone share the recipe, or tell what are the official plans to support this (I can wait, knowing that it is upcoming).

Thank you in advance.

submitted by /u/NickNau
[link] [comments]

Read the full article at

r/LocalLLaMA

Visit Source β†—