llama.cpp ngram on RAM/SSD?
r/LocalLLaMA
September 11, 2026 at 10:39 AM
π I've been out of the loop for some time. Is there already an official way to offload ngram to RAM or SSD in something like Unsloth Studio? Interested in running Qwen3.8-Flash-Next on 72GB VRAM, but naiive attempts failed because even at Q4 it seems to load God knows what to God knows where. Would appreciate if someone share the recipe, or tell what are the official plans to support this (I can wait, knowing that it is upcoming). Thank you in advance. submitted by /u/NickNau [link] [...
I've been out of the loop for some time.
Is there already an official way to offload ngram to RAM or SSD in something like Unsloth Studio?
Interested in running Qwen3.8-Flash-Next on 72GB VRAM, but naiive attempts failed because even at Q4 it seems to load God knows what to God knows where.
Would appreciate if someone share the recipe, or tell what are the official plans to support this (I can wait, knowing that it is upcoming).
Thank you in advance.
[link] [comments]
Read the full article at
r/LocalLLaMA
More in Models & Releases
Models & Releases
Nvidia's RTX 5090 vanishes from online retail in the US β third-party sellers now demand as much as $9,500 for Nvidia's fastest GPU
submitted by /u/Norwood_Reaper_ [link] [comments]
Jensen Huang puts Trump on speakerphone onstage to announce robots wonβt take over the world
Nvidia CEO Jensen Huang took a call from President Trump on Monday while onstage at the All-In Podcast's All-In Summit. It's not the first time Huang has taken a call from the president during work, but this time he put Trump on speakerp...
Models & Releases
Base-10's Charlie O'Neill on why Kimi and GLM are "almost objectively" better than Opus 5
Edit: Spelled Baseten not Base-10 Full episode of this available at https://www.youtube.com/watch?v=PrSf7IOYu-I It's interesting to see how Dwarkesh has had to come around to the evidence that we are well on our way to creating AGI and e...
Models & Releases
NVIDIA Unveils RTX PRO 5500 "Blackwell" Workstation GPU with 84 GB GDDR7 Memory
submitted by /u/Lumpy_Phase_9539 [link] [comments]