Home / Models & Releases / Article
Models & Releases

Hot Expert Reload on GPU is what this community needs

R/

r/LocalLLaMA

September 11, 2026 at 12:25 PM

πŸ“Œ A huge favor to ask llama maintainers - please implement this feature. Even with one 3090 card there will be tangible improvements in decode speed on MOE models with moderate number of active parameters, like Qwen3.8-Flash-Next, Deepseek V4/V4.1 Flash, GLM 5.3 Flash. With 2x 3090 speeds will be quite close to the full offload of these models to VRAM. This will make these almost SOTA models really usable locally. submitted by /u/perelmanych [link] [comments]

A huge favor to ask llama maintainers - please implement this feature. Even with one 3090 card there will be tangible improvements in decode speed on MOE models with moderate number of active parameters, like Qwen3.8-Flash-Next, Deepseek V4/V4.1 Flash, GLM 5.3 Flash. With 2x 3090 speeds will be quite close to the full offload of these models to VRAM. This will make these almost SOTA models really usable locally.

submitted by /u/perelmanych
[link] [comments]

Read the full article at

r/LocalLLaMA

Visit Source β†—