33 signals
HOlO V1 IS LIVEone ranked AI digest a day, scored in publicREAD HOW IT WORKS →HOlO V2 STARTSyour X account, your signals, every day
signalDEV Community2026-09-29

How to pick --n-cpu-moe in llama.cpp: Qwen3.6 35B-A3B on 12, 16 and 24 GB GPUs

This article explains how to calculate the optimal --n-cpu-moe flag for llama.cpp when running MoE models like Qwen3.6 35B-A3B on consumer GPUs. It derives the required VRAM from GGUF tensor sizes and KV cache, providing concrete examples for 12, 16, and 24 GB cards. For instance, a 16 GB GPU with 32K context uses --n-cpu-moe 13, yielding 15.8 GiB VRAM and 31-53 tok/s.

for who
Developers and enthusiasts running large MoE models locally on consumer GPUs
why now
New Qwen3.6 MoE models now run on 12-16GB GPUs via llama.cpp's CPU offload flag.
what changes
They can now compute the correct offload setting from the GGUF file, avoiding trial-and-error restarts and optimizing VRAM usage and speed.
to do
Follow the step-by-step method or use the author's free planner to find the best --n-cpu-moe value for your GPU and context.
key points
  • Step-by-step: read tensor sizes from GGUF header, add KV cache, find smallest N that fits
  • Qwen3.6 35B-A3B: 40 layers, only 10 attention, KV cache 20 KB per token
  • Examples: 12 GB uses N=22 at 32K, 16 GB uses N=13, 24 GB uses N=0
#llama.cpp#moe#vram optimization#local deployment
score
score 9 out of 10. 0-10: how dense the facts are, multiplied by how much you can do with them after reading. 8+ means the topic's evidence bar is met: benchmarks and availability for a new model, amount and investors for a funding round, revenue figures for a solo-money story. Below 5 an item does not enter the digest. A press release scores 3 or less, a reprint loses 2, anything older than 14 days loses 1, a headline that misleads loses 3.
read the source