Qwen3.6 35B-A3B on RTX 3060
Yes — Qwen3.6 35B-A3B runs on the RTX 3060 (12 GB). Fastest community-measured result: 38.9 tok/s, with 9.8 GB peak VRAM.
Qwen3.6-35B-A3B on RTX 3060: 38.9 tok/s from a 12GB card via MoE expert CPU-offload
- name
- Qwen3.6 35B-A3B
- slug
- qwen3-6-35b-a3b-mtp-ud-q4-k-xl-gguf
- vertical
- multimodal
- status
- active
- repo
- huggingface.co ↗
| Task | Quant | Speed | VRAM | Works | Confidence | Source | Verified |
|---|---|---|---|---|---|---|---|
| llm | Q4_K_M (Unsloth UD) | 38.9tok/s | 9.8GB | ✓ | smeltcore.com· community | 2026-08-08 |
- llmQ4_K_M (Unsloth UD)38.9 tok/s
Expert-offload run: routed experts pushed to system RAM via --n-cpu-moe, attention + shared weights on GPU. llama.cpp b10088 (67b9b0e), CUDA sm_86, flash-attn on. Command: llama-bench -m MODEL -ngl 99 -ncmoe 24 -fa 1 -p 512 -n 128 -d 0,4096,8192 -r 3 Rig: i7-7700, 32GB DDR4-2133 dual-channel (~34 GB/s), PCIe 3.0 x16, headless. Generation speed is FLAT through 8K context: 38.9 / 38.5 / 38.2 tok/s at 0 / 4K / 8K (pp512: 413 / 394 / 386). Expert-streaming cost over DDR4 is constant per token. Offload sweep (empty ctx): -ncmoe 40 = 28.1 tok/s @ 2.5GB VRAM; -ncmoe 32 = 31.8 @ 6.1GB; -ncmoe 24 = 38.9 @ 9.8GB (recommended); -ncmoe 20 = 42.6 @ 11.7GB but OOMs once real context loads. Caveat: headless card. A 3060 also driving a display loses ~0.5-1GB and needs +1-2 on -ncmoe. Recommend the 38.9 figure, not the 42.6 ceiling.
recorded from smeltcore.com ↗ · verified 2026-08-08
Can you run Qwen3.6 35B-A3B on RTX 3060?
Yes — Qwen3.6 35B-A3B runs on the RTX 3060 (12 GB). Fastest community-measured result: 38.9 tok/s, with 9.8 GB peak VRAM.
How much VRAM does Qwen3.6 35B-A3B need on RTX 3060?
Measured peak VRAM is 9.8 GB — about 2.2 GB of headroom on the 12 GB card.
Which quantizations have been tested for Qwen3.6 35B-A3B on RTX 3060?
Q4_K_M (Unsloth UD) — measured in community benchmarks.
How fast is Qwen3.6 35B-A3B on RTX 3060?
Up to 38.9 tok/s (llm), the fastest community-measured result.
Are there step-by-step instructions for Qwen3.6 35B-A3B on RTX 3060?
Yes — a step-by-step recipe documents Qwen3.6 35B-A3B on the RTX 3060, linked at the top of this page.