self-hosted/ai
§01·compatibility · /check

Qwen3.6 35B-A3B on RTX 3060

Yes — Qwen3.6 35B-A3B runs on the RTX 3060 (12 GB). Fastest community-measured result: 38.9 tok/s, with 9.8 GB peak VRAM.

✓ runsmultimodalactive30 series12GB VRAM
step-by-step recipe for this pair

Qwen3.6-35B-A3B on RTX 3060: 38.9 tok/s from a 12GB card via MoE expert CPU-offload

llmadvanced12GB+
model
name
Qwen3.6 35B-A3B
slug
qwen3-6-35b-a3b-mtp-ud-q4-k-xl-gguf
vertical
multimodal
status
active
open detail ↗
gpu
name
RTX 3060
slug
rtx-3060
vram
12 GB
series
30
open detail ↗
§02·benchmarks
TaskQuantSpeedVRAMWorksConfidenceSourceVerified
llmQ4_K_M (Unsloth UD)38.9tok/s9.8GB✓smeltcore.com· community2026-08-08
§03·how it was measured
  • llmQ4_K_M (Unsloth UD)38.9 tok/s

    Expert-offload run: routed experts pushed to system RAM via --n-cpu-moe, attention + shared weights on GPU. llama.cpp b10088 (67b9b0e), CUDA sm_86, flash-attn on. Command: llama-bench -m MODEL -ngl 99 -ncmoe 24 -fa 1 -p 512 -n 128 -d 0,4096,8192 -r 3 Rig: i7-7700, 32GB DDR4-2133 dual-channel (~34 GB/s), PCIe 3.0 x16, headless. Generation speed is FLAT through 8K context: 38.9 / 38.5 / 38.2 tok/s at 0 / 4K / 8K (pp512: 413 / 394 / 386). Expert-streaming cost over DDR4 is constant per token. Offload sweep (empty ctx): -ncmoe 40 = 28.1 tok/s @ 2.5GB VRAM; -ncmoe 32 = 31.8 @ 6.1GB; -ncmoe 24 = 38.9 @ 9.8GB (recommended); -ncmoe 20 = 42.6 @ 11.7GB but OOMs once real context loads. Caveat: headless card. A 3060 also driving a display loses ~0.5-1GB and needs +1-2 on -ncmoe. Recommend the 38.9 figure, not the 42.6 ceiling.

    recorded from smeltcore.com ↗ · verified 2026-08-08

§04·more Qwen3.6 35B-A3B recipes
§05·common questions
Can you run Qwen3.6 35B-A3B on RTX 3060?

Yes — Qwen3.6 35B-A3B runs on the RTX 3060 (12 GB). Fastest community-measured result: 38.9 tok/s, with 9.8 GB peak VRAM.

How much VRAM does Qwen3.6 35B-A3B need on RTX 3060?

Measured peak VRAM is 9.8 GB — about 2.2 GB of headroom on the 12 GB card.

Which quantizations have been tested for Qwen3.6 35B-A3B on RTX 3060?

Q4_K_M (Unsloth UD) — measured in community benchmarks.

How fast is Qwen3.6 35B-A3B on RTX 3060?

Up to 38.9 tok/s (llm), the fastest community-measured result.

Are there step-by-step instructions for Qwen3.6 35B-A3B on RTX 3060?

Yes — a step-by-step recipe documents Qwen3.6 35B-A3B on the RTX 3060, linked at the top of this page.