§01·compatibility · /check
Qwen3.6 35B-A3B on RTX 4070
Yes — Qwen3.6 35B-A3B runs on the RTX 4070 (12 GB). Fastest community-measured result: 80.8 tokens/s.
✓ runsmultimodalactive40 series12GB VRAM
step-by-step recipe for this pair
Qwen3.6-35B-A3B on RTX 4070: 80 tok/s + 128K context from a 12GB card via MoE CPU-offload
llmadvanced12GB+
model
- name
- Qwen3.6 35B-A3B
- slug
- qwen3-6-35b-a3b-mtp-ud-q4-k-xl-gguf
- vertical
- multimodal
- status
- active
- repo
- huggingface.co ↗
§02·benchmarks
| Task | Quant | Speed | VRAM | Works | Confidence | Source | Verified |
|---|---|---|---|---|---|---|---|
| llm | Q4_K_XL | 80.8tokens/s | ✓ | reddit.com· reddit | 2026-05-13 |
§03·how it was measured
- llmQ4_K_XL80.8 tokens/s
Benchmark: code_python. Achieved using llama.cpp with MTP support and 128K context. Using -fitt 1536 to balance GPU/CPU load.
recorded from reddit.com ↗ · verified 2026-05-13
§04·more Qwen3.6 35B-A3B recipes
§05·common questions
Can you run Qwen3.6 35B-A3B on RTX 4070?
Yes — Qwen3.6 35B-A3B runs on the RTX 4070 (12 GB). Fastest community-measured result: 80.8 tokens/s.
Which quantizations have been tested for Qwen3.6 35B-A3B on RTX 4070?
Q4_K_XL — measured in community benchmarks.
How fast is Qwen3.6 35B-A3B on RTX 4070?
Up to 80.8 tokens/s (llm), the fastest community-measured result.
Are there step-by-step instructions for Qwen3.6 35B-A3B on RTX 4070?
Yes — a step-by-step recipe documents Qwen3.6 35B-A3B on the RTX 4070, linked at the top of this page.