self-hosted/ai
§01·compatibility · /check

Qwen3.6 35B-A3B on RTX 4070

Yes — Qwen3.6 35B-A3B runs on the RTX 4070 (12 GB). Fastest community-measured result: 80.8 tokens/s.

✓ runsmultimodalactive40 series12GB VRAM
step-by-step recipe for this pair

Qwen3.6-35B-A3B on RTX 4070: 80 tok/s + 128K context from a 12GB card via MoE CPU-offload

llmadvanced12GB+
model
name
Qwen3.6 35B-A3B
slug
qwen3-6-35b-a3b-mtp-ud-q4-k-xl-gguf
vertical
multimodal
status
active
open detail ↗
gpu
name
RTX 4070
slug
rtx-4070
vram
12 GB
series
40
open detail ↗
§02·benchmarks
TaskQuantSpeedVRAMWorksConfidenceSourceVerified
llmQ4_K_XL80.8tokens/s✓reddit.com· reddit2026-05-13
§03·how it was measured
  • llmQ4_K_XL80.8 tokens/s

    Benchmark: code_python. Achieved using llama.cpp with MTP support and 128K context. Using -fitt 1536 to balance GPU/CPU load.

    recorded from reddit.com ↗ · verified 2026-05-13

§04·more Qwen3.6 35B-A3B recipes
§05·common questions
Can you run Qwen3.6 35B-A3B on RTX 4070?

Yes — Qwen3.6 35B-A3B runs on the RTX 4070 (12 GB). Fastest community-measured result: 80.8 tokens/s.

Which quantizations have been tested for Qwen3.6 35B-A3B on RTX 4070?

Q4_K_XL — measured in community benchmarks.

How fast is Qwen3.6 35B-A3B on RTX 4070?

Up to 80.8 tokens/s (llm), the fastest community-measured result.

Are there step-by-step instructions for Qwen3.6 35B-A3B on RTX 4070?

Yes — a step-by-step recipe documents Qwen3.6 35B-A3B on the RTX 4070, linked at the top of this page.