self-hosted/ai
§01·model · /models

Qwen3.6 35B-A3B

multimodalactiveApache-2.0

35B Mixture-of-Experts LLM by Alibaba — Qwen3.6-35B-A3B, about 3B active per token, with multi-token prediction. The Unsloth Q4_K_XL GGUF is what turns it into a 12 GB proposition: both cards covered here are 12 GB boards, with 80.8 tokens/s recorded on an RTX 4070 and 38.9 tok/s on a 3060. Apache-2.0.

§02·GPUs that run this model
2 total

benchmarked·~ runs via recipe (not benchmarked)· untested·doesn't fit