§01·model · /models
Qwen3.6 35B-A3B
multimodalactiveApache-2.0
35B Mixture-of-Experts LLM by Alibaba — Qwen3.6-35B-A3B, about 3B active per token, with multi-token prediction. The Unsloth Q4_K_XL GGUF is what turns it into a 12 GB proposition: both cards covered here are 12 GB boards, with 80.8 tokens/s recorded on an RTX 4070 and 38.9 tok/s on a 3060. Apache-2.0.
§02·GPUs that run this model
2 total✓ benchmarked·~ runs via recipe (not benchmarked)·— untested·✕doesn't fit