§01·model · /models
Qwen3 30B-A3B
llmactiveApache-2.0
30B Mixture-of-Experts LLM by Alibaba — Qwen3-30B-A3B, roughly 3B active per token, with hybrid thinking modes. That design is what puts a 30B on a 12 GB RTX 4070 at all; four cards carry a recipe here. At Q4_K it is recorded at 153.6 tokens/s on a 3090 and 226.1 on a 5090, against 35.1 and 61.4 for the dense Qwen3 32B on those same cards. Apache-2.0.
Download· 4 variants
§02·same family
5 other models · by nameOther models grouped with Qwen3 30B-A3B. Sizes and modalities can differ, and nothing here says whether one fits your GPU — open a model for its own compatibility table.
§03·GPUs that run this model
4 total✓ benchmarked·~ runs via recipe (not benchmarked)·— untested·✕doesn't fit