self-hosted/ai
§01·model · /models

Qwen3 30B-A3B

llmactiveApache-2.0

30B Mixture-of-Experts LLM by Alibaba — Qwen3-30B-A3B, roughly 3B active per token, with hybrid thinking modes. That design is what puts a 30B on a 12 GB RTX 4070 at all; four cards carry a recipe here. At Q4_K it is recorded at 153.6 tokens/s on a 3090 and 226.1 on a 5090, against 35.1 and 61.4 for the dense Qwen3 32B on those same cards. Apache-2.0.

Download· 4 variants
§02·same family
5 other models · by name

Other models grouped with Qwen3 30B-A3B. Sizes and modalities can differ, and nothing here says whether one fits your GPU — open a model for its own compatibility table.

§03·GPUs that run this model
4 total

benchmarked·~ runs via recipe (not benchmarked)· untested·doesn't fit