self-hosted/ai
§01·model · /models

Qwen3 14B

llmactiveApache-2.0

14B dense LLM by Alibaba (Qwen3) with hybrid thinking / non-thinking modes, Apache-2.0. Twenty cards in this catalogue carry a published recipe for it, and the smallest is a 12 GB board at 4-bit GGUF — that is the practical floor for a 14B, not a comfortable one. Measured here at 123.8 tok/s on an RTX 5090 and 80.6 tok/s on a 16 GB RTX 5080, so the middle of the range keeps interactive speed rather than merely fitting.

Download· 4 variants
§02·same family
5 other models · by name

Other models grouped with Qwen3 14B. Sizes and modalities can differ, and nothing here says whether one fits your GPU — open a model for its own compatibility table.

§03·GPUs that run this model
20 total

benchmarked·~ runs via recipe (not benchmarked)· untested·doesn't fit