self-hosted/ai
§01·model · /models

Qwen3-8B

llmactiveApache-2.0

8B LLM by Alibaba (Qwen3) with hybrid thinking / non-thinking modes, Apache-2.0. Recipes on 20 cards here, smallest 12 GB. Measured at 200.4 tok/s on an RTX 5090 and 129.1 tok/s on a 16 GB RTX 5080 — fast enough that the bottleneck in a chat UI stops being the model. The usual reason to pick this over the 14B is not VRAM but headroom for context and a second process on the card.

Download· 4 variants
§02·same family
5 other models · by name

Other models grouped with Qwen3-8B. Sizes and modalities can differ, and nothing here says whether one fits your GPU — open a model for its own compatibility table.

§03·GPUs that run this model
20 total

benchmarked·~ runs via recipe (not benchmarked)· untested·doesn't fit