self-hosted/ai
§01·model · /models

Qwen3 32B

llmactiveApache-2.0

32B dense LLM by Alibaba with hybrid thinking and non-thinking modes. Seven machines carry a recipe here from a 24 GB floor, three of them Apple — M3 Max, M4 Max and M2 Max — so there is a Metal route as well as a CUDA one. Recorded at 35.1 tokens/s on a 3090 and 61.4 on a 5090 at Q4_K. Apache-2.0.

Download· 4 variants
§02·same family
5 other models · by name

Other models grouped with Qwen3 32B. Sizes and modalities can differ, and nothing here says whether one fits your GPU — open a model for its own compatibility table.

§03·GPUs that run this model
7 total

benchmarked·~ runs via recipe (not benchmarked)· untested·doesn't fit