§01·model · /models
Qwen3 32B
llmactiveApache-2.0
32B dense LLM by Alibaba with hybrid thinking and non-thinking modes. Seven machines carry a recipe here from a 24 GB floor, three of them Apple — M3 Max, M4 Max and M2 Max — so there is a Metal route as well as a CUDA one. Recorded at 35.1 tokens/s on a 3090 and 61.4 on a 5090 at Q4_K. Apache-2.0.
Download· 4 variants
§02·same family
5 other models · by nameOther models grouped with Qwen3 32B. Sizes and modalities can differ, and nothing here says whether one fits your GPU — open a model for its own compatibility table.
§03·GPUs that run this model
7 total| GPU | VRAM | Series | Best speed | Min VRAM | Works | Evidence | |
|---|---|---|---|---|---|---|---|
| RTX 5090 | 32GB | 50 | 61.4tokens/s | ✓ | 3benchesrecipe | check ↗ | |
| RTX 3090 Ti | 24GB | 30 | 38tokens/s | 24GB | ✓ | 2benchesrecipe | check ↗ |
| RTX 3090 | 24GB | 30 | 35.1tokens/s | 24GB | ✓ | 2benchesrecipe | check ↗ |
| Apple M2 Max | 64GB | apple | ~ | recipe | check ↗ | ||
| Apple M3 Max | 48GB | apple | ~ | recipe | check ↗ | ||
| Apple M4 Max | 48GB | apple | ~ | recipe | check ↗ | ||
| RTX 4090 | 24GB | 40 | ~ | recipe | check ↗ |
- VRAM
- 32GB
- Best speed
- 61.4tokens/s
- Min VRAM
- Evidence
- 3benchesrecipe
- VRAM
- 24GB
- Best speed
- 38tokens/s
- Min VRAM
- 24GB
- Evidence
- 2benchesrecipe
- VRAM
- 24GB
- Best speed
- 35.1tokens/s
- Min VRAM
- 24GB
- Evidence
- 2benchesrecipe
- VRAM
- 64GB
- Best speed
- Min VRAM
- Evidence
- recipe
- VRAM
- 48GB
- Best speed
- Min VRAM
- Evidence
- recipe
- VRAM
- 48GB
- Best speed
- Min VRAM
- Evidence
- recipe
- VRAM
- 24GB
- Best speed
- Min VRAM
- Evidence
- recipe
✓ benchmarked·~ runs via recipe (not benchmarked)·— untested·✕doesn't fit