§01·compatibility · /check
Qwen3 32B on RTX 5090
Yes — Qwen3 32B runs on the RTX 5090 (32 GB). Fastest community-measured generation speed: 61.4 tokens/s.
✓ runsllmactive50 series32GB VRAM
step-by-step recipe for this pair
Qwen3-32B on RTX 5090: Q6_K_XL GGUF via llama.cpp (with AWQ-INT4 + 128K context alternative)
llmintermediate29GB+
§02·benchmarks
| Task | Quant | Speed | VRAM | Works | Confidence | Source | Verified |
|---|---|---|---|---|---|---|---|
| llm | Q4_K | 2931.3prefill tokens/s | ✓ | hardware-corner.net· web | 2026-05-15 | ||
| llm | Q4_K | 61.4tokens/s | ✓ | hardware-corner.net· web | 2026-05-15 | ||
| llm | Q4_K_XL | 50.92tokens/s | ✓ | hardware-corner.net· web | 2026-05-15 |
§03·how it was measured
- llmQ4_K2931.3 prefill tokens/s
4k Context
recorded from hardware-corner.net ↗ · verified 2026-05-15
- llmQ4_K61.4 tokens/s
4k Context
recorded from hardware-corner.net ↗ · verified 2026-05-15
- llmQ4_K_XL50.92 tokens/s
16K context
recorded from hardware-corner.net ↗ · verified 2026-05-15
§04·more Qwen3 32B recipes
- llmintermediate19GB+
Qwen3-32B on Apple M3 Max: 32B local chat with MLX 4-bit in 48 GB unified memory
- llmintermediate19GB+
Qwen3-32B on Apple M4 Max: 32B local chat with MLX 4-bit in 48 GB unified memory
- llmintermediate19GB+
Qwen3-32B on Apple M2 Max: 32B local chat with MLX 4-bit in 64 GB unified memory
- llmintermediate22GB+
Qwen3-32B on RTX 3090 Ti: UD-Q4_K_XL GGUF via llama.cpp
§05·common questions
Can you run Qwen3 32B on RTX 5090?
Yes — Qwen3 32B runs on the RTX 5090 (32 GB). Fastest community-measured generation speed: 61.4 tokens/s.
Which quantizations have been tested for Qwen3 32B on RTX 5090?
Q4_K, Q4_K_XL — measured in community benchmarks.
How fast is Qwen3 32B on RTX 5090?
Up to 61.4 tokens/s (llm), the fastest community-measured generation speed.
Are there step-by-step instructions for Qwen3 32B on RTX 5090?
Yes — a step-by-step recipe documents Qwen3 32B on the RTX 5090, linked at the top of this page.