§01·compatibility · /check
Qwen3 32B on RTX 3090
Yes — Qwen3 32B runs on the RTX 3090 (24 GB). Fastest community-measured result: 1087.9 prefill tokens/s, with 24 GB peak VRAM.
✓ runsllmactive30 series24GB VRAM
step-by-step recipe for this pair
Qwen3-32B on RTX 3090: UD-Q4_K_XL GGUF via llama.cpp
llmintermediate22GB+
§02·benchmarks
| Task | Quant | Speed | VRAM | Works | Confidence | Source | Verified |
|---|---|---|---|---|---|---|---|
| llm | Q4_K | 1087.9prefill tokens/s | 24GB | ✓ | hardware-corner.net· web | 2026-05-15 | |
| llm | Q4_K | 35.1tokens/s | 24GB | ✓ | hardware-corner.net· web | 2026-05-15 |
§03·more Qwen3 32B recipes
- llmintermediate19GB+
Qwen3-32B on Apple M3 Max: 32B local chat with MLX 4-bit in 48 GB unified memory
- llmintermediate19GB+
Qwen3-32B on Apple M4 Max: 32B local chat with MLX 4-bit in 48 GB unified memory
- llmintermediate19GB+
Qwen3-32B on Apple M2 Max: 32B local chat with MLX 4-bit in 64 GB unified memory
- llmintermediate22GB+
Qwen3-32B on RTX 3090 Ti: UD-Q4_K_XL GGUF via llama.cpp
§04·common questions
Can you run Qwen3 32B on RTX 3090?
Yes — Qwen3 32B runs on the RTX 3090 (24 GB). Fastest community-measured result: 1087.9 prefill tokens/s, with 24 GB peak VRAM.
How much VRAM does Qwen3 32B need on RTX 3090?
Measured peak VRAM is 24 GB.
Which quantizations have been tested for Qwen3 32B on RTX 3090?
Q4_K — measured in community benchmarks.
How fast is Qwen3 32B on RTX 3090?
Up to 1087.9 prefill tokens/s (llm), the fastest community-measured result.
Are there step-by-step instructions for Qwen3 32B on RTX 3090?
Yes — a step-by-step recipe documents Qwen3 32B on the RTX 3090, linked at the top of this page.