§01·compatibility · /check
Qwen3 32B on RTX 3090 Ti
Yes — Qwen3 32B runs on the RTX 3090 Ti (24 GB). Fastest community-measured result: 1238.8 prefill tokens/s, with 24 GB peak VRAM.
✓ runsllmactive30 series24GB VRAM
step-by-step recipe for this pair
Qwen3-32B on RTX 3090 Ti: UD-Q4_K_XL GGUF via llama.cpp
llmintermediate22GB+
§02·benchmarks
| Task | Quant | Speed | VRAM | Works | Confidence | Source | Verified |
|---|---|---|---|---|---|---|---|
| llm | Q4_K | 1238.8prefill tokens/s | 24GB | ✓ | hardware-corner.net· web | 2026-05-15 | |
| llm | Q4_K | 38tokens/s | 24GB | ✓ | hardware-corner.net· web | 2026-05-15 |
§03·more Qwen3 32B recipes
- llmintermediate19GB+
Qwen3-32B on Apple M3 Max: 32B local chat with MLX 4-bit in 48 GB unified memory
- llmintermediate19GB+
Qwen3-32B on Apple M4 Max: 32B local chat with MLX 4-bit in 48 GB unified memory
- llmintermediate19GB+
Qwen3-32B on Apple M2 Max: 32B local chat with MLX 4-bit in 64 GB unified memory
- llmintermediate29GB+
Qwen3-32B on RTX 5090: Q6_K_XL GGUF via llama.cpp (with AWQ-INT4 + 128K context alternative)
§04·common questions
Can you run Qwen3 32B on RTX 3090 Ti?
Yes — Qwen3 32B runs on the RTX 3090 Ti (24 GB). Fastest community-measured result: 1238.8 prefill tokens/s, with 24 GB peak VRAM.
How much VRAM does Qwen3 32B need on RTX 3090 Ti?
Measured peak VRAM is 24 GB.
Which quantizations have been tested for Qwen3 32B on RTX 3090 Ti?
Q4_K — measured in community benchmarks.
How fast is Qwen3 32B on RTX 3090 Ti?
Up to 1238.8 prefill tokens/s (llm), the fastest community-measured result.
Are there step-by-step instructions for Qwen3 32B on RTX 3090 Ti?
Yes — a step-by-step recipe documents Qwen3 32B on the RTX 3090 Ti, linked at the top of this page.