self-hosted/ai
§01·compatibility · /check

Qwen3 32B on RTX 3090 Ti

Yes — Qwen3 32B runs on the RTX 3090 Ti (24 GB). Fastest community-measured generation speed: 38 tokens/s.

✓ runsllmactive30 series24GB VRAM
step-by-step recipe for this pair

Qwen3-32B on RTX 3090 Ti: UD-Q4_K_XL GGUF via llama.cpp

llmintermediate22GB+
model
name
Qwen3 32B
slug
qwen3-32b
vertical
llm
status
active
open detail ↗
gpu
name
RTX 3090 Ti
slug
rtx-3090-ti
vram
24 GB
series
30
open detail ↗
§02·benchmarks
TaskQuantSpeedVRAMWorksConfidenceSourceVerified
llmQ4_K1238.8prefill tokens/s✓hardware-corner.net· web2026-05-15
llmQ4_K38tokens/s✓hardware-corner.net· web2026-05-15
§03·how it was measured
  • llmQ4_K1238.8 prefill tokens/s

    Prompt processing speed at 4k context

    recorded from hardware-corner.net ↗ · verified 2026-05-15

  • llmQ4_K38 tokens/s

    Generation speed at 4k context. Also explicitly mentioned 32.6 t/s at 16k context.

    recorded from hardware-corner.net ↗ · verified 2026-05-15

§04·more Qwen3 32B recipes
§05·common questions
Can you run Qwen3 32B on RTX 3090 Ti?

Yes — Qwen3 32B runs on the RTX 3090 Ti (24 GB). Fastest community-measured generation speed: 38 tokens/s.

Which quantizations have been tested for Qwen3 32B on RTX 3090 Ti?

Q4_K — measured in community benchmarks.

How fast is Qwen3 32B on RTX 3090 Ti?

Up to 38 tokens/s (llm), the fastest community-measured generation speed.

Are there step-by-step instructions for Qwen3 32B on RTX 3090 Ti?

Yes — a step-by-step recipe documents Qwen3 32B on the RTX 3090 Ti, linked at the top of this page.