§01·compatibility · /check
gpt-oss 20B on RTX 3090 Ti
Yes — gpt-oss 20B runs on the RTX 3090 Ti (24 GB). Fastest community-measured generation speed: 160.3 tokens/s.
✓ runsllmactive30 series24GB VRAM
step-by-step recipe for this pair
gpt-oss 20B on RTX 3090 Ti: MXFP4 Chat at 160 tok/s via Ollama or vLLM
llmbeginner16GB+
§02·benchmarks
| Task | Quant | Speed | VRAM | Works | Confidence | Source | Verified |
|---|---|---|---|---|---|---|---|
| llm | MXFP4 | 4805.4prefill tokens/s | ✓ | hardware-corner.net· web | 2026-05-15 | ||
| llm | MXFP4 | 160.3tokens/s | ✓ | hardware-corner.net· web | 2026-05-15 |
§03·how it was measured
- llmMXFP44805.4 prefill tokens/s
Prompt processing speed at 4k context
recorded from hardware-corner.net ↗ · verified 2026-05-15
- llmMXFP4160.3 tokens/s
Generation speed at 4k context
recorded from hardware-corner.net ↗ · verified 2026-05-15
§04·more gpt-oss 20B recipes
- llmbeginner16GB+
gpt-oss 20B on RTX 5060 Ti: MXFP4 Chat at 92 tok/s via Ollama or vLLM
- llmbeginner16GB+
gpt-oss 20B on RTX 4060 Ti 16GB: MXFP4 chat at 63 tok/s via Ollama or vLLM
- llmintermediate13GB+
gpt-oss 20B on Apple M2 Pro: running 20B in 16 GB unified memory with the wired-limit raise
- llmbeginner13GB+
gpt-oss 20B on Apple M3 Max: native-MXFP4 chat in 48 GB unified memory with MLX
§05·common questions
Can you run gpt-oss 20B on RTX 3090 Ti?
Yes — gpt-oss 20B runs on the RTX 3090 Ti (24 GB). Fastest community-measured generation speed: 160.3 tokens/s.
Which quantizations have been tested for gpt-oss 20B on RTX 3090 Ti?
MXFP4 — measured in community benchmarks.
How fast is gpt-oss 20B on RTX 3090 Ti?
Up to 160.3 tokens/s (llm), the fastest community-measured generation speed.
Are there step-by-step instructions for gpt-oss 20B on RTX 3090 Ti?
Yes — a step-by-step recipe documents gpt-oss 20B on the RTX 3090 Ti, linked at the top of this page.