self-hosted/ai
§01·compatibility · /check

gpt-oss 20B on RTX 3090 Ti

Yes — gpt-oss 20B runs on the RTX 3090 Ti (24 GB). Fastest community-measured result: 4805.4 prefill tokens/s, with 24 GB peak VRAM.

runsllmactive30 series24GB VRAM
step-by-step recipe for this pair

gpt-oss 20B on RTX 3090 Ti: MXFP4 Chat at 160 tok/s via Ollama or vLLM

llmbeginner16GB+
model
name
gpt-oss 20B
slug
gpt-oss-20b
vertical
llm
status
active
open detail ↗
gpu
name
RTX 3090 Ti
slug
rtx-3090-ti
vram
24 GB
series
30
open detail ↗
§02·benchmarks
TaskQuantSpeedVRAMWorksConfidenceSourceVerified
llmMXFP44805.4prefill tokens/s24GBhardware-corner.net· web2026-05-15
llmMXFP4160.3tokens/s24GBhardware-corner.net· web2026-05-15
§03·more gpt-oss 20B recipes
§04·common questions
Can you run gpt-oss 20B on RTX 3090 Ti?

Yes — gpt-oss 20B runs on the RTX 3090 Ti (24 GB). Fastest community-measured result: 4805.4 prefill tokens/s, with 24 GB peak VRAM.

How much VRAM does gpt-oss 20B need on RTX 3090 Ti?

Measured peak VRAM is 24 GB.

Which quantizations have been tested for gpt-oss 20B on RTX 3090 Ti?

MXFP4 — measured in community benchmarks.

How fast is gpt-oss 20B on RTX 3090 Ti?

Up to 4805.4 prefill tokens/s (llm), the fastest community-measured result.

Are there step-by-step instructions for gpt-oss 20B on RTX 3090 Ti?

Yes — a step-by-step recipe documents gpt-oss 20B on the RTX 3090 Ti, linked at the top of this page.