self-hosted/ai
§01·compatibility · /check

gpt-oss 20B on RTX 4060 Ti 16GB

Yes — gpt-oss 20B runs on the RTX 4060 Ti 16GB (16 GB). Fastest community-measured generation speed: 63.2 tokens/s.

✓ runsllmactive40 series16GB VRAM
step-by-step recipe for this pair

gpt-oss 20B on RTX 4060 Ti 16GB: MXFP4 chat at 63 tok/s via Ollama or vLLM

llmbeginner16GB+
model
name
gpt-oss 20B
slug
gpt-oss-20b
vertical
llm
status
active
open detail ↗
gpu
name
RTX 4060 Ti 16GB
slug
rtx-4060-ti-16gb
vram
16 GB
series
40
open detail ↗
§02·benchmarks
TaskQuantSpeedVRAMWorksConfidenceSourceVerified
llmMXFP43274.2prefill tokens/s✓hardware-corner.net· web2026-05-15
llmMXFP463.2tokens/s✓hardware-corner.net· web2026-05-15
§03·how it was measured
  • llmMXFP43274.2 prefill tokens/s

    4k context, prompt processing speed

    recorded from hardware-corner.net ↗ · verified 2026-05-15

  • llmMXFP463.2 tokens/s

    4k context, generation speed

    recorded from hardware-corner.net ↗ · verified 2026-05-15

§04·more gpt-oss 20B recipes
§05·common questions
Can you run gpt-oss 20B on RTX 4060 Ti 16GB?

Yes — gpt-oss 20B runs on the RTX 4060 Ti 16GB (16 GB). Fastest community-measured generation speed: 63.2 tokens/s.

Which quantizations have been tested for gpt-oss 20B on RTX 4060 Ti 16GB?

MXFP4 — measured in community benchmarks.

How fast is gpt-oss 20B on RTX 4060 Ti 16GB?

Up to 63.2 tokens/s (llm), the fastest community-measured generation speed.

Are there step-by-step instructions for gpt-oss 20B on RTX 4060 Ti 16GB?

Yes — a step-by-step recipe documents gpt-oss 20B on the RTX 4060 Ti 16GB, linked at the top of this page.