self-hosted/ai
§01·compatibility · /check

gpt-oss 20B on RTX 3060

Yes — gpt-oss 20B runs on the RTX 3060 (12 GB). Fastest community-measured result: 64 tokens/s.

runsllmactive30 series12GB VRAM
step-by-step recipe for this pair

gpt-oss 20B on RTX 3060: MXFP4 Chat at 64 tok/s in 12 GB via llama.cpp Expert Offload

llmintermediate12GB+
model
name
gpt-oss 20B
slug
gpt-oss-20b
vertical
llm
status
active
open detail ↗
gpu
name
RTX 3060
slug
rtx-3060
vram
12 GB
series
30
open detail ↗
§02·benchmarks
TaskQuantSpeedVRAMWorksConfidenceSourceVerified
llmMXFP464tokens/sgithub.com· web2026-06-13
§03·how it was measured
  • llmMXFP464 tokens/s

    Token generation @ 16K with --n-cpu-moe 2 (official llama.cpp guide #15396, first-party RTX 3060 12GB / cc 8.6 / sm_86, QuantiusBenignus; ~600 MB VRAM spare)

    recorded from github.com · verified 2026-06-13

§04·more gpt-oss 20B recipes
§05·common questions
Can you run gpt-oss 20B on RTX 3060?

Yes — gpt-oss 20B runs on the RTX 3060 (12 GB). Fastest community-measured result: 64 tokens/s.

Which quantizations have been tested for gpt-oss 20B on RTX 3060?

MXFP4 — measured in community benchmarks.

How fast is gpt-oss 20B on RTX 3060?

Up to 64 tokens/s (llm), the fastest community-measured result.

Are there step-by-step instructions for gpt-oss 20B on RTX 3060?

Yes — a step-by-step recipe documents gpt-oss 20B on the RTX 3060, linked at the top of this page.