§01·model · /models
gpt-oss 20B
llmactiveApache-2.0
20B open-weight LLM by OpenAI (gpt-oss) for local reasoning and agentic use, Apache-2.0. Twenty-three cards here carry a published recipe, down to a 12 GB floor. Measured at 298.2 tok/s generation on an RTX 5090, and 6,364 prefill tok/s on a 16 GB RTX 4080 SUPER: for agent loops the prefill number is the one that decides whether a long tool history feels instant or sluggish, since the prompt is re-read far more often than the answer is written.
§02·same family
1 other modelOther models grouped with gpt-oss 20B. Sizes and modalities can differ, and nothing here says whether one fits your GPU — open a model for its own compatibility table.
§03·GPUs that run this model
23 total| GPU | VRAM | Series | Best speed | Min VRAM | Works | Evidence | |
|---|---|---|---|---|---|---|---|
| RTX 4080 Super | 16GB | 40 | 6364prefill tokens/s | ✓ | 1benchrecipe | check ↗ | |
| RTX 5090 | 32GB | 50 | 298.2tokens/s | ✓ | 3benchesrecipe | check ↗ | |
| RTX 5080 | 16GB | 50 | 172.4tokens/s | ✓ | 2benchesrecipe | check ↗ | |
| RTX 3090 Ti | 24GB | 30 | 160.3tokens/s | ✓ | 2benchesrecipe | check ↗ | |
| RTX 5070 Ti | 16GB | 50 | 156tokens/s | ✓ | 2benchesrecipe | check ↗ | |
| RTX 3090 | 24GB | 30 | 147.5tokens/s | ✓ | 2benchesrecipe | check ↗ | |
| RTX 4080 | 16GB | 40 | 136.5tokens/s | ✓ | 2benchesrecipe | check ↗ | |
| RTX 4070 Ti Super | 16GB | 40 | 128.9tokens/s | ✓ | 2benchesrecipe | check ↗ | |
| RTX 5060 Ti | 16GB | 50 | 92.1tokens/s | ✓ | 2benchesrecipe | check ↗ | |
| RTX 3060 | 12GB | 30 | 64tokens/s | ✓ | 1benchrecipe | check ↗ | |
| RTX 4060 Ti 16GB | 16GB | 40 | 63.2tokens/s | ✓ | 2benchesrecipe | check ↗ | |
| Apple M2 Max | 64GB | apple | ~ | recipe | check ↗ | ||
| Apple M2 Pro | 16GB | apple | ~ | recipe | check ↗ | ||
| Apple M3 Max | 48GB | apple | ~ | recipe | check ↗ | ||
| Apple M4 Max | 48GB | apple | ~ | recipe | check ↗ | ||
| RTX 3080 Ti | 12GB | 30 | ~ | recipe | check ↗ | ||
| RTX 4070 | 12GB | 40 | ~ | recipe | check ↗ | ||
| RTX 4070 Super | 12GB | 40 | ~ | recipe | check ↗ | ||
| RTX 4070 Ti | 12GB | 40 | ~ | recipe | check ↗ | ||
| RTX 4090 | 24GB | 40 | ~ | recipe | check ↗ | ||
| RTX 5070 | 12GB | 50 | ~ | recipe | check ↗ | ||
| RX 7800 XT | 16GB | amd | ~ | recipe | check ↗ | ||
| RX 7900 XTX | 24GB | amd | ~ | recipe | check ↗ |
- VRAM
- 16GB
- Best speed
- 6364prefill tokens/s
- Min VRAM
- Evidence
- 1bench
- VRAM
- 32GB
- Best speed
- 298.2tokens/s
- Min VRAM
- Evidence
- 3benches
- VRAM
- 16GB
- Best speed
- 172.4tokens/s
- Min VRAM
- Evidence
- 2benches
- VRAM
- 24GB
- Best speed
- 160.3tokens/s
- Min VRAM
- Evidence
- 2benches
- VRAM
- 16GB
- Best speed
- 156tokens/s
- Min VRAM
- Evidence
- 2benches
- VRAM
- 24GB
- Best speed
- 147.5tokens/s
- Min VRAM
- Evidence
- 2benches
- VRAM
- 16GB
- Best speed
- 136.5tokens/s
- Min VRAM
- Evidence
- 2benches
- VRAM
- 16GB
- Best speed
- 128.9tokens/s
- Min VRAM
- Evidence
- 2benches
- VRAM
- 16GB
- Best speed
- 92.1tokens/s
- Min VRAM
- Evidence
- 2benches
- VRAM
- 12GB
- Best speed
- 64tokens/s
- Min VRAM
- Evidence
- 1bench
- VRAM
- 16GB
- Best speed
- 63.2tokens/s
- Min VRAM
- Evidence
- 2benches
- VRAM
- 64GB
- Best speed
- Min VRAM
- VRAM
- 16GB
- Best speed
- Min VRAM
- VRAM
- 48GB
- Best speed
- Min VRAM
- VRAM
- 48GB
- Best speed
- Min VRAM
- VRAM
- 12GB
- Best speed
- Min VRAM
- VRAM
- 12GB
- Best speed
- Min VRAM
- VRAM
- 12GB
- Best speed
- Min VRAM
- VRAM
- 12GB
- Best speed
- Min VRAM
- VRAM
- 24GB
- Best speed
- Min VRAM
- VRAM
- 12GB
- Best speed
- Min VRAM
- VRAM
- 16GB
- Best speed
- Min VRAM
- VRAM
- 24GB
- Best speed
- Min VRAM
✓ benchmarked·~ runs via recipe (not benchmarked)·— untested·✕doesn't fit