§01·compatibility · /check
Qwen3 14B on RTX 3080 Ti
Yes — Qwen3 14B runs on the RTX 3080 Ti (12 GB). Fastest community-measured generation speed: 69.9 tokens/s.
✓ runsllmactive30 series12GB VRAM
step-by-step recipe for this pair
Qwen3-14B on RTX 3080 Ti: Q4_K_M GGUF via Ollama or llama.cpp
llmbeginner12GB+
§02·benchmarks
| Task | Quant | Speed | VRAM | Works | Confidence | Source | Verified |
|---|---|---|---|---|---|---|---|
| llm | Q4_K | 2600.1prefill tokens/s | ✓ | hardware-corner.net· web | 2026-05-15 | ||
| llm | Q4_K | 69.9tokens/s | ✓ | hardware-corner.net· web | 2026-05-15 |
§03·how it was measured
- llmQ4_K2600.1 prefill tokens/s
Prompt processing speed at 4k context
recorded from hardware-corner.net ↗ · verified 2026-05-15
- llmQ4_K69.9 tokens/s
Generation speed at 4k context
recorded from hardware-corner.net ↗ · verified 2026-05-15
§04·more Qwen3 14B recipes
- llmbeginner12GB+
Qwen3-14B on RTX 4060 Ti 16GB: Q4_K_M GGUF via Ollama or llama.cpp
- llmintermediate9GB+
Qwen3-14B on Apple M2 Pro: the strongest LLM that fits a 16 GB unified-memory Mac, via MLX 4-bit
- llmbeginner9GB+
Qwen3-14B on RX 7800 XT: ROCm via Ollama or llama.cpp-HIP
- llmbeginner12GB+
Qwen3-14B on RTX 5060 Ti: Q4_K_M GGUF via Ollama or llama.cpp
§05·common questions
Can you run Qwen3 14B on RTX 3080 Ti?
Yes — Qwen3 14B runs on the RTX 3080 Ti (12 GB). Fastest community-measured generation speed: 69.9 tokens/s.
Which quantizations have been tested for Qwen3 14B on RTX 3080 Ti?
Q4_K — measured in community benchmarks.
How fast is Qwen3 14B on RTX 3080 Ti?
Up to 69.9 tokens/s (llm), the fastest community-measured generation speed.
Are there step-by-step instructions for Qwen3 14B on RTX 3080 Ti?
Yes — a step-by-step recipe documents Qwen3 14B on the RTX 3080 Ti, linked at the top of this page.