§01·compatibility · /check
Llama 3.3 70B on RTX 4090
Yes — Llama 3.3 70B runs on the RTX 4090 (24 GB). Fastest community-measured result: 17.79 tokens/s.
✓ runsllmactive40 series24GB VRAM
step-by-step recipe for this pair
Llama 3.3 70B on RTX 4090: 70B-Class Chat on One 24 GB Card (Q4 Offload or Fully-On-GPU IQ2)
llmadvanced24GB+
§02·benchmarks
| Task | Quant | Speed | VRAM | Works | Confidence | Source | Verified |
|---|---|---|---|---|---|---|---|
| llm | Q4_K_XL | 17.79tokens/s | ✓ | hardware-corner.net· web | 2026-06-26 |
§03·how it was measured
- llmQ4_K_XL17.79 tokens/s
16K context
recorded from hardware-corner.net ↗ · verified 2026-06-26
§04·more Llama 3.3 70B recipes
§05·common questions
Can you run Llama 3.3 70B on RTX 4090?
Yes — Llama 3.3 70B runs on the RTX 4090 (24 GB). Fastest community-measured result: 17.79 tokens/s.
Which quantizations have been tested for Llama 3.3 70B on RTX 4090?
Q4_K_XL — measured in community benchmarks.
How fast is Llama 3.3 70B on RTX 4090?
Up to 17.79 tokens/s (llm), the fastest community-measured result.
Are there step-by-step instructions for Llama 3.3 70B on RTX 4090?
Yes — a step-by-step recipe documents Llama 3.3 70B on the RTX 4090, linked at the top of this page.