§01·compatibility · /check
Llama 3.3 70B on RTX 4090
Yes — Llama 3.3 70B runs on the RTX 4090 (24 GB). Fastest community-measured result: 17.79 tokens/s.
✓ runsllmactive40 series24GB VRAM
§02·benchmarks
| Task | Quant | Speed | VRAM | Works | Confidence | Source | Verified |
|---|---|---|---|---|---|---|---|
| llm | Q4_K_XL | 17.79tokens/s | ✓ | hardware-corner.net· web | 2026-06-26 |
§03·common questions
Can you run Llama 3.3 70B on RTX 4090?
Yes — Llama 3.3 70B runs on the RTX 4090 (24 GB). Fastest community-measured result: 17.79 tokens/s.
Which quantizations have been tested for Llama 3.3 70B on RTX 4090?
Q4_K_XL — measured in community benchmarks.
How fast is Llama 3.3 70B on RTX 4090?
Up to 17.79 tokens/s (llm), the fastest community-measured result.
Are there step-by-step instructions for Llama 3.3 70B on RTX 4090?
Yes — a step-by-step recipe documents Llama 3.3 70B on the RTX 4090; see the recipes listed below.
§04·related recipes
- llmadvanced24GB+
Llama 3.3 70B on RTX 4090: 70B-Class Chat on One 24 GB Card (Q4 Offload or Fully-On-GPU IQ2)
- llmadvanced40GB+
Llama 3.3 70B on Apple M3 Max: 70B-class chat in 48 GB unified memory with MLX
- llmadvanced40GB+
Llama 3.3 70B on Apple M4 Max: 70B-class chat in 48 GB unified memory with MLX
- llmintermediate40GB+
Llama 3.3 70B on Apple M2 Max: 70B-class chat in 64 GB unified memory with MLX