§01·model · /models
Llama 3.3 70B
llmactiveLlama 3.3 Community License
70B instruction-tuned LLM by Meta — Llama 3.3, its strongest 70B, with quality approaching the 405B tier. Four machines carry a recipe here and three are Apple laptops with 48 to 64 GB of unified memory; on NVIDIA it comes down to a single 24 GB RTX 4090, recorded at 17.79 tokens/s. At this size unified memory is the practical route. Llama 3.3 Community License.
§02·GPUs that run this model
4 total| GPU | VRAM | Series | Best speed | Min VRAM | Works | Evidence | |
|---|---|---|---|---|---|---|---|
| RTX 4090 | 24GB | 40 | 17.79tokens/s | ✓ | 1benchrecipe | check ↗ | |
| Apple M2 Max | 64GB | apple | ~ | recipe | check ↗ | ||
| Apple M3 Max | 48GB | apple | ~ | recipe | check ↗ | ||
| Apple M4 Max | 48GB | apple | ~ | recipe | check ↗ |
✓ benchmarked·~ runs via recipe (not benchmarked)·— untested·✕doesn't fit