self-hosted/ai
§01·model · /models

Llama 3.3 70B

llmactiveLlama 3.3 Community License

70B instruction-tuned LLM by Meta — Llama 3.3, its strongest 70B, with quality approaching the 405B tier. Four machines carry a recipe here and three are Apple laptops with 48 to 64 GB of unified memory; on NVIDIA it comes down to a single 24 GB RTX 4090, recorded at 17.79 tokens/s. At this size unified memory is the practical route. Llama 3.3 Community License.

§02·GPUs that run this model
4 total

benchmarked·~ runs via recipe (not benchmarked)· untested·doesn't fit