self-hosted/ai
§01·model · /models

Llama 3.1 8B

llmactiveLlama 3.1 Community License

8B instruction-tuned LLM by Meta (Llama 3.1) with 128K context — the workhorse small Llama, under the Llama 3.1 Community License. Twenty-three cards here have a recipe, and the floor is 8 GB, which is what keeps it relevant next to newer 8B models: measured 57.3 tok/s on an 8 GB RTX 3060 Ti and 60.0 tok/s on a 12 GB RTX 4070 Ti. Note the licence is Meta's own, not Apache — it carries conditions that Apache-2.0 models in this catalogue do not.

§02·same family
1 other model

Other models grouped with Llama 3.1 8B. Sizes and modalities can differ, and nothing here says whether one fits your GPU — open a model for its own compatibility table.

§03·GPUs that run this model
23 total

benchmarked·~ runs via recipe (not benchmarked)· untested·doesn't fit