self-hosted/ai
§01·model · /models

llama2 13b

llmactiveLlama 2 Community License

13B chat LLM by Meta — the larger Llama 2 chat model, a legacy generation that is still widely fine-tuned. One recipe covers it, on an 8 GB RTX 3060 Ti, and it is the slowest of the eleven models measured on that card: 9.25 tokens/s at 4 bits, against 73.07 for Llama 2 7B on the very same hardware. Llama 2 Community License.

§02·same family
1 other model

Other models grouped with llama2 13b. Sizes and modalities can differ, and nothing here says whether one fits your GPU — open a model for its own compatibility table.

§03·GPUs that run this model
1 total

benchmarked·~ runs via recipe (not benchmarked)· untested·doesn't fit