§01·model · /models
llama2 13b
llmactiveLlama 2 Community License
13B chat LLM by Meta — the larger Llama 2 chat model, a legacy generation that is still widely fine-tuned. One recipe covers it, on an 8 GB RTX 3060 Ti, and it is the slowest of the eleven models measured on that card: 9.25 tokens/s at 4 bits, against 73.07 for Llama 2 7B on the very same hardware. Llama 2 Community License.
§02·same family
1 other modelOther models grouped with llama2 13b. Sizes and modalities can differ, and nothing here says whether one fits your GPU — open a model for its own compatibility table.
§03·GPUs that run this model
1 total| GPU | VRAM | Series | Best speed | Min VRAM | Works | Evidence | |
|---|---|---|---|---|---|---|---|
| RTX 3060 Ti | 8GB | 30 | 9.25tokens/s | 8GB | ✓ | 1benchrecipe | check ↗ |
✓ benchmarked·~ runs via recipe (not benchmarked)·— untested·✕doesn't fit