§01·model · /models
llama2 7b
llmactiveLlama 2 Community License
7B chat LLM by Meta — the Llama 2 baseline, kept here as a reference point rather than a recommendation. It is also the fastest of the eleven models measured on the same 8 GB RTX 3060 Ti in one shared 4-bit Ollama run, at 73.07 tokens/s, which is what a small model of an older generation buys you. Llama 2 Community License.
§02·same family
1 other modelOther models grouped with llama2 7b. Sizes and modalities can differ, and nothing here says whether one fits your GPU — open a model for its own compatibility table.
§03·GPUs that run this model
1 total| GPU | VRAM | Series | Best speed | Min VRAM | Works | Evidence | |
|---|---|---|---|---|---|---|---|
| RTX 3060 Ti | 8GB | 30 | 73.07tokens/s | 8GB | ✓ | 1benchrecipe | check ↗ |
✓ benchmarked·~ runs via recipe (not benchmarked)·— untested·✕doesn't fit