self-hosted/ai
§01·model · /models

llama2 7b

llmactiveLlama 2 Community License

7B chat LLM by Meta — the Llama 2 baseline, kept here as a reference point rather than a recommendation. It is also the fastest of the eleven models measured on the same 8 GB RTX 3060 Ti in one shared 4-bit Ollama run, at 73.07 tokens/s, which is what a small model of an older generation buys you. Llama 2 Community License.

§02·same family
1 other model

Other models grouped with llama2 7b. Sizes and modalities can differ, and nothing here says whether one fits your GPU — open a model for its own compatibility table.

§03·GPUs that run this model
1 total

benchmarked·~ runs via recipe (not benchmarked)· untested·doesn't fit