§01·model · /models
qwen2 7b
llmactiveApache-2.0
7B instruction-tuned LLM by Alibaba, the Qwen2 generation — since superseded by Qwen2.5 and Qwen3, both also in this catalogue. It is here mainly as a comparison point: one recipe, on an 8 GB RTX 3060 Ti, where a shared 4-bit Ollama run records 63.73 tokens/s, third of the nine models measured on that same card. Apache-2.0.
Download· 4 variants
§02·GPUs that run this model
1 total| GPU | VRAM | Series | Best speed | Min VRAM | Works | Evidence | |
|---|---|---|---|---|---|---|---|
| RTX 3060 Ti | 8GB | 30 | 63.73tokens/s | 8GB | ✓ | 1benchrecipe | check ↗ |
✓ benchmarked·~ runs via recipe (not benchmarked)·— untested·✕doesn't fit