§01·model · /models
Qwen2.5 7B
llmactiveApache-2.0
7B instruction-tuned LLM by Alibaba, the Qwen2.5 generation — the step between Qwen2 and Qwen3, both of which are also catalogued here. One recipe covers it, on an 8 GB RTX 3060 Ti, where the shared 4-bit Ollama run records 58.13 tokens/s: slower than the Qwen2 it replaced on that same card, which is the usual trade when a generation adds capability rather than speed. Apache-2.0.
§02·same family
1 other modelOther models grouped with Qwen2.5 7B. Sizes and modalities can differ, and nothing here says whether one fits your GPU — open a model for its own compatibility table.
§03·GPUs that run this model
1 total| GPU | VRAM | Series | Best speed | Min VRAM | Works | Evidence | |
|---|---|---|---|---|---|---|---|
| RTX 3060 Ti | 8GB | 30 | 58.13tokens/s | 8GB | ✓ | 1benchrecipe | check ↗ |
✓ benchmarked·~ runs via recipe (not benchmarked)·— untested·✕doesn't fit