§01·model · /models
llava 7b
multimodalactiveLlama 2 Community License
7B open vision-language model — LLaVA 1.5, a Vicuna/Llama-2 LLM paired with a CLIP visual encoder for image understanding and visual Q&A. One of the models that opened this space, and still light enough to matter: the single recipe here runs it on an 8 GB RTX 3060 Ti, with 72 tokens/s recorded at 4-bit. Llama 2 Community License.
§02·GPUs that run this model
1 total| GPU | VRAM | Series | Best speed | Min VRAM | Works | Evidence | |
|---|---|---|---|---|---|---|---|
| RTX 3060 Ti | 8GB | 30 | 72tokens/s | 8GB | ✓ | 1benchrecipe | check ↗ |
✓ benchmarked·~ runs via recipe (not benchmarked)·— untested·✕doesn't fit