Qwen3.8 27B
27B dense vision-language model by Alibaba, Apache-2.0, reading images and video alongside text. Its layer stack is hybrid — 48 of the 64 layers run Gated DeltaNet linear attention and only 16 keep a KV cache — so context costs roughly a quarter of what the dimensions suggest: 32K works out at about 2.1 GB where an all-attention 27B of the same shape would spend 8.6 GB, and a 128K window costs about what that model would spend on 32K. The full 262,144-token context still wants around 17 GB of KV on top of the weights. Ollama serves it as qwen3.8:27b at 18 GB with vision, tools and thinking enabled; llama.cpp runs ggml-org's own GGUF alongside a 0.93 GB mmproj projector for the image path. The 4-bit footprint depends on who built it — Q4_K_M spans 16.81 GB (lmstudio-community) to 18.97 GB (ggml-org).
| GPU | VRAM | Series | Best speed | Min VRAM | Works | Evidence | |
|---|---|---|---|---|---|---|---|
| RTX 5060 Ti | 16GB | 50 | 28.49tokens/s | ✓ | 2benchesrecipe | check ↗ | |
| Apple M2 Max | 64GB | apple | ~ | recipe | check ↗ | ||
| Apple M3 Max | 48GB | apple | ~ | recipe | check ↗ | ||
| Apple M4 Max | 48GB | apple | ~ | recipe | check ↗ | ||
| RTX 3090 | 24GB | 30 | ~ | recipe | check ↗ | ||
| RTX 3090 Ti | 24GB | 30 | ~ | recipe | check ↗ | ||
| RTX 4060 Ti 16GB | 16GB | 40 | ~ | recipe | check ↗ | ||
| RTX 4070 Ti Super | 16GB | 40 | ~ | recipe | check ↗ | ||
| RTX 4080 | 16GB | 40 | ~ | recipe | check ↗ | ||
| RTX 4080 Super | 16GB | 40 | ~ | recipe | check ↗ | ||
| RTX 4090 | 24GB | 40 | ~ | recipe | check ↗ | ||
| RTX 5070 Ti | 16GB | 50 | ~ | recipe | check ↗ | ||
| RTX 5080 | 16GB | 50 | ~ | recipe | check ↗ | ||
| RTX 5090 | 32GB | 50 | ~ | recipe | check ↗ | ||
| RX 7800 XT | 16GB | amd | ~ | recipe | check ↗ | ||
| RX 7900 XTX | 24GB | amd | ~ | recipe | check ↗ |
- VRAM
- 16GB
- Best speed
- 28.49tokens/s
- Min VRAM
- Evidence
- 2benches
- VRAM
- 64GB
- Best speed
- Min VRAM
- VRAM
- 48GB
- Best speed
- Min VRAM
- VRAM
- 48GB
- Best speed
- Min VRAM
- VRAM
- 24GB
- Best speed
- Min VRAM
- VRAM
- 24GB
- Best speed
- Min VRAM
- VRAM
- 16GB
- Best speed
- Min VRAM
- VRAM
- 16GB
- Best speed
- Min VRAM
- VRAM
- 16GB
- Best speed
- Min VRAM
- VRAM
- 16GB
- Best speed
- Min VRAM
- VRAM
- 24GB
- Best speed
- Min VRAM
- VRAM
- 16GB
- Best speed
- Min VRAM
- VRAM
- 16GB
- Best speed
- Min VRAM
- VRAM
- 32GB
- Best speed
- Min VRAM
- VRAM
- 16GB
- Best speed
- Min VRAM
- VRAM
- 24GB
- Best speed
- Min VRAM
✓ benchmarked·~ runs via recipe (not benchmarked)·— untested·✕doesn't fit