self-hosted/ai
§01·model · /models

Qwen3.5 27B

multimodalactiveApache-2.0

27B vision-language model by Alibaba (Qwen3.5), Apache-2.0 — it reads images, not only text. Two cards here carry a published recipe: 33.5 tok/s on an RTX 3090 and 58.8 tok/s on an RTX 5090. Both are 24 GB or larger, but the weights do not demand that much. Only 16 of the 64 layers keep a KV cache and the other 48 run linear attention, so 32K of context costs about 2.1 GB rather than the 8.6 GB the same model would need if every layer kept one, and Unsloth's IQ4_XS build is 14.98 GB. A 16 GB board is untested here rather than ruled out — it is the 0.93 GB vision projector on top that pushes that configuration past the limit.

§02·same family
1 other model

Other models grouped with Qwen3.5 27B. Sizes and modalities can differ, and nothing here says whether one fits your GPU — open a model for its own compatibility table.

§03·GPUs that run this model
2 total

benchmarked·~ runs via recipe (not benchmarked)· untested·doesn't fit