self-hosted/ai
§01·model · /models

Bonsai 27B

multimodalactiveApache-2.0

PrismML's 1-bit build of Qwen3.6-27B — a 27B multimodal (text + image) reasoning model compressed to a 3.54 GB GGUF, which puts a 27B-class model within reach of any 8 GB card. Weights are binary with an FP16 scale per 128-weight group (~1.125 bits/weight); the vision tower ships separately as a 0.59 GB mmproj file.

Runs on upstream llama.cpp — the Q1_0 type is merged, so the vendor's fork is not required — but not on Ollama, whose Q1_0 support was never merged. 262K context, thinking-mode reasoning, Apache-2.0.

PrismML reports ~89.5% of the FP16 baseline; independent testing shows the loss is uneven, with knowledge and coding holding up while math and multilingual degrade sharply.

Download· 2 variants
§02·GPUs that run this model
6 total

benchmarked·~ runs via recipe (not benchmarked)· untested·doesn't fit