self-hosted/ai
§01·model · /models

Gemma 4 26B MoE

multimodalactiveApache-2.0

26B Mixture-of-Experts LLM by Google — Gemma 4 A4B, roughly 4B active per token, instruction-tuned and multimodal. Four cards carry a recipe here from a 24 GB floor: Q4_K_M for chat on the 3090, 3090 Ti and 4090, and a Q8_0 quality tier on the 5090. A 12 GB RTX 3060 has a benchmark row but no recipe — a community-submitted run at 37.2 tok/s with a 10.92 GB peak. Apache-2.0.

§02·same family
3 other models · by name

Other models grouped with Gemma 4 26B MoE. Sizes and modalities can differ, and nothing here says whether one fits your GPU — open a model for its own compatibility table.

§03·GPUs that run this model
5 total

benchmarked·~ runs via recipe (not benchmarked)· untested·doesn't fit