§01·model · /models
Gemma 4 26B MoE
multimodalactiveApache-2.0
26B Mixture-of-Experts LLM by Google — Gemma 4 A4B, roughly 4B active per token, instruction-tuned and multimodal. Four cards carry a recipe here from a 24 GB floor: Q4_K_M for chat on the 3090, 3090 Ti and 4090, and a Q8_0 quality tier on the 5090. A 12 GB RTX 3060 has a benchmark row but no recipe — a community-submitted run at 37.2 tok/s with a 10.92 GB peak. Apache-2.0.
§02·same family
3 other models · by nameOther models grouped with Gemma 4 26B MoE. Sizes and modalities can differ, and nothing here says whether one fits your GPU — open a model for its own compatibility table.
§03·GPUs that run this model
5 total✓ benchmarked·~ runs via recipe (not benchmarked)·— untested·✕doesn't fit