self-hosted/ai
§01·model · /models

Gemma4 31B

multimodalactiveApache-2.0

31B dense LLM by Google, from the Gemma 4 family. Two cards carry a recipe here — a 24 GB RTX 3090 and a 32 GB 5090 — and the figures recorded on them are what dense costs: 34.7 tokens/s and 61.1 tokens/s at Q4_K. The MoE sibling in this same catalogue, Qwen3.5 35B-A3B, is larger on paper and runs roughly two and a half to three times faster on those exact two cards. Apache-2.0.

§02·same family
3 other models · by name

Other models grouped with Gemma4 31B. Sizes and modalities can differ, and nothing here says whether one fits your GPU — open a model for its own compatibility table.

§03·GPUs that run this model
2 total

benchmarked·~ runs via recipe (not benchmarked)· untested·doesn't fit