§01·model · /models
Gemma4 31B
multimodalactiveApache-2.0
31B dense LLM by Google, from the Gemma 4 family. Two cards carry a recipe here — a 24 GB RTX 3090 and a 32 GB 5090 — and the figures recorded on them are what dense costs: 34.7 tokens/s and 61.1 tokens/s at Q4_K. The MoE sibling in this same catalogue, Qwen3.5 35B-A3B, is larger on paper and runs roughly two and a half to three times faster on those exact two cards. Apache-2.0.
§02·same family
3 other models · by nameOther models grouped with Gemma4 31B. Sizes and modalities can differ, and nothing here says whether one fits your GPU — open a model for its own compatibility table.
§03·GPUs that run this model
2 total✓ benchmarked·~ runs via recipe (not benchmarked)·— untested·✕doesn't fit