self-hosted/ai
§01·compatibility · /check

Gemma4 31B on RTX 5090

Yes — Gemma4 31B runs on the RTX 5090 (32 GB). Fastest community-measured generation speed: 61.1 tokens/s.

✓ runsmultimodalactive50 series32GB VRAM
step-by-step recipe for this pair

Gemma 4 31B on RTX 5090: dense 31B local chat at 61 tok/s, with Q5/Q6 quality headroom in 32 GB

llmintermediate20GB+
model
name
Gemma4 31B
slug
gemma4-31b
vertical
multimodal
status
active
open detail ↗
gpu
name
RTX 5090
slug
rtx-5090
vram
32 GB
series
50
open detail ↗
§02·benchmarks
TaskQuantSpeedVRAMWorksConfidenceSourceVerified
llmQ4_K3395prefill tokens/s✓hardware-corner.net· web2026-05-15
llmQ4_K61.1tokens/s✓hardware-corner.net· web2026-05-15
§03·how it was measured
§04·more Gemma4 31B recipes
§05·common questions
Can you run Gemma4 31B on RTX 5090?

Yes — Gemma4 31B runs on the RTX 5090 (32 GB). Fastest community-measured generation speed: 61.1 tokens/s.

Which quantizations have been tested for Gemma4 31B on RTX 5090?

Q4_K — measured in community benchmarks.

How fast is Gemma4 31B on RTX 5090?

Up to 61.1 tokens/s (llm), the fastest community-measured generation speed.

Are there step-by-step instructions for Gemma4 31B on RTX 5090?

Yes — a step-by-step recipe documents Gemma4 31B on the RTX 5090, linked at the top of this page.