§01·compatibility · /check
Gemma 4 26B MoE on RTX 5090
Yes — Gemma 4 26B MoE runs on the RTX 5090 (32 GB). Fastest community-measured generation speed: 180.3 tokens/s.
✓ runsmultimodalactive50 series32GB VRAM
step-by-step recipe for this pair
Gemma 4 26B A4B-it on RTX 5090: Q8_0 Quality Tier via ggml-org GGUF + llama.cpp
llmintermediate29GB+
model
- name
- Gemma 4 26B MoE
- slug
- gemma4-26b
- vertical
- multimodal
- status
- active
- repo
- huggingface.co ↗
§02·benchmarks
| Task | Quant | Speed | VRAM | Works | Confidence | Source | Verified |
|---|---|---|---|---|---|---|---|
| llm | Q4_K | 8799.2prefill tokens/s | ✓ | hardware-corner.net· web | 2026-05-15 | ||
| llm | Q4_K | 180.3tokens/s | ✓ | hardware-corner.net· web | 2026-05-15 |
§03·how it was measured
- llmQ4_K8799.2 prefill tokens/s
4k Context
recorded from hardware-corner.net ↗ · verified 2026-05-15
- llmQ4_K180.3 tokens/s
4k Context
recorded from hardware-corner.net ↗ · verified 2026-05-15
§04·more Gemma 4 26B MoE recipes
- llmintermediate18GB+
Gemma 4 26B A4B-it on RTX 3090 Ti: Local Multimodal Chat via Q4_K_M GGUF + llama.cpp
- llmintermediate18GB+
Gemma 4 26B A4B-it on RTX 3090: Local Multimodal Chat via Q4_K_M GGUF + llama.cpp
- llmintermediate18GB+
Gemma 4 26B A4B-it on RTX 4090: Local Multimodal Chat via Q4_K_M GGUF + llama.cpp
§05·common questions
Can you run Gemma 4 26B MoE on RTX 5090?
Yes — Gemma 4 26B MoE runs on the RTX 5090 (32 GB). Fastest community-measured generation speed: 180.3 tokens/s.
Which quantizations have been tested for Gemma 4 26B MoE on RTX 5090?
Q4_K — measured in community benchmarks.
How fast is Gemma 4 26B MoE on RTX 5090?
Up to 180.3 tokens/s (llm), the fastest community-measured generation speed.
Are there step-by-step instructions for Gemma 4 26B MoE on RTX 5090?
Yes — a step-by-step recipe documents Gemma 4 26B MoE on the RTX 5090, linked at the top of this page.