§01·compatibility · /check
Gemma 4 26B MoE on RTX 3060
Yes — Gemma 4 26B MoE runs on the RTX 3060 (12 GB). Fastest community-measured result: 37.2 tok/s.
✓ runsmultimodalactive30 series12GB VRAM
model
- name
- Gemma 4 26B MoE
- slug
- gemma4-26b
- vertical
- multimodal
- status
- active
- repo
- huggingface.co ↗
§02·benchmarks
| Task | Quant | Speed | VRAM | Works | Confidence | Source | Verified |
|---|---|---|---|---|---|---|---|
| chat | Q4_K_M (Unsloth UD) | 37.2tok/s | ✓ | smeltcore.com· community | 2026-08-09 |
§03·common questions
Can you run Gemma 4 26B MoE on RTX 3060?
Yes — Gemma 4 26B MoE runs on the RTX 3060 (12 GB). Fastest community-measured result: 37.2 tok/s.
Which quantizations have been tested for Gemma 4 26B MoE on RTX 3060?
Q4_K_M (Unsloth UD) — measured in community benchmarks.
How fast is Gemma 4 26B MoE on RTX 3060?
Up to 37.2 tok/s (chat), the fastest community-measured result.
Where can I find step-by-step recipes for Gemma 4 26B MoE?
No recipe targets the RTX 3060 specifically yet, but 4 published recipes cover Gemma 4 26B MoE on other GPUs — a solid starting point. See the recipes listed below.
§04·more Gemma 4 26B MoE recipes
- llmintermediate18GB+
Gemma 4 26B A4B-it on RTX 3090 Ti: Local Multimodal Chat via Q4_K_M GGUF + llama.cpp
- llmintermediate29GB+
Gemma 4 26B A4B-it on RTX 5090: Q8_0 Quality Tier via ggml-org GGUF + llama.cpp
- llmintermediate18GB+
Gemma 4 26B A4B-it on RTX 3090: Local Multimodal Chat via Q4_K_M GGUF + llama.cpp
- llmintermediate18GB+
Gemma 4 26B A4B-it on RTX 4090: Local Multimodal Chat via Q4_K_M GGUF + llama.cpp