Gemma 4 26B MoE on RTX 3060
Yes — Gemma 4 26B MoE runs on the RTX 3060 (12 GB). Fastest community-measured result: 37.2 tok/s, with 10.92 GB peak VRAM.
- name
- Gemma 4 26B MoE
- slug
- gemma4-26b
- vertical
- multimodal
- status
- active
- repo
- huggingface.co ↗
| Task | Quant | Speed | VRAM | Works | Confidence | Source | Verified |
|---|---|---|---|---|---|---|---|
| llm | Q4_K_M (Unsloth UD) | 37.2tok/s | 10.92GB | ✓ | smeltcore.com· community | 2026-08-09 |
- llmQ4_K_M (Unsloth UD)37.2 tok/s
Expert-offload run: routed experts pushed to system RAM via --n-cpu-moe (-ncmoe 12), attention + shared weights on GPU. llama.cpp b10088, CUDA sm_86, flash-attn on. Rig: i7-7700, 32GB DDR4-2133 dual-channel (~34 GB/s), headless 3060. Weights 15.8 GB (UD-Q4_K_M, 15.77 GiB), 30 layers, 128 experts, 8 active (~4B). Generation speed by context depth: 37.2 / 36.1 / 35.8 tok/s at 0 / 4K / 8K. Offload ladder: -ncmoe 12 = 37.23 tok/s at 11,179 MiB peak; -ncmoe 30 = 21.20 tok/s; -ncmoe 8 and below fail to load outright. That peak is 11,179 of the card's 12,288 MiB — roughly 1.1 GB spare, which is why the lower rungs fail rather than merely slow down. Measured 2026-07-22. The submitter published the full ladder at insiderllm.com/guides/gemma-4-2gb-ram-ssd-streaming/, which is also where the quantization label comes from.
recorded from smeltcore.com ↗ · verified 2026-08-09
- llmintermediate18GB+
Gemma 4 26B A4B-it on RTX 3090 Ti: Local Multimodal Chat via Q4_K_M GGUF + llama.cpp
- llmintermediate29GB+
Gemma 4 26B A4B-it on RTX 5090: Q8_0 Quality Tier via ggml-org GGUF + llama.cpp
- llmintermediate18GB+
Gemma 4 26B A4B-it on RTX 3090: Local Multimodal Chat via Q4_K_M GGUF + llama.cpp
- llmintermediate18GB+
Gemma 4 26B A4B-it on RTX 4090: Local Multimodal Chat via Q4_K_M GGUF + llama.cpp
No step-by-step recipe covers Gemma 4 26B MoE on the RTX 3060 yet, but 39 other models have a step-by-step recipe written for this exact card. Other Multimodal models first.
- multimodal12GB+
Fara1.5-4B
open recipe → - multimodal12GB+
Fara1.5-9B
open recipe → - multimodal6GB+
Gemma 4 E4B-IT
open recipe → - multimodal4GB+
MiniMind-O
open recipe → - multimodal12GB+
MOSS-Audio
open recipe → - multimodal9.8GB+
Qwen3.6 35B-A3B
open recipe →
Can you run Gemma 4 26B MoE on RTX 3060?
Yes — Gemma 4 26B MoE runs on the RTX 3060 (12 GB). Fastest community-measured result: 37.2 tok/s, with 10.92 GB peak VRAM.
How much VRAM does Gemma 4 26B MoE need on RTX 3060?
Measured peak VRAM is 10.92 GB — about 1.1 GB of headroom on the 12 GB card.
Which quantizations have been tested for Gemma 4 26B MoE on RTX 3060?
Q4_K_M (Unsloth UD) — measured in community benchmarks.
How fast is Gemma 4 26B MoE on RTX 3060?
Up to 37.2 tok/s (llm), the fastest community-measured result.
Where can I find step-by-step recipes for Gemma 4 26B MoE?
No recipe targets the RTX 3060 specifically yet, but 4 published recipes cover Gemma 4 26B MoE on other GPUs — a solid starting point, linked on this page.