self-hosted/ai
§01·compatibility · /check

Gemma 4 26B MoE on RTX 3060

Yes — Gemma 4 26B MoE runs on the RTX 3060 (12 GB). Fastest community-measured result: 37.2 tok/s, with 10.92 GB peak VRAM.

✓ runsmultimodalactive30 series12GB VRAM
model
name
Gemma 4 26B MoE
slug
gemma4-26b
vertical
multimodal
status
active
open detail ↗
gpu
name
RTX 3060
slug
rtx-3060
vram
12 GB
series
30
open detail ↗
§02·benchmarks
TaskQuantSpeedVRAMWorksConfidenceSourceVerified
llmQ4_K_M (Unsloth UD)37.2tok/s10.92GB✓smeltcore.com· community2026-08-09
§03·how it was measured
  • llmQ4_K_M (Unsloth UD)37.2 tok/s

    Expert-offload run: routed experts pushed to system RAM via --n-cpu-moe (-ncmoe 12), attention + shared weights on GPU. llama.cpp b10088, CUDA sm_86, flash-attn on. Rig: i7-7700, 32GB DDR4-2133 dual-channel (~34 GB/s), headless 3060. Weights 15.8 GB (UD-Q4_K_M, 15.77 GiB), 30 layers, 128 experts, 8 active (~4B). Generation speed by context depth: 37.2 / 36.1 / 35.8 tok/s at 0 / 4K / 8K. Offload ladder: -ncmoe 12 = 37.23 tok/s at 11,179 MiB peak; -ncmoe 30 = 21.20 tok/s; -ncmoe 8 and below fail to load outright. That peak is 11,179 of the card's 12,288 MiB — roughly 1.1 GB spare, which is why the lower rungs fail rather than merely slow down. Measured 2026-07-22. The submitter published the full ladder at insiderllm.com/guides/gemma-4-2gb-ram-ssd-streaming/, which is also where the quantization label comes from.

    recorded from smeltcore.com ↗ · verified 2026-08-09

§04·more Gemma 4 26B MoE recipes
§05·what runs on the RTX 3060

No step-by-step recipe covers Gemma 4 26B MoE on the RTX 3060 yet, but 39 other models have a step-by-step recipe written for this exact card. Other Multimodal models first.

all 39 recipes tested on the RTX 3060 ↗
§06·request a recipe

Nobody has written up Gemma 4 26B MoE on the RTX 3060 yet. Ask for it and it goes on the list — requests are how we decide what to measure next.

Email is optional — leave it blank and the request still counts.

§07·common questions
Can you run Gemma 4 26B MoE on RTX 3060?

Yes — Gemma 4 26B MoE runs on the RTX 3060 (12 GB). Fastest community-measured result: 37.2 tok/s, with 10.92 GB peak VRAM.

How much VRAM does Gemma 4 26B MoE need on RTX 3060?

Measured peak VRAM is 10.92 GB — about 1.1 GB of headroom on the 12 GB card.

Which quantizations have been tested for Gemma 4 26B MoE on RTX 3060?

Q4_K_M (Unsloth UD) — measured in community benchmarks.

How fast is Gemma 4 26B MoE on RTX 3060?

Up to 37.2 tok/s (llm), the fastest community-measured result.

Where can I find step-by-step recipes for Gemma 4 26B MoE?

No recipe targets the RTX 3060 specifically yet, but 4 published recipes cover Gemma 4 26B MoE on other GPUs — a solid starting point, linked on this page.