§01·spec · /gpus
RTX 4060
nvidia40 series8GB VRAM
14 open-weights AI models run on the RTX 4060 — see which fit, how fast they go, and the VRAM each needs.
§02·models that run on this GPU
14 totalLLM · 5
| Model | Best speed | Min VRAM | Works | Evidence | |
|---|---|---|---|---|---|
| Gemma 4 12B | 8GB | ✓ | recipe | check ↗ | |
| Mistral Nemo 12B | 8GB | ✓ | recipe | check ↗ | |
| Nanbeige4.2 3B | 8GB | ✓ | recipe | check ↗ | |
| Ornith 1.0 9B | 8GB | ✓ | recipe | check ↗ | |
| Qwen3-4B | 4GB | ✓ | recipe | check ↗ |
Multimodal · 3
| Model | Best speed | Min VRAM | Works | Evidence | |
|---|---|---|---|---|---|
| Fara1.5-4B | 8GB | ✓ | recipe | check ↗ | |
| Gemma 4 E4B-IT | 6GB | ✓ | recipe | check ↗ | |
| MiniMind-O | 4GB | ✓ | recipe | check ↗ |
TTS · 4
Music · 1
| Model | Best speed | Min VRAM | Works | Evidence | |
|---|---|---|---|---|---|
| Foundation-1 | 8GB | ✓ | recipe | check ↗ |
§03·tested recipes
showing 6 of 14- llmintermediate8GB+recipe
Nanbeige4.2-3B on RTX 4060: 32K Context on 8 GB Despite the Looped-Transformer KV Tax
- multimodaladvanced8GB+recipe
Fara1.5-4B on RTX 4060: an 8GB Browser Computer-Use Agent via llama.cpp Vision
- llmintermediate8GB+recipe
Gemma 4 12B on RTX 4060: Local Private Assistant via llama.cpp / Ollama (8GB)
- llmintermediate8GB+recipe
Mistral Nemo 12B on RTX 4060: Local Private Assistant via llama.cpp / Ollama (8GB)
- llmintermediate8GB+recipe
Ornith 1.0 9B on RTX 4060 (8GB): Local Agentic Coding at the Fit Boundary via llama.cpp + OpenHands
- multimodalbeginner6GB+recipe
Gemma 4 E4B on RTX 4060: Multimodal Inference via Q4_K_M GGUF (llama.cpp or Ollama)