§01·tested · /gpus/rtx-5070
Recipes tested on RTX 5070
29 community-tested setups — recipes and guides whose author ran them on this exact card.
nvidia50 series12GB VRAM
- multimodalbeginner6GB+recipe
Gemma 4 E4B on RTX 5070: Multimodal Inference via Q4_K_M GGUF (llama.cpp or Ollama — BF16 will not fit)
- llmbeginner10GB+recipe
Llama 3.1 8B on RTX 5070: Local Chat via Ollama or llama.cpp + Unsloth UD-Q4_K_XL GGUF
- imageintermediate11GB+recipe
Chroma1-Base (V48) on RTX 5070: Uncensored 8.9B FLUX.1-Schnell De-Distillation via GGUF Q4_K_M in ComfyUI
- imageintermediate12GB+recipe
Qwen-Image on RTX 5070: 20B Text-to-Image via ComfyUI GGUF Q3 (Blackwell sm_120, 12 GB)
- imageintermediate12GB+recipe
Juggernaut Z on RTX 5070: Cinematic Photoreal Fine-Tune of Z-Image Base via FP8 in ComfyUI
- imageintermediate12GB+recipe
SenseNova U1 (8B-MoT) on RTX 5070: VAE-Free Unified Image Gen + Understanding via Q4 GGUF + Layer Offload
- imageintermediate12GB+recipe
ERNIE-Image-Turbo on RTX 5070: 8-step text-to-image via GGUF in ComfyUI
- videointermediate8GB+recipe
Wan 2.2 TI2V-5B on RTX 5070: 720p Text/Image-to-Video in ComfyUI
- videointermediate8GB+recipe
LightX2V on RTX 5070: 4-Step Text-to-Video with Distilled Wan2.1-14B via Blackwell-Native FP8 + Offload
- ttsintermediate8GB+recipe
ACE-Step 1.5 XL on RTX 5070: Text-to-Music Generation via the 8 GB Optimization Path
- ttsintermediate12GB+recipe
MOSS-Audio 4B-Instruct on RTX 5070: local audio understanding in a tight 12 GB
- ttsintermediate8GB+recipe
Foundation-1 on RTX 5070: Structured Music Sample Generation
- ttsintermediate4GB+recipe
OmniVoice on RTX 5070: Zero-Shot Voice Cloning Across 646 Languages
- ttsintermediate10GB+recipe
Voxtral Mini 3B on RTX 5070: local speech understanding in ~9.5 GB
- ttsintermediate8GB+recipe
Qwen3-TTS 1.7B-Base on RTX 5070: Multilingual Voice Cloning in 10 Languages
- ttsbeginner8GB+recipe
VoxCPM2 on RTX 5070: 30-Language 48kHz Voice Cloning in ~8 GB VRAM
- ttsintermediate5GB+recipe
OpenAudio S1 Mini on RTX 5070: 13-Language Distilled TTS in ~5 GB VRAM
- ttsbeginner5GB+recipe
VoxCPM-0.5B on RTX 5070: Zero-Shot Voice Cloning TTS in ~5 GB VRAM
- ttsbeginner2GB+recipe
Kokoro TTS on RTX 5070: 82M-Parameter Text-to-Speech, 54 Voices, ~10 GB Free to Colocate a Second Model
- specializedintermediate3GB+recipe
KiMoDo on RTX 5070: Text-to-3D-Motion Generation Guide
- specializedbeginner4GB+recipe
SAM 3 on RTX 5070: Promptable Image and Video Segmentation
- multimodalintermediate4GB+recipe
MiniMind-O on RTX 5070: 0.1B Omni Model with Headroom to Spare
- 3dintermediate10GB+recipe
Hunyuan3D-2.1 on RTX 5070: Image-to-Mesh 3D Generation (Shape-Only)
- imagebeginner7GB+recipe
Anima 2B on RTX 5070: Native ComfyUI Anime Text-to-Image
- imagebeginner8GB+recipe
Flux.2 Klein 4B on RTX 5070: Blackwell-Native FP8 4-Step Text-to-Image at ~8.4 GB
- imageintermediate10GB+recipe
HiDream-O1-Image on RTX 5070: 2048×2048 Text-to-Image with FP8 in ComfyUI
- llmintermediate12GB+recipe
gpt-oss 20B on RTX 5070: MXFP4 Chat in 12 GB via llama.cpp Expert Offload
- llmbeginner12GB+recipe
Qwen3-14B on RTX 5070: Q4_K_M GGUF via Ollama or llama.cpp
- llmbeginner12GB+recipe
Qwen3-8B on RTX 5070: Q4_K_M GGUF via Ollama or llama.cpp