§01·index · /recipes
Recipes
930 community-tested setups for running open-weights AI models on real consumer GPUs.page 3 of 10
- llmintermediate12GB+
Ornith 1.0 9B on RTX 3080 Ti: A Local Agentic-Coding Model in 12GB via llama.cpp + OpenHands
- llmintermediate8GB+
Ornith 1.0 9B on RTX 5060 (8GB): Local Agentic Coding at the Fit Boundary via llama.cpp + OpenHands
- llmintermediate8GB+
Ornith 1.0 9B on RTX 4060 Ti 8GB: Local Agentic Coding at the Fit Boundary via llama.cpp + OpenHands
- llmintermediate8GB+
Ornith 1.0 9B on RTX 3060 Ti (8GB): Local Agentic Coding at the Fit Boundary via llama.cpp + OpenHands
- llmintermediate12GB+
Ornith 1.0 9B on RTX 5070: A Local Agentic-Coding Model in 12GB via llama.cpp + OpenHands
- llmintermediate8GB+
Ornith 1.0 9B on RTX 4060 (8GB): Local Agentic Coding at the Fit Boundary via llama.cpp + OpenHands
- llmintermediate12GB+
Ornith 1.0 9B on RTX 3060 (12GB): A Local Agentic-Coding Model on a Budget Card via llama.cpp + OpenHands
- llmintermediate12GB+
Ornith 1.0 9B on RTX 4080: Max-Fidelity Local Agentic Coding in 16GB via llama.cpp + OpenHands
- llmadvanced48GB+
Ornith 1.0 35B on Apple M4 Max: Local Agentic Coding via llama.cpp Metal + OpenHands (48GB Unified Memory)
- llmadvanced32GB+
Ornith 1.0 35B on RTX 5090: Comfortable Local Agentic Coding via llama.cpp + OpenHands (32GB Tier)
- llmadvanced24GB+
Ornith 1.0 35B on RTX 3090 Ti: Local Agentic Coding via llama.cpp + OpenHands (24GB Entry Tier)
- llmadvanced24GB+
Ornith 1.0 35B on RTX 4090: Local Agentic Coding via llama.cpp + OpenHands (24GB Entry Tier)
- llmintermediate12GB+
Ornith 1.0 9B on RTX 4070: A Local Agentic-Coding Model in 12GB via llama.cpp + OpenHands
- llmadvanced24GB+
Ornith 1.0 35B on RTX 3090: Local Agentic Coding via llama.cpp + OpenHands (24GB Entry Tier)
- multimodalintermediate24GB+
Qwen3.5-35B-A3B on RTX 5090: Blackwell MXFP4 MoE Chat at 165 tok/s
- multimodalintermediate20GB+
Qwen3.5 27B on RTX 5090: Q4_K GGUF local chat via llama.cpp
- videoadvanced14GB+
LTX-2.3 on RTX 4060 Ti 16GB: 22B Audio-Video at the 16 GB Floor via Distilled GGUF + Streamed Encoder
- llmadvanced24GB+
Llama 3.3 70B on RTX 4090: 70B-Class Chat on One 24 GB Card (Q4 Offload or Fully-On-GPU IQ2)
- llmbeginner12GB+
Qwen3-14B on RTX 4060 Ti 16GB: Q4_K_M GGUF via Ollama or llama.cpp
- llmbeginner8GB+
Llama 3.1 8B on RTX 3060 Ti: Local Chat via Ollama or llama.cpp + Unsloth UD-Q4_K_XL GGUF
- llmintermediate20GB+
Gemma 4 31B on RTX 5090: dense 31B local chat at 61 tok/s, with Q5/Q6 quality headroom in 32 GB
- videointermediate10GB+
CogVideoX 1.5 on RTX 4070 Super: 1360x768 Text-to-Video with Diffusers
- imageintermediate8GB+
Anima 2B on RTX 5060: 8GB Anime Text-to-Image via INT8 ConvRot in ComfyUI
- llmbeginner16GB+
gpt-oss 20B on RTX 5060 Ti: MXFP4 Chat at 92 tok/s via Ollama or vLLM
- llmbeginner16GB+
gpt-oss 20B on RTX 4060 Ti 16GB: MXFP4 chat at 63 tok/s via Ollama or vLLM
- imagebeginner24GB+
Flux.1 Dev on RTX 5090: ComfyUI Image Generation Guide
- imagebeginner24GB+
Flux.1 Dev on RTX 3090 Ti: ComfyUI Image Generation Guide
- llmintermediate19GB+
Qwen3-30B-A3B on RTX 5090: 226 tok/s MoE Chat with Room to Spare
- llmintermediate24GB+
Qwen3-30B-A3B on RTX 3090 Ti: 167 tok/s MoE Chat That Fits the Full 24 GB Card
- llmintermediate24GB+
Qwen3-30B-A3B on RTX 3090: Full-GPU MoE Chat at 153 tok/s
- llmadvanced32GB+
Llama 3.1 70B on RTX 5090: Fitting a 70B Model on One 32 GB Card with IQ3 GGUF
- imagebeginner4GB+
Stable Diffusion 1.5 on RTX 3060: 512x512 Image Generation with 12 GB to Spare
- llmadvanced12GB+
Qwen3.6-35B-A3B on RTX 4070: 80 tok/s + 128K context from a 12GB card via MoE CPU-offload
- llmbeginner9GB+
Qwen2.5-14B-Instruct on RX 7900 XTX: a fast local chat LLM via Ollama (ROCm)
- llmbeginner2GB+
Llama 3.2 1B on RX 7900 XTX: Fast Local Chat via Ollama (ROCm)
- multimodalintermediate24GB+
Qwen3.5-35B-A3B on RTX 3090: MXFP4 MoE Chat at 111 tok/s
- multimodalintermediate24GB+
Qwen3.5-27B on RTX 3090: Q4_K GGUF local chat via llama.cpp
- llmintermediate24GB+
Gemma 4 31B on RTX 3090: dense 31B local chat via Q4_K_M GGUF in Ollama / llama.cpp
- llmbeginner8GB+
StableLM 2 12B on RTX 3060 Ti: Run a Multilingual Base LLM at the 8 GB Wall
- llmbeginner8GB+
Llama 2 13B on RTX 3060 Ti: The Slow Ceiling of What 8 GB Holds
- llmbeginner8GB+
Gemma 7B on RTX 3060 Ti: Run a Local Chat LLM at the 8 GB Floor
- llmbeginner8GB+
Falcon2 11B on RTX 3060 Ti: Run TII's 11B LLM Local at the 8 GB Ceiling
- llmbeginner8GB+
WizardLM-2 7B on RTX 3060 Ti: Run Microsoft's Evol-Instruct Chat LLM Locally at the 8 GB Floor
- llmbeginner8GB+
Qwen2 7B on RTX 3060 Ti: Run a Local Chat LLM at the 8 GB Floor
- llmbeginner8GB+
Llama 2 7B on RTX 3060 Ti: Run a Local Chatbot at the 8 GB Floor
- llmbeginner8GB+
Gemma 2 9B on RTX 3060 Ti: A Capable Local LLM at the 8 GB Floor
- llmbeginner8GB+
Qwen2.5 7B on RTX 3060 Ti: Run a Local LLM at the 8 GB Floor
- videoadvanced24GB+
LTX-2 (19B) on RTX 4090: Audio-Video in the 24 GB Gap — FP8 + Quantized Gemma, Below the 32 GB Native Floor
- videointermediate10GB+
LTX-Video 2B on RTX 4060 Ti 16GB: Fast Local Text-to-Video
- multimodalbeginner8GB+
LLaVA 7B on RTX 3060 Ti: Ask Questions About Images Locally at the 8 GB Floor
- videointermediate6GB+
AnimateDiff on RTX 3060 Ti: Low-VRAM SD1.5 Animation on 8GB
- videoadvanced20GB+
LTX-2.3 on RTX 4090: 22B Audio-Video in the 24 GB Gap — Roomier GGUF, Still Below the 32 GB Native Floor
- imageintermediate12GB+
Krea 2 Turbo on RTX 5070 (12GB) via ComfyUI + GGUF: 8-Step Text-to-Image
- imageintermediate12GB+
Krea 2 Turbo on RTX 3080 Ti (12GB) via ComfyUI + GGUF: 8-Step Text-to-Image
- imageintermediate12GB+
Krea 2 Turbo on RTX 3060 (12GB) via ComfyUI + GGUF: 8-Step Text-to-Image
- imageintermediate12GB+
Krea 2 Turbo on RTX 4070 Ti (12GB) via ComfyUI + GGUF: 8-Step Text-to-Image
- imageintermediate12GB+
Krea 2 Turbo on RTX 4070 SUPER (12GB) via ComfyUI + GGUF: 8-Step Text-to-Image
- imageintermediate12GB+
Krea 2 Turbo on RTX 4070 (12GB) via ComfyUI + GGUF: 8-Step Text-to-Image
- videoadvanced14GB+
LTX-2.3 on Apple M2 Pro (16 GB): experimental 22B audio-video in unified memory via the MLX --low-ram port
- ttsbeginner2GB+
VoxCPM-0.5B on Apple M2 Pro: Zero-Shot Voice Cloning TTS in Unified Memory (MPS)
- ttsbeginner1GB+
Kokoro TTS on Apple M2 Pro: 82M Text-to-Speech, 54 Voices, Native MLX-Audio
- imagebeginner6GB+
Z-Image Turbo on Apple M2 Pro: 8-step 1024x1024 text-to-image in 16 GB unified memory with mflux
- multimodalbeginner5GB+
Gemma 4 E4B on Apple M2 Pro: local vision-language inference in 16 GB unified memory with MLX-VLM
- llmintermediate13GB+
gpt-oss 20B on Apple M2 Pro: running 20B in 16 GB unified memory with the wired-limit raise
- llmintermediate9GB+
Qwen3-14B on Apple M2 Pro: the strongest LLM that fits a 16 GB unified-memory Mac, via MLX 4-bit
- llmbeginner5GB+
Qwen3-8B on Apple M2 Pro: local 8B chat with MLX 4-bit in 16 GB unified memory
- llmbeginner5GB+
Llama 3.1 8B on Apple M2 Pro: your first local LLM in 16 GB unified memory with MLX
- 3dadvanced17GB+
TRELLIS.2-4B on Apple M3 Max: image-to-3D in unified memory via the community Metal port
- videoadvanced34GB+
LTX-2.3 on Apple M3 Max: experimental 22B audio-video in unified memory via Draw Things (Metal) or MLX
- ttsbeginner2GB+
VoxCPM-0.5B on Apple M3 Max: Zero-Shot Voice Cloning TTS in Unified Memory (MPS)
- ttsbeginner1GB+
Kokoro TTS on Apple M3 Max: 82M Text-to-Speech, 54 Voices, Native MLX-Audio
- imageintermediate23GB+
Qwen-Image on Apple M3 Max: 20B text-to-image in unified memory with mflux
- imagebeginner17GB+
Z-Image Turbo on Apple M3 Max: 8-step 1024x1024 text-to-image in unified memory with mflux
- multimodalbeginner5GB+
Gemma 4 E4B on Apple M3 Max: local vision-language inference in unified memory with MLX-VLM
- llmbeginner5GB+
Llama 3.1 8B on Apple M3 Max: the easy local-LLM on-ramp in unified memory with MLX
- llmbeginner13GB+
gpt-oss 20B on Apple M3 Max: native-MXFP4 chat in 48 GB unified memory with MLX
- llmintermediate19GB+
Qwen3-32B on Apple M3 Max: 32B local chat with MLX 4-bit in 48 GB unified memory
- llmadvanced40GB+
Llama 3.3 70B on Apple M3 Max: 70B-class chat in 48 GB unified memory with MLX
- imageintermediate24GB+
Krea 2 Turbo on Apple M3 Max: 8-Step Text-to-Image in Unified Memory via ComfyUI (MPS)
- imageintermediate24GB+
Krea 2 Turbo on Apple M4 Max: 8-Step Text-to-Image in Unified Memory via ComfyUI (MPS)
- imageintermediate24GB+
Krea 2 Turbo on Apple M2 Max: 8-Step Text-to-Image in Unified Memory via ComfyUI (MPS)
- imageintermediate16GB+
Krea 2 Turbo on RX 7800 XT via ComfyUI + ROCm: GGUF Text-to-Image in 16GB
- imageintermediate16GB+
Krea 2 Turbo on RX 7900 XTX via ComfyUI + ROCm: GGUF Text-to-Image in 24GB
- imageintermediate16GB+
Krea 2 Turbo (FP8) on RTX 4090 via ComfyUI: 8-Step Text-to-Image, with the Raw Tier in Reach
- imageintermediate16GB+
Krea 2 Turbo (FP8) on RTX 3090 Ti via ComfyUI: 8-Step Text-to-Image, with the Raw Tier in Reach
- imageintermediate16GB+
Krea 2 Turbo (FP8) on RTX 3090 via ComfyUI: 8-Step Text-to-Image, with the Raw Tier in Reach
- imageintermediate16GB+
Krea 2 Turbo (FP8) on RTX 4060 Ti 16GB via ComfyUI: 8-Step Text-to-Image in 16GB
- imageintermediate16GB+
Krea 2 Turbo (FP8) on RTX 4070 Ti Super via ComfyUI: 8-Step Text-to-Image in 16GB
- imageintermediate16GB+
Krea 2 Turbo (FP8) on RTX 4080 Super via ComfyUI: 8-Step Text-to-Image in 16GB
- imageintermediate16GB+
Krea 2 Turbo (FP8) on RTX 4080 via ComfyUI: 8-Step Text-to-Image in 16GB
- imageintermediate16GB+
Krea 2 on RTX 5090 via ComfyUI: 8-Step Turbo + Resident Full-Quality Raw in 32GB
- imageintermediate16GB+
Krea 2 Turbo (FP8) on RTX 5070 Ti via ComfyUI: 8-Step Text-to-Image in 16GB
- imageintermediate16GB+
Krea 2 Turbo (FP8) on RTX 5080 via ComfyUI: 8-Step Text-to-Image in 16GB
- imageintermediate16GB+
Krea 2 Turbo (FP8) on RTX 5060 Ti via ComfyUI: 8-Step Text-to-Image in 16GB
- 3dadvanced17GB+
TRELLIS.2-4B on Apple M4 Max: image-to-3D in unified memory via the community Metal port
- videoadvanced34GB+
LTX-2.3 on Apple M4 Max: experimental 22B audio-video in unified memory via Draw Things (Metal) or MLX
- ttsbeginner2GB+
VoxCPM-0.5B on Apple M4 Max: Zero-Shot Voice Cloning TTS in Unified Memory (MPS)
- ttsbeginner1GB+
Kokoro TTS on Apple M4 Max: 82M Text-to-Speech, 54 Voices, Native MLX-Audio
- imageintermediate23GB+
Qwen-Image on Apple M4 Max: 20B text-to-image in unified memory with mflux
- imagebeginner17GB+
Z-Image Turbo on Apple M4 Max: 8-step 1024x1024 text-to-image in unified memory with mflux