§01·index · /recipes
Recipes
950 community-tested setups for running open-weights AI models on real consumer GPUs.page 1 of 10
- multimodaladvanced64GB+
Muse Glimmer 30B on Apple M2 Max: 8-bit MLX with vision and the DFlash drafter
- multimodaladvanced36GB+
Muse Glimmer 30B on Apple M3 Max: ExecuTorch Metal agent server with vision and DFlash
- multimodalintermediate24GB+
Muse Glimmer 30B on RTX 4090: vision and DFlash in llama.cpp at full 131K context
- multimodalintermediate24GB+
Muse Glimmer 30B on RTX 3090 Ti: vision + DFlash speculation at full 131K context
- multimodalintermediate48GB+
Qwen3.8-27B on Apple M4 Max: 4-bit MLX Vision-Language at the Full 262K Context
- multimodalintermediate48GB+
Qwen3.8-27B on Apple M3 Max: 4-bit MLX Vision-Language at the Full 262K Context
- multimodaladvanced16GB+
Qwen3.8-27B on RX 7800 XT: a vision-capable 27B inside 16 GB on ROCm
- multimodalintermediate24GB+
Qwen3.8-27B on RX 7900 XTX: 128K-context vision chat on ROCm with llama.cpp
- multimodaladvanced16GB+
Qwen3.8-27B on RTX 5060 Ti: a vision-capable 27B in 16 GB on a 128-bit bus
- multimodaladvanced16GB+
Qwen3.8-27B on RTX 4060 Ti 16GB: a vision-capable 27B on 288 GB/s
- multimodaladvanced16GB+
Qwen3.8-27B on RTX 4070 Ti Super: a vision-capable 27B inside 16 GB with llama.cpp
- multimodaladvanced16GB+
Qwen3.8-27B on RTX 5080: a vision-capable 27B inside 16 GB with llama.cpp
- multimodaladvanced16GB+
Qwen3.8-27B on RTX 4080 Super: a vision-capable 27B inside 16 GB with llama.cpp
- multimodaladvanced16GB+
Qwen3.8-27B on RTX 4080: a vision-capable 27B inside 16 GB with llama.cpp
- multimodalintermediate24GB+
Qwen3.8-27B on RTX 4090: 128K-context vision chat via llama.cpp Q4_K_M
- multimodalintermediate24GB+
Qwen3.8-27B on RTX 3090 Ti: 128K-context vision chat via llama.cpp Q4_K_M
- multimodalintermediate48GB+
Qwen3.8-27B on Apple M2 Max: 8-bit MLX Vision-Language with MTP Speculative Decoding
- multimodaladvanced16GB+
Qwen3.8-27B on RTX 5070 Ti: a vision-capable 27B inside 16 GB with llama.cpp
- multimodalintermediate32GB+
Qwen3.8-27B on RTX 5090: the full 262K context window on one card
- multimodalintermediate24GB+
Qwen3.8-27B on RTX 3090: 128K-context vision chat via llama.cpp GGUF
- videoadvanced16GB+
LTX-2.5 on RTX 4070 Ti SUPER: 22B audio-video in 16 GB the plain Ti does not have
- videoadvanced16GB+
LTX-2.5 on RTX 4080 SUPER: 22B audio-video in 16 GB, and why the SUPER badge changes nothing
- videoadvanced16GB+
LTX-2.5 on RTX 4080: 22B audio-video in 16 GB via a community Q3_K_M GGUF on Ada
- videoadvanced16GB+
LTX-2.5 on RTX 5070 Ti: 22B audio-video in 16 GB, and the decode trap three owners hit
- videoadvanced16GB+
LTX-2.5 on RTX 5080: 22B audio-video in 16 GB, and what the wider bus cannot buy
- videoadvanced16GB+
LTX-2.5 on RTX 4060 Ti 16GB: 22B audio-video in 16 GB via a community Q3_K_M GGUF
- videoadvanced16GB+
LTX-2.5 on RTX 5060 Ti: 22B audio-video in 16 GB via a community Q3_K_M GGUF
- multimodaladvanced36GB+
Muse Glimmer 30B on Apple M4 Max: ExecuTorch Metal agent server with vision and DFlash
- multimodaladvanced24GB+
Muse Glimmer 30B on RX 7900 XTX: ROCm llama.cpp with vision and DFlash speculative decoding
- multimodalintermediate24GB+
Muse Glimmer 30B on RTX 3090: vision + DFlash speculation at full 131K context
- multimodalintermediate32GB+
Muse Glimmer 30B on RTX 5090: K-Quant-Dynamic, vision and DFlash in llama.cpp
- videoadvanced24GB+
LTX-2.5 on RTX 3090 Ti: 22B audio-video on Ampere via the official ComfyUI INT8 checkpoint
- videoadvanced24GB+
LTX-2.5 on RTX 3090: 22B audio-video on Ampere via the official ComfyUI INT8 checkpoint
- videoadvanced24GB+
LTX-2.5 on RTX 4090: Fitting the 22B Audio-Video DiT in 24 GB via ComfyUI int8
- videoadvanced24GB+
LTX-2.5 on RTX 5090: 22B Audio-Video in 32 GB via ComfyUI's INT8 Path
- llmintermediate36GB+
Nanbeige4.2-3B on Apple M4 Max: 546 GB/s for a stack that streams twice
- llmintermediate36GB+
Nanbeige4.2-3B on Apple M3 Max: the full 262,144-token context in 36 GiB
- llmintermediate16GB+
Nanbeige4.2-3B on RX 7800 XT: 96K context on a 16 GB gfx1101 ROCm build
- llmintermediate24GB+
Nanbeige4.2-3B on RTX 4090: the full 256K context on an sm_89 llama.cpp build
- llmintermediate24GB+
Nanbeige4.2-3B on RTX 3090 Ti: the full 256K context in VRAM with llama.cpp
- llmintermediate16GB+
Nanbeige4.2-3B on RTX 5080: the fastest 16 GB card stops at the same context
- llmintermediate16GB+
Nanbeige4.2-3B on RTX 5070 Ti: 96K context, and what twice the bandwidth cannot buy
- llmintermediate16GB+
Nanbeige4.2-3B on RTX 5060 Ti: the whole 16 GB context ladder on a 128-bit bus
- llmintermediate16GB+
Nanbeige4.2-3B on RTX 4080 SUPER: 16 GB Is the Limit, Not the Silicon
- llmintermediate16GB+
Nanbeige4.2-3B on RTX 4080: 96K Context, and What Each Token Costs in Bytes
- llmintermediate16GB+
Nanbeige4.2-3B on RTX 4070 Ti SUPER: 96K Context on 16 GB, One Rung at a Time
- llmintermediate12GB+
Nanbeige4.2-3B on RTX 5070: Q8_0 weights, 65,536-token cache, CUDA 12.8 floor
- llmintermediate12GB+
Nanbeige4.2-3B on RTX 4070 Ti: Q8_0 weights and a 65,536-token cache in 12 GB
- llmintermediate12GB+
Nanbeige4.2-3B on RTX 4070 SUPER: Q8_0 weights, 65,536-token cache in 12 GB
- llmintermediate12GB+
Nanbeige4.2-3B on RTX 4070: Q8_0 weights and a 65,536-token cache in 12 GB
- llmintermediate12GB+
Nanbeige4.2-3B on RTX 3080 Ti: Q8_0 weights and a 65,536-token cache in 12 GB
- llmintermediate8GB+
Nanbeige4.2-3B on RTX 5060: the One 8 GB Card With a Real 65,536-Token Run Behind It
- llmintermediate8GB+
Nanbeige4.2-3B on RTX 4060 Ti 8GB: an 8 GB Budget for 65,536 New Tokens
- llmintermediate8GB+
Nanbeige4.2-3B on RTX 3060 Ti: 32K Context in 8 GB, and What sm_86 Does Not Change
- llmintermediate16GB+
Nanbeige4.2-3B on Apple M2 Pro: the context ladder for 16 GB of unified memory
- llmintermediate24GB+
Nanbeige4.2-3B on RX 7900 XTX: the full 256K context on a ROCm llama.cpp build
- llmintermediate32GB+
Nanbeige4.2-3B on RTX 5090: 256K context with a near-lossless q8_0 KV cache
- llmintermediate16GB+
Nanbeige4.2-3B on RTX 4060 Ti 16GB: 96K Context Without a 4-bit KV Cache
- llmintermediate12GB+
Nanbeige4.2-3B on RTX 3060: Q8_0 weights and a 65,536-token cache in 12 GB
- llmintermediate16GB+
Nanbeige4.2-3B on Apple M2 Max: 128K-context agentic LLM via llama.cpp Metal
- llmintermediate24GB+
Nanbeige4.2-3B on RTX 3090: the full 256K context in VRAM with llama.cpp
- llmintermediate8GB+
Nanbeige4.2-3B on RTX 4060: 32K Context on 8 GB Despite the Looped-Transformer KV Tax
- videoadvanced12GB+
MiniMax H3 on RTX 5080: 16 GB, and what the extra silicon cannot buy
- videoadvanced12GB+
MiniMax H3 on RTX 5070 Ti: 16 GB, where only one big module fits
- videoadvanced12GB+
MiniMax H3 on RTX 5070: 12 GB, where nvfp4 is native and unused
- videoadvanced12GB+
MiniMax H3 on RTX 5060 Ti: 16 GB, two published runs, and the gap between them
- videoadvanced12GB+
MiniMax H3 on RTX 4090: 24 GB video+audio, and where Ada's FP8 shows up
- videoadvanced12GB+
MiniMax H3 on RTX 4080 SUPER: 16 GB, and the one 4080 number that is nearly yours
- videoadvanced12GB+
MiniMax H3 on RTX 4080: 16 GB Ada, and the only number this card has
- videoadvanced12GB+
MiniMax H3 on RTX 4070 Ti SUPER: Ada at 16 GB, and the Turbo LoRA question
- videoadvanced12GB+
MiniMax H3 on RTX 4070 Ti: 12 GB Ada, and the 16 GB card wearing your card's name
- videoadvanced12GB+
MiniMax H3 on RTX 4070 SUPER: 12 GB, where the host machine decides the run
- videoadvanced12GB+
MiniMax H3 on RTX 4060 Ti 16GB: the capacity fits, the bus is the question
- videoadvanced12GB+
MiniMax H3 on RTX 3090 Ti: 24 GB video+audio, and what sm_86 takes off the table
- videoadvanced12GB+
MiniMax H3 on RTX 3080 Ti: 12 GB video+audio in ComfyUI on Ampere sm_86
- videoadvanced12GB+
MiniMax H3 on RTX 3060: 12 GB video+audio in ComfyUI, measured on this card
- videoadvanced16GB+
MiniMax H3 on RX 7800 XT: 16 GB video+audio in ComfyUI on ROCm (gfx1101)
- videoadvanced12GB+
MiniMax H3 on RTX 5090: 32 GB video+audio, and what nvfp4 does not buy you
- videoadvanced64GB+
MiniMax H3 on Apple M2 Max: native MLX video with synchronised audio
- videoadvanced12GB+
MiniMax H3 on RTX 4070: 12 GB video+audio via ComfyUI VRAM offload
- videoadvanced12GB+
MiniMax H3 on RTX 3090: 5-second audio+video in ComfyUI on pruned int8
- multimodaladvanced48GB+
Fara1.5-27B on Apple M4 Max: Browser Computer-Use Agent on llama.cpp Metal
- multimodaladvanced24GB+
Fara1.5-9B on RX 7900 XTX: Unquantised Browser Computer-Use Agent on ROCm
- multimodaladvanced16GB+
Fara1.5-9B on RX 7800 XT: a Q8_0 Browser Computer-Use Agent on ROCm gfx1101
- multimodaladvanced24GB+
Fara1.5-9B on RTX 5090: BF16 Computer-Use Agent on Blackwell sm_120
- multimodaladvanced16GB+
Fara1.5-9B on RTX 5080: a Q8_0 Computer-Use Agent on Blackwell sm_120
- multimodaladvanced12GB+
Fara1.5-9B on RTX 5070: a Blackwell Browser Computer-Use Agent with llama.cpp
- multimodaladvanced16GB+
Fara1.5-9B on RTX 5070 Ti: a Q8_0 Computer-Use Agent on Blackwell sm_120
- multimodaladvanced16GB+
Fara1.5-9B on RTX 5060 Ti 16GB: a Q8_0 Computer-Use Agent on Blackwell
- multimodaladvanced24GB+
Fara1.5-9B on RTX 4090: BF16 Computer-Use Agent with Vision on Ada sm_89
- multimodaladvanced16GB+
Fara1.5-9B on RTX 4080: a Q8_0 Browser Computer-Use Agent with llama.cpp
- multimodaladvanced16GB+
Fara1.5-9B on RTX 4080 Super: a Q8_0 Browser Computer-Use Agent with llama.cpp
- multimodaladvanced12GB+
Fara1.5-9B on RTX 4070 Ti: a Q5_K_M Browser Computer-Use Agent with llama.cpp
- multimodaladvanced16GB+
Fara1.5-9B on RTX 4070 Ti Super: a Q8_0 Browser Computer-Use Agent with llama.cpp
- multimodaladvanced12GB+
Fara1.5-9B on RTX 4070 Super: a Local Browser Computer-Use Agent in 12GB
- multimodaladvanced16GB+
Fara1.5-9B on RTX 4060 Ti 16GB: a Q8_0 Browser Computer-Use Agent with llama.cpp
- multimodaladvanced24GB+
Fara1.5-9B on RTX 3090: BF16 Computer-Use Agent with Vision under llama.cpp
- multimodaladvanced24GB+
Fara1.5-9B on RTX 3090 Ti: BF16 Browser Computer-Use Agent on Ampere sm_86
- multimodaladvanced12GB+
Fara1.5-9B on RTX 3080 Ti: a Browser Computer-Use Agent at Q5_K_M with llama.cpp
- multimodaladvanced12GB+
Fara1.5-9B on RTX 3060: a 12GB Browser Computer-Use Agent at Q5_K_M with llama.cpp