§01·index · /recipes
Recipes
1009 community-tested setups for running open-weights AI models on real consumer GPUs.page 1 of 11
- imageadvanced24GB+
Qwen-Image-2.1 on RX 7900 XTX: int8 ComfyUI on ROCm 7.2.4, resident in 24 GB, output check
- imageintermediate32GB+
Qwen-Image-2.1 on Apple M4 Max 48 GB: 8-bit mflux, Third-Party 64 GB Timings, ComfyUI Edits
- imageintermediate32GB+
Qwen-Image-2.1 on Apple M3 Max 48 GB: 8-bit mflux Inside the 36 GiB Budget, Edits in ComfyUI
- imageintermediate64GB+
Qwen-Image-2.1 on Apple M2 Max 64 GB: Full bf16 in mflux, --low-ram for 2K, ComfyUI Edits
- imageintermediate8GB+
Qwen-Image-2.1 on RTX 4060 Ti 8GB: What Fits, What Streams per Step, Which 4-Bit DiTs Run
- imageintermediate8GB+
Qwen-Image-2.1 on RTX 4060: 8 GB ComfyUI Install, What Crosses the x8 Link, Ada 4-Bit DiTs
- imageintermediate8GB+
Qwen-Image-2.1 on RTX 3060 Ti 8 GB: what streams per OS, and 4-bit DiTs Ampere runs natively
- imageintermediate12GB+
Qwen-Image-2.1 on RTX 5090: what stays resident in 32 GB, int8 vs bf16, and third-party timings
- imageintermediate12GB+
Qwen-Image-2.1 on RTX 3090 Ti: int8 ComfyUI setup, what stays resident in 24 GB, 2K and edits
- imageintermediate12GB+
Qwen-Image-2.1 on RTX 3090: int8 ComfyUI setup, 24 GB fit per OS, third-party 3090 timings
- imageintermediate12GB+
Qwen-Image-2.1 on RTX 5070: int8 template in 12 GB, 2K, editing, and what native NVFP4 changes
- imageintermediate12GB+
Qwen-Image-2.1 on RTX 4070 Ti: int8 ComfyUI in 12 GB, with third-party 4 MP and edit timings
- imageintermediate12GB+
Qwen-Image-2.1 on the 12 GB RTX 4070 SUPER: what fits, int8 ComfyUI setup, 2K and editing
- imageintermediate12GB+
Qwen-Image-2.1 on any RTX 4070 board: int8 ComfyUI install in 12 GB, 2K and image editing
- imageintermediate12GB+
Qwen-Image-2.1 on RTX 3080 Ti: int8 template, 12 GB budget from owner logs, 2K and editing
- imageintermediate12GB+
Qwen-Image-2.1 on RTX 4080 SUPER: int8 ComfyUI template, 2K and editing, same fit as the 4080
- imageintermediate12GB+
Qwen-Image-2.1 on RTX 4080: ComfyUI int8 template in 16 GB, encoder and DiT take turns
- imageintermediate12GB+
Qwen-Image-2.1 on RTX 4070 Ti SUPER, the 16 GB Ti: ComfyUI int8 template, 2K and editing
- imageintermediate12GB+
Qwen-Image-2.1 on RTX 5080: the int8 ComfyUI template and the edit timings from a ComfyUI PR
- imageintermediate12GB+
Qwen-Image-2.1 on RTX 5070 Ti: int8 template, 16 GB budget and an edit-grid sweep by one user
- imageintermediate12GB+
Qwen-Image-2.1 on RTX 4060 Ti 16GB: the x8 PCIe link, per prompt change and per edit step
- imageintermediate12GB+
Qwen-Image-2.1 on RTX 5060 Ti: the int8 ComfyUI template, 2K output and editing in 16 GB
- imageintermediate12GB+
Qwen-Image-2.1 on RTX 4090: all three models resident at 1 MP in ComfyUI
- imageintermediate12GB+
Qwen-Image-2.1 on RTX 3060: int8 text-to-image and editing in ComfyUI (Ampere sm_86, 12 GB)
- imageintermediate8GB+
Qwen-Image-2.1 on RTX 5060 8 GB: Which Stages Fit, What Streams, and the NVFP4 Swap
- musicintermediate16GB+
YuE2-3B on RTX 4060 Ti 16GB: the fit transfers, the timings do not
- musicintermediate16GB+
YuE2-3B on RTX 4070 Ti Super: the 16GB fit is arithmetic, not luck
- musicintermediate16GB+
YuE2-3B on RTX 4080 Super: full-length songs inside the 16GB clamp
- musicintermediate16GB+
YuE2-3B on RTX 4080: the measured 16GB fit, carried across to Ada
- musicintermediate16GB+
YuE2-3B on RTX 5070 Ti: same pins as the measured card, twice the bus
- musicintermediate16GB+
YuE2-3B on RTX 5080: a measured 16GB fit, carried from a slower Blackwell card
- musicintermediate16GB+
YuE2-3B on RTX 5060 Ti: 16GB measured, worst case included
- musicintermediate16GB+
YuE2-3B on RTX 5090: the 32GB card the library caps at 22 GiB
- musicintermediate16GB+
YuE2-3B on RTX 4090: full songs with an editable score
- musicintermediate16GB+
YuE2-3B on RTX 3090: full songs on Ampere without the FP8 path
- musicintermediate16GB+
YuE2-3B on RTX 3090 Ti: full songs, and the host RAM it needs
- llmintermediate24GB+
Apodex 1.1 mini on RTX 3090: 128K-Context Agent Model in Q4_K_M GGUF
- llmadvanced12GB+
Apodex 1.1 mini on RTX 5070: a 35B-A3B agent in 12 GB via expert offload on Blackwell
- llmadvanced12GB+
Apodex 1.1 mini on RTX 4070 Ti: a 35B-A3B agent in 12 GB, and what Ada can and cannot buy
- llmadvanced12GB+
Apodex 1.1 mini on RTX 4070: a 35B-A3B agent in 12 GB via expert offload on Ada
- llmadvanced12GB+
Apodex 1.1 mini on RTX 4070 SUPER: 35B-A3B in 12 GB, and why extra shaders do not help
- llmadvanced12GB+
Apodex 1.1 mini on RTX 3080 Ti: a 35B-A3B agent in 12 GB, where the card is not the bottleneck
- llmadvanced12GB+
Apodex 1.1 mini on RTX 3060: a 35B-A3B agent model on 12GB via expert offload
- llmadvanced16GB+
Apodex 1.1 mini on RX 7800 XT: a 36B agent in 16 GB with ROCm and llama.cpp expert offload
- llmadvanced16GB+
Apodex 1.1 mini on RTX 5060 Ti: a 36B agent at 128K, and what GDDR7 buys on a 128-bit bus
- llmadvanced16GB+
Apodex 1.1 mini on RTX 5080: a 36B agent at 128K on Blackwell, where your DIMMs decide the pace
- llmadvanced16GB+
Apodex 1.1 mini on RTX 5070 Ti: a 36B agent at 128K, and where DDR5 flips the bottleneck
- llmadvanced16GB+
Apodex 1.1 mini on RTX 4060 Ti 16GB: a 36B agent at 128K, where the card is the slower half
- llmadvanced16GB+
Apodex 1.1 mini on RTX 4070 Ti Super: a 36B agent in 16 GB via llama.cpp expert offload
- llmadvanced16GB+
Apodex 1.1 mini on RTX 4080 SUPER: a 36B agent in 16 GB, and what the SUPER badge buys
- llmadvanced16GB+
Apodex 1.1 mini on RTX 4080: a 36B agent in 16 GB via llama.cpp expert offload
- llmadvanced48GB+
Apodex 1.1 mini on Apple M3 Max: 4-bit MLX at 131K, and why 48 GB stops short of 262K
- llmadvanced64GB+
Apodex 1.1 mini on Apple M2 Max: 6-bit MLX at 131K, and why 262K is a different tier
- llmintermediate24GB+
Apodex 1.1 mini on RTX 4090: 128K Agent Context in Q4_K_M, and the 24 GB Ceiling
- llmintermediate24GB+
Apodex 1.1 mini on RTX 3090 Ti: 128K Agent Context in Q4_K_M, and the CUDA-Graph Caveat
- llmadvanced24GB+
Apodex 1.1 mini on RX 7900 XTX: a 128K-Context Agent Server on llama.cpp-HIP
- llmadvanced36GB+
Apodex 1.1 mini on Apple M4 Max: a 262K-context agent server on MLX
- llmintermediate32GB+
Apodex 1.1 mini on RTX 5090: Q5_K_M Agentic LLM at the Full 262K Context
- llmadvanced12GB+
Qwen3.6-35B-A3B on RTX 3060: 38.9 tok/s from a 12GB card via MoE expert CPU-offload
- multimodaladvanced64GB+
Muse Glimmer 30B on Apple M2 Max: 8-bit MLX with vision and the DFlash drafter
- multimodaladvanced36GB+
Muse Glimmer 30B on Apple M3 Max: ExecuTorch Metal agent server with vision and DFlash
- multimodalintermediate24GB+
Muse Glimmer 30B on RTX 4090: vision and DFlash in llama.cpp at full 131K context
- multimodalintermediate24GB+
Muse Glimmer 30B on RTX 3090 Ti: vision + DFlash speculation at full 131K context
- multimodalintermediate48GB+
Qwen3.8-27B on Apple M4 Max: 4-bit MLX Vision-Language at the Full 262K Context
- multimodalintermediate48GB+
Qwen3.8-27B on Apple M3 Max: 4-bit MLX Vision-Language at the Full 262K Context
- multimodaladvanced16GB+
Qwen3.8-27B on RX 7800 XT: a vision-capable 27B inside 16 GB on ROCm
- multimodalintermediate24GB+
Qwen3.8-27B on RX 7900 XTX: 128K-context vision chat on ROCm with llama.cpp
- multimodaladvanced16GB+
Qwen3.8-27B on RTX 5060 Ti: a vision-capable 27B in 16 GB on a 128-bit bus
- multimodaladvanced16GB+
Qwen3.8-27B on RTX 4060 Ti 16GB: a vision-capable 27B on 288 GB/s
- multimodaladvanced16GB+
Qwen3.8-27B on RTX 4070 Ti Super: a vision-capable 27B inside 16 GB with llama.cpp
- multimodaladvanced16GB+
Qwen3.8-27B on RTX 5080: a vision-capable 27B inside 16 GB with llama.cpp
- multimodaladvanced16GB+
Qwen3.8-27B on RTX 4080 Super: a vision-capable 27B inside 16 GB with llama.cpp
- multimodaladvanced16GB+
Qwen3.8-27B on RTX 4080: a vision-capable 27B inside 16 GB with llama.cpp
- multimodalintermediate24GB+
Qwen3.8-27B on RTX 4090: 128K-context vision chat via llama.cpp Q4_K_M
- multimodalintermediate24GB+
Qwen3.8-27B on RTX 3090 Ti: 128K-context vision chat via llama.cpp Q4_K_M
- multimodalintermediate48GB+
Qwen3.8-27B on Apple M2 Max: 8-bit MLX Vision-Language with MTP Speculative Decoding
- multimodaladvanced16GB+
Qwen3.8-27B on RTX 5070 Ti: a vision-capable 27B inside 16 GB with llama.cpp
- multimodalintermediate32GB+
Qwen3.8-27B on RTX 5090: the full 262K context window on one card
- multimodalintermediate24GB+
Qwen3.8-27B on RTX 3090: 128K-context vision chat via llama.cpp GGUF
- videoadvanced16GB+
LTX-2.5 on RTX 4070 Ti SUPER: 22B audio-video in 16 GB the plain Ti does not have
- videoadvanced16GB+
LTX-2.5 on RTX 4080 SUPER: 22B audio-video in 16 GB, and why the SUPER badge changes nothing
- videoadvanced16GB+
LTX-2.5 on RTX 4080: 22B audio-video in 16 GB via a community Q3_K_M GGUF on Ada
- videoadvanced16GB+
LTX-2.5 on RTX 5070 Ti: 22B audio-video in 16 GB, and the decode trap three owners hit
- videoadvanced16GB+
LTX-2.5 on RTX 5080: 22B audio-video in 16 GB, and what the wider bus cannot buy
- videoadvanced16GB+
LTX-2.5 on RTX 4060 Ti 16GB: 22B audio-video in 16 GB via a community Q3_K_M GGUF
- videoadvanced16GB+
LTX-2.5 on RTX 5060 Ti: 22B audio-video in 16 GB via a community Q3_K_M GGUF
- multimodaladvanced36GB+
Muse Glimmer 30B on Apple M4 Max: ExecuTorch Metal agent server with vision and DFlash
- multimodaladvanced24GB+
Muse Glimmer 30B on RX 7900 XTX: ROCm llama.cpp with vision and DFlash speculative decoding
- multimodalintermediate24GB+
Muse Glimmer 30B on RTX 3090: vision + DFlash speculation at full 131K context
- multimodalintermediate32GB+
Muse Glimmer 30B on RTX 5090: K-Quant-Dynamic, vision and DFlash in llama.cpp
- videoadvanced24GB+
LTX-2.5 on RTX 3090 Ti: 22B audio-video on Ampere via the official ComfyUI INT8 checkpoint
- videoadvanced24GB+
LTX-2.5 on RTX 3090: 22B audio-video on Ampere via the official ComfyUI INT8 checkpoint
- videoadvanced24GB+
LTX-2.5 on RTX 4090: Fitting the 22B Audio-Video DiT in 24 GB via ComfyUI int8
- videoadvanced24GB+
LTX-2.5 on RTX 5090: 22B Audio-Video in 32 GB via ComfyUI's INT8 Path
- llmintermediate36GB+
Nanbeige4.2-3B on Apple M4 Max: 546 GB/s for a stack that streams twice
- llmintermediate36GB+
Nanbeige4.2-3B on Apple M3 Max: the full 262,144-token context in 36 GiB
- llmintermediate16GB+
Nanbeige4.2-3B on RX 7800 XT: 96K context on a 16 GB gfx1101 ROCm build
- llmintermediate24GB+
Nanbeige4.2-3B on RTX 4090: the full 256K context on an sm_89 llama.cpp build
- llmintermediate24GB+
Nanbeige4.2-3B on RTX 3090 Ti: the full 256K context in VRAM with llama.cpp
- llmintermediate16GB+
Nanbeige4.2-3B on RTX 5080: the fastest 16 GB card stops at the same context