§01·index · /recipes
Recipes
1009 community-tested setups for running open-weights AI models on real consumer GPUs.page 2 of 11
- llmintermediate16GB+
Nanbeige4.2-3B on RTX 5070 Ti: 96K context, and what twice the bandwidth cannot buy
- llmintermediate16GB+
Nanbeige4.2-3B on RTX 5060 Ti: the whole 16 GB context ladder on a 128-bit bus
- llmintermediate16GB+
Nanbeige4.2-3B on RTX 4080 SUPER: 16 GB Is the Limit, Not the Silicon
- llmintermediate16GB+
Nanbeige4.2-3B on RTX 4080: 96K Context, and What Each Token Costs in Bytes
- llmintermediate16GB+
Nanbeige4.2-3B on RTX 4070 Ti SUPER: 96K Context on 16 GB, One Rung at a Time
- llmintermediate12GB+
Nanbeige4.2-3B on RTX 5070: Q8_0 weights, 65,536-token cache, CUDA 12.8 floor
- llmintermediate12GB+
Nanbeige4.2-3B on RTX 4070 Ti: Q8_0 weights and a 65,536-token cache in 12 GB
- llmintermediate12GB+
Nanbeige4.2-3B on RTX 4070 SUPER: Q8_0 weights, 65,536-token cache in 12 GB
- llmintermediate12GB+
Nanbeige4.2-3B on RTX 4070: Q8_0 weights and a 65,536-token cache in 12 GB
- llmintermediate12GB+
Nanbeige4.2-3B on RTX 3080 Ti: Q8_0 weights and a 65,536-token cache in 12 GB
- llmintermediate8GB+
Nanbeige4.2-3B on RTX 5060: the One 8 GB Card With a Real 65,536-Token Run Behind It
- llmintermediate8GB+
Nanbeige4.2-3B on RTX 4060 Ti 8GB: an 8 GB Budget for 65,536 New Tokens
- llmintermediate8GB+
Nanbeige4.2-3B on RTX 3060 Ti: 32K Context in 8 GB, and What sm_86 Does Not Change
- llmintermediate16GB+
Nanbeige4.2-3B on Apple M2 Pro: the context ladder for 16 GB of unified memory
- llmintermediate24GB+
Nanbeige4.2-3B on RX 7900 XTX: the full 256K context on a ROCm llama.cpp build
- llmintermediate32GB+
Nanbeige4.2-3B on RTX 5090: 256K context with a near-lossless q8_0 KV cache
- llmintermediate16GB+
Nanbeige4.2-3B on RTX 4060 Ti 16GB: 96K Context Without a 4-bit KV Cache
- llmintermediate12GB+
Nanbeige4.2-3B on RTX 3060: Q8_0 weights and a 65,536-token cache in 12 GB
- llmintermediate16GB+
Nanbeige4.2-3B on Apple M2 Max: 128K-context agentic LLM via llama.cpp Metal
- llmintermediate24GB+
Nanbeige4.2-3B on RTX 3090: the full 256K context in VRAM with llama.cpp
- llmintermediate8GB+
Nanbeige4.2-3B on RTX 4060: 32K Context on 8 GB Despite the Looped-Transformer KV Tax
- videoadvanced12GB+
MiniMax H3 on RTX 5080: 16 GB, and what the extra silicon cannot buy
- videoadvanced12GB+
MiniMax H3 on RTX 5070 Ti: 16 GB, where only one big module fits
- videoadvanced12GB+
MiniMax H3 on RTX 5070: 12 GB, where nvfp4 is native and unused
- videoadvanced12GB+
MiniMax H3 on RTX 5060 Ti: 16 GB, two published runs, and the gap between them
- videoadvanced12GB+
MiniMax H3 on RTX 4090: 24 GB video+audio, and where Ada's FP8 shows up
- videoadvanced12GB+
MiniMax H3 on RTX 4080 SUPER: 16 GB, and the one 4080 number that is nearly yours
- videoadvanced12GB+
MiniMax H3 on RTX 4080: 16 GB Ada, and the only number this card has
- videoadvanced12GB+
MiniMax H3 on RTX 4070 Ti SUPER: Ada at 16 GB, and the Turbo LoRA question
- videoadvanced12GB+
MiniMax H3 on RTX 4070 Ti: 12 GB Ada, and the 16 GB card wearing your card's name
- videoadvanced12GB+
MiniMax H3 on RTX 4070 SUPER: 12 GB, where the host machine decides the run
- videoadvanced12GB+
MiniMax H3 on RTX 4060 Ti 16GB: the capacity fits, the bus is the question
- videoadvanced12GB+
MiniMax H3 on RTX 3090 Ti: 24 GB video+audio, and what sm_86 takes off the table
- videoadvanced12GB+
MiniMax H3 on RTX 3080 Ti: 12 GB video+audio in ComfyUI on Ampere sm_86
- videoadvanced12GB+
MiniMax H3 on RTX 3060: 12 GB video+audio in ComfyUI, measured on this card
- videoadvanced16GB+
MiniMax H3 on RX 7800 XT: 16 GB video+audio in ComfyUI on ROCm (gfx1101)
- videoadvanced12GB+
MiniMax H3 on RTX 5090: 32 GB video+audio, and what nvfp4 does not buy you
- videoadvanced64GB+
MiniMax H3 on Apple M2 Max: native MLX video with synchronised audio
- videoadvanced12GB+
MiniMax H3 on RTX 4070: 12 GB video+audio via ComfyUI VRAM offload
- videoadvanced12GB+
MiniMax H3 on RTX 3090: 5-second audio+video in ComfyUI on pruned int8
- multimodaladvanced48GB+
Fara1.5-27B on Apple M4 Max: Browser Computer-Use Agent on llama.cpp Metal
- multimodaladvanced24GB+
Fara1.5-9B on RX 7900 XTX: Unquantised Browser Computer-Use Agent on ROCm
- multimodaladvanced16GB+
Fara1.5-9B on RX 7800 XT: a Q8_0 Browser Computer-Use Agent on ROCm gfx1101
- multimodaladvanced24GB+
Fara1.5-9B on RTX 5090: BF16 Computer-Use Agent on Blackwell sm_120
- multimodaladvanced16GB+
Fara1.5-9B on RTX 5080: a Q8_0 Computer-Use Agent on Blackwell sm_120
- multimodaladvanced12GB+
Fara1.5-9B on RTX 5070: a Blackwell Browser Computer-Use Agent with llama.cpp
- multimodaladvanced16GB+
Fara1.5-9B on RTX 5070 Ti: a Q8_0 Computer-Use Agent on Blackwell sm_120
- multimodaladvanced16GB+
Fara1.5-9B on RTX 5060 Ti 16GB: a Q8_0 Computer-Use Agent on Blackwell
- multimodaladvanced24GB+
Fara1.5-9B on RTX 4090: BF16 Computer-Use Agent with Vision on Ada sm_89
- multimodaladvanced16GB+
Fara1.5-9B on RTX 4080: a Q8_0 Browser Computer-Use Agent with llama.cpp
- multimodaladvanced16GB+
Fara1.5-9B on RTX 4080 Super: a Q8_0 Browser Computer-Use Agent with llama.cpp
- multimodaladvanced12GB+
Fara1.5-9B on RTX 4070 Ti: a Q5_K_M Browser Computer-Use Agent with llama.cpp
- multimodaladvanced16GB+
Fara1.5-9B on RTX 4070 Ti Super: a Q8_0 Browser Computer-Use Agent with llama.cpp
- multimodaladvanced12GB+
Fara1.5-9B on RTX 4070 Super: a Local Browser Computer-Use Agent in 12GB
- multimodaladvanced16GB+
Fara1.5-9B on RTX 4060 Ti 16GB: a Q8_0 Browser Computer-Use Agent with llama.cpp
- multimodaladvanced24GB+
Fara1.5-9B on RTX 3090: BF16 Computer-Use Agent with Vision under llama.cpp
- multimodaladvanced24GB+
Fara1.5-9B on RTX 3090 Ti: BF16 Browser Computer-Use Agent on Ampere sm_86
- multimodaladvanced12GB+
Fara1.5-9B on RTX 3080 Ti: a Browser Computer-Use Agent at Q5_K_M with llama.cpp
- multimodaladvanced12GB+
Fara1.5-9B on RTX 3060: a 12GB Browser Computer-Use Agent at Q5_K_M with llama.cpp
- multimodaladvanced32GB+
Fara1.5-9B on Apple M4 Max: Browser Computer-Use Agent on llama.cpp Metal
- multimodaladvanced32GB+
Fara1.5-9B on Apple M3 Max: Browser Computer-Use Agent on llama.cpp Metal
- multimodaladvanced16GB+
Fara1.5-9B on Apple M2 Pro: Browser Computer-Use Agent in 16GB Unified Memory
- multimodaladvanced32GB+
Fara1.5-9B on Apple M2 Max: Browser Computer-Use Agent on llama.cpp Metal
- multimodaladvanced16GB+
Fara1.5-4B on RX 7900 XTX: Unquantised Browser Computer-Use Agent on ROCm
- multimodaladvanced16GB+
Fara1.5-4B on RX 7800 XT: Unquantised Browser Computer-Use Agent on ROCm gfx1101
- multimodaladvanced24GB+
Fara1.5-4B on RTX 5090: BF16 Computer-Use Agent on Blackwell sm_120 with llama.cpp
- multimodaladvanced16GB+
Fara1.5-4B on RTX 5080: bf16 Weights and mmproj Vision, No Quantisation Step
- multimodaladvanced12GB+
Fara1.5-4B on RTX 5070: Blackwell Browser Computer-Use Agent at Q8_0 with llama.cpp
- multimodaladvanced16GB+
Fara1.5-4B on RTX 5070 Ti: bf16 Computer-Use Agent with llama.cpp on sm_120
- multimodaladvanced8GB+
Fara1.5-4B on RTX 5060: Blackwell Browser Computer-Use Agent via llama.cpp
- multimodaladvanced16GB+
Fara1.5-4B on RTX 5060 Ti: Unquantised bf16 Browser Agent on 16GB Blackwell
- multimodaladvanced16GB+
Fara1.5-4B on RTX 4090: BF16 Computer-Use Agent with a 131K Window on llama.cpp
- multimodaladvanced16GB+
Fara1.5-4B on RTX 4080: the Vendor's BF16 Precision for a Computer-Use Agent
- multimodaladvanced16GB+
Fara1.5-4B on RTX 4080 Super: BF16 Browser Computer-Use Agent under llama-server
- multimodaladvanced12GB+
Fara1.5-4B on RTX 4070: a Q8_0 Browser Computer-Use Agent with llama.cpp
- multimodaladvanced12GB+
Fara1.5-4B on RTX 4070 Ti: a Near-Lossless Local Browser Agent with llama.cpp
- multimodaladvanced16GB+
Fara1.5-4B on RTX 4070 Ti Super: a BF16 Browser Computer-Use Agent with llama.cpp
- multimodaladvanced12GB+
Fara1.5-4B on RTX 4070 Super: a 65K-Context Computer-Use Agent in 12GB
- multimodaladvanced8GB+
Fara1.5-4B on RTX 4060: an 8GB Browser Computer-Use Agent via llama.cpp Vision
- multimodaladvanced8GB+
Fara1.5-4B on RTX 4060 Ti 8GB: Local Browser Computer-Use Agent with llama.cpp
- multimodaladvanced16GB+
Fara1.5-4B on RTX 4060 Ti 16GB: BF16 Computer-Use Agent with llama.cpp
- multimodaladvanced16GB+
Fara1.5-4B on RTX 3090: BF16 Computer-Use Agent at a 131K Context with llama.cpp
- multimodaladvanced16GB+
Fara1.5-4B on RTX 3090 Ti: BF16 Browser Computer-Use Agent on Ampere sm_86
- multimodaladvanced12GB+
Fara1.5-4B on RTX 3080 Ti: Browser Computer-Use Agent at Q8_0 in 12GB with llama.cpp
- multimodaladvanced12GB+
Fara1.5-4B on RTX 3060: a 12GB Browser Computer-Use Agent at Q8_0 with llama.cpp
- multimodaladvanced24GB+
Fara1.5-4B on Apple M4 Max: Browser Computer-Use Agent at Full 262K Context
- multimodaladvanced24GB+
Fara1.5-4B on Apple M3 Max: Browser Computer-Use Agent on llama.cpp Metal
- multimodaladvanced16GB+
Fara1.5-4B on Apple M2 Pro: Browser Computer-Use Agent in 16GB Unified Memory
- multimodaladvanced24GB+
Fara1.5-4B on Apple M2 Max: Browser Computer-Use Agent on llama.cpp Metal
- multimodaladvanced24GB+
Fara1.5-27B on RX 7900 XTX: Q4_K_M Browser Computer-Use Agent on ROCm gfx1100
- multimodaladvanced32GB+
Fara1.5-27B on RTX 5090: Q6_K Computer-Use Agent on Blackwell sm_120
- multimodaladvanced24GB+
Fara1.5-27B on RTX 4090: Q4_K_M Browser Computer-Use Agent on Ada sm_89
- multimodaladvanced24GB+
Fara1.5-27B on RTX 3090 Ti: Q4_K_M Browser Computer-Use Agent on Ampere sm_86
- multimodaladvanced48GB+
Fara1.5-27B on Apple M3 Max: Browser Computer-Use Agent on llama.cpp Metal
- multimodaladvanced64GB+
Fara1.5-27B on Apple M2 Max: Browser Computer-Use Agent at the Full 262K Context
- multimodaladvanced24GB+
Fara1.5-27B on RTX 3090: local browser computer-use agent with llama.cpp
- multimodaladvanced12GB+
Fara1.5-9B on RTX 4070: a Local Browser Computer-Use Agent with llama.cpp
- multimodaladvanced8GB+
Fara1.5-4B on RTX 3060 Ti: Browser Computer-Use Agent with llama.cpp Vision
- multimodaladvanced24GB+
Agents-A1 35B-A3B on RTX 3090: 128K Agentic Serving via llama.cpp, and What Vision Costs on 24 GB
- multimodalintermediate14GB+
Agents-A1 4B on Apple M2 Max: First-Party GGUF + Vision on Metal at the Full 262K Context