§01·spec · /gpus
Apple M3 Max
appleapple series48GB unified
25 open-weights AI models run on the Apple M3 Max — see which fit, how fast they go, and the VRAM each needs.
Apple Silicon uses unified memory shared between CPU and GPU. Of this machine’s 48 GB the GPU addresses 36.0 GiB by default — exactly three quarters — and the limit is raisable via Metal’s wired-memory setting.
§02·models that run on this GPU
25 totalLLM · 14
| Model | Best speed | Min VRAM | Works | Evidence | |
|---|---|---|---|---|---|
| Devstral Small 2 (24B) | 48GB | ✓ | recipe | check ↗ | |
| Gemma 4 12B | 48GB | ✓ | recipe | check ↗ | |
| gpt-oss 20B | 13GB | ✓ | recipe | check ↗ | |
| Laguna XS 2.1 | 24GB | ✓ | recipe | check ↗ | |
| Llama 3.1 8B | 5GB | ✓ | recipe | check ↗ | |
| Llama 3.3 70B | 40GB | ✓ | recipe | check ↗ | |
| Mistral Nemo 12B | 48GB | ✓ | recipe | check ↗ | |
| Mistral Small 3.2 24B | 48GB | ✓ | recipe | check ↗ | |
| Nanbeige4.2 3B | 36GB | ✓ | recipe | check ↗ | |
| North Mini Code 1.0 | 48GB | ✓ | recipe | check ↗ | |
| Ornith 1.0 35B | 48GB | ✓ | recipe | check ↗ | |
| Phi-4 | 48GB | ✓ | recipe | check ↗ | |
| Qwen3 32B | 19GB | ✓ | recipe | check ↗ | |
| Qwen3-Next 80B-A3B | 48GB | ✓ | recipe | check ↗ |
- Best speed
- Min VRAM
- 48GB
- Evidence
- recipe
- Best speed
- Min VRAM
- 48GB
- Evidence
- recipe
- Best speed
- Min VRAM
- 13GB
- Evidence
- recipe
- Best speed
- Min VRAM
- 24GB
- Evidence
- recipe
- Best speed
- Min VRAM
- 5GB
- Evidence
- recipe
- Best speed
- Min VRAM
- 40GB
- Evidence
- recipe
- Best speed
- Min VRAM
- 48GB
- Evidence
- recipe
- Best speed
- Min VRAM
- 48GB
- Evidence
- recipe
- Best speed
- Min VRAM
- 36GB
- Evidence
- recipe
- Best speed
- Min VRAM
- 48GB
- Evidence
- recipe
- Best speed
- Min VRAM
- 48GB
- Evidence
- recipe
- Best speed
- Min VRAM
- 48GB
- Evidence
- recipe
- Best speed
- Min VRAM
- 19GB
- Evidence
- recipe
- Best speed
- Min VRAM
- 48GB
- Evidence
- recipe
Multimodal · 4
| Model | Best speed | Min VRAM | Works | Evidence | |
|---|---|---|---|---|---|
| Fara1.5-27B | 48GB | ✓ | recipe | check ↗ | |
| Fara1.5-4B | 24GB | ✓ | recipe | check ↗ | |
| Fara1.5-9B | 32GB | ✓ | recipe | check ↗ | |
| Gemma 4 E4B-IT | 5GB | ✓ | recipe | check ↗ |
Image · 3
| Model | Best speed | Min VRAM | Works | Evidence | |
|---|---|---|---|---|---|
| Krea 2 | 24GB | ✓ | recipe | check ↗ | |
| Qwen-Image | 23GB | ✓ | recipe | check ↗ | |
| Z-Image Turbo | 17GB | ✓ | recipe | check ↗ |
Video · 1
3D · 1
| Model | Best speed | Min VRAM | Works | Evidence | |
|---|---|---|---|---|---|
| TRELLIS.2-4B | 17GB | ✓ | recipe | check ↗ |
§03·tested recipes
showing 6 of 25- llmintermediate36GB+recipe
Nanbeige4.2-3B on Apple M3 Max: the full 262,144-token context in 36 GiB
- multimodaladvanced32GB+recipe
Fara1.5-9B on Apple M3 Max: Browser Computer-Use Agent on llama.cpp Metal
- multimodaladvanced24GB+recipe
Fara1.5-4B on Apple M3 Max: Browser Computer-Use Agent on llama.cpp Metal
- multimodaladvanced48GB+recipe
Fara1.5-27B on Apple M3 Max: Browser Computer-Use Agent on llama.cpp Metal
- llmadvanced48GB+recipe
Qwen3-Next 80B-A3B on Apple M3 Max: an 80B MoE Assistant in 48GB via a Sub-Q4 GGUF
- llmintermediate48GB+recipe
Gemma 4 12B on Apple M3 Max: Local Private Assistant via llama.cpp / Ollama (48GB)