§01·spec · /gpus
Apple M3 Max
appleapple series48GB unified
27 open-weights AI models run on the Apple M3 Max — see which fit, how fast they go, and the VRAM each needs.
Apple Silicon uses unified memory shared between CPU and GPU. Of this machine’s 48 GB the GPU addresses 36.0 GiB by default — exactly three quarters — and the limit is raisable via Metal’s wired-memory setting.
§02·models that run on this GPU
27 totalLLM · 14
| Model | Best speed | Min VRAM | Works | Evidence | |
|---|---|---|---|---|---|
| Devstral Small 2 (24B) | 48GB | ✓ | recipe | check ↗ | |
| Gemma 4 12B | 48GB | ✓ | recipe | check ↗ | |
| gpt-oss 20B | 13GB | ✓ | recipe | check ↗ | |
| Laguna XS 2.1 | 24GB | ✓ | recipe | check ↗ | |
| Llama 3.1 8B | 5GB | ✓ | recipe | check ↗ | |
| Llama 3.3 70B | 40GB | ✓ | recipe | check ↗ | |
| Mistral Nemo 12B | 48GB | ✓ | recipe | check ↗ | |
| Mistral Small 3.2 24B | 48GB | ✓ | recipe | check ↗ | |
| Nanbeige4.2 3B | 36GB | ✓ | recipe | check ↗ | |
| North Mini Code 1.0 | 48GB | ✓ | recipe | check ↗ | |
| Ornith 1.0 35B | 48GB | ✓ | recipe | check ↗ | |
| Phi-4 | 48GB | ✓ | recipe | check ↗ | |
| Qwen3 32B | 19GB | ✓ | recipe | check ↗ | |
| Qwen3-Next 80B-A3B | 48GB | ✓ | recipe | check ↗ |
- Best speed
- Min VRAM
- 48GB
- Evidence
- recipe
- Best speed
- Min VRAM
- 48GB
- Evidence
- recipe
- Best speed
- Min VRAM
- 13GB
- Evidence
- recipe
- Best speed
- Min VRAM
- 24GB
- Evidence
- recipe
- Best speed
- Min VRAM
- 5GB
- Evidence
- recipe
- Best speed
- Min VRAM
- 40GB
- Evidence
- recipe
- Best speed
- Min VRAM
- 48GB
- Evidence
- recipe
- Best speed
- Min VRAM
- 48GB
- Evidence
- recipe
- Best speed
- Min VRAM
- 36GB
- Evidence
- recipe
- Best speed
- Min VRAM
- 48GB
- Evidence
- recipe
- Best speed
- Min VRAM
- 48GB
- Evidence
- recipe
- Best speed
- Min VRAM
- 48GB
- Evidence
- recipe
- Best speed
- Min VRAM
- 19GB
- Evidence
- recipe
- Best speed
- Min VRAM
- 48GB
- Evidence
- recipe
Multimodal · 6
| Model | Best speed | Min VRAM | Works | Evidence | |
|---|---|---|---|---|---|
| Fara1.5-27B | 48GB | ✓ | recipe | check ↗ | |
| Fara1.5-4B | 24GB | ✓ | recipe | check ↗ | |
| Fara1.5-9B | 32GB | ✓ | recipe | check ↗ | |
| Gemma 4 E4B-IT | 5GB | ✓ | recipe | check ↗ | |
| Muse Glimmer 30B | 36GB | ✓ | recipe | check ↗ | |
| Qwen3.8 27B | 48GB | ✓ | recipe | check ↗ |
Image · 3
| Model | Best speed | Min VRAM | Works | Evidence | |
|---|---|---|---|---|---|
| Krea 2 | 24GB | ✓ | recipe | check ↗ | |
| Qwen-Image | 23GB | ✓ | recipe | check ↗ | |
| Z-Image Turbo | 17GB | ✓ | recipe | check ↗ |
Video · 1
3D · 1
| Model | Best speed | Min VRAM | Works | Evidence | |
|---|---|---|---|---|---|
| TRELLIS.2-4B | 17GB | ✓ | recipe | check ↗ |
§03·tested recipes
showing 6 of 27- multimodaladvanced36GB+recipe
Muse Glimmer 30B on Apple M3 Max: ExecuTorch Metal agent server with vision and DFlash
- multimodalintermediate48GB+recipe
Qwen3.8-27B on Apple M3 Max: 4-bit MLX Vision-Language at the Full 262K Context
- llmintermediate36GB+recipe
Nanbeige4.2-3B on Apple M3 Max: the full 262,144-token context in 36 GiB
- multimodaladvanced32GB+recipe
Fara1.5-9B on Apple M3 Max: Browser Computer-Use Agent on llama.cpp Metal
- multimodaladvanced24GB+recipe
Fara1.5-4B on Apple M3 Max: Browser Computer-Use Agent on llama.cpp Metal
- multimodaladvanced48GB+recipe
Fara1.5-27B on Apple M3 Max: Browser Computer-Use Agent on llama.cpp Metal