§01·spec · /gpus
Apple M4 Max
appleapple series48GB unified
21 open-weights AI models run on the Apple M4 Max — see which fit, how fast they go, and the VRAM each needs.
Apple Silicon uses unified memory shared between CPU and GPU. Of this machine’s 48 GB the GPU addresses 36.0 GiB by default — exactly three quarters — and the limit is raisable via Metal’s wired-memory setting.
§02·models that run on this GPU
21 totalLLM · 7
| Model | Best speed | Min VRAM | Works | Evidence | |
|---|---|---|---|---|---|
| Apodex 1.1 mini | 36GB | ✓ | recipe | check ↗ | |
| gpt-oss 20B | 13GB | ✓ | recipe | check ↗ | |
| Llama 3.1 8B | 5GB | ✓ | recipe | check ↗ | |
| Llama 3.3 70B | 40GB | ✓ | recipe | check ↗ | |
| Nanbeige4.2 3B | 36GB | ✓ | recipe | check ↗ | |
| Ornith 1.0 35B | 48GB | ✓ | recipe | check ↗ | |
| Qwen3 32B | 19GB | ✓ | recipe | check ↗ |
Multimodal · 6
| Model | Best speed | Min VRAM | Works | Evidence | |
|---|---|---|---|---|---|
| Fara1.5-27B | 48GB | ✓ | recipe | check ↗ | |
| Fara1.5-4B | 24GB | ✓ | recipe | check ↗ | |
| Fara1.5-9B | 32GB | ✓ | recipe | check ↗ | |
| Gemma 4 E4B-IT | 5GB | ✓ | recipe | check ↗ | |
| Muse Glimmer 30B | 36GB | ✓ | recipe | check ↗ | |
| Qwen3.8 27B | 48GB | ✓ | recipe | check ↗ |
Image · 4
| Model | Best speed | Min VRAM | Works | Evidence | |
|---|---|---|---|---|---|
| Krea 2 | 24GB | ✓ | recipe | check ↗ | |
| Qwen-Image | 23GB | ✓ | recipe | check ↗ | |
| Qwen-Image-2.1 | 32GB | ✓ | recipe | check ↗ | |
| Z-Image Turbo | 17GB | ✓ | recipe | check ↗ |
Video · 1
3D · 1
| Model | Best speed | Min VRAM | Works | Evidence | |
|---|---|---|---|---|---|
| TRELLIS.2-4B | 17GB | ✓ | recipe | check ↗ |
§03·tested recipes
showing 6 of 21- imageintermediate32GB+recipe
Qwen-Image-2.1 on Apple M4 Max 48 GB: 8-bit mflux, Third-Party 64 GB Timings, ComfyUI Edits
- llmadvanced36GB+recipe
Apodex 1.1 mini on Apple M4 Max: a 262K-context agent server on MLX
- multimodalintermediate48GB+recipe
Qwen3.8-27B on Apple M4 Max: 4-bit MLX Vision-Language at the Full 262K Context
- multimodaladvanced36GB+recipe
Muse Glimmer 30B on Apple M4 Max: ExecuTorch Metal agent server with vision and DFlash
- llmintermediate36GB+recipe
Nanbeige4.2-3B on Apple M4 Max: 546 GB/s for a stack that streams twice
- multimodaladvanced48GB+recipe
Fara1.5-27B on Apple M4 Max: Browser Computer-Use Agent on llama.cpp Metal