self-hosted/ai
§01·model · /models

Llama 3.2 1B

llmactiveLlama 3.2 Community License

1B instruction-tuned LLM by Meta, sized for edge and on-device use — small enough that it is a realistic CPU or laptop target, not just a GPU one. The single recipe here happens to run it on a 24 GB Radeon RX 7900 XTX under ROCm, recorded at 106 tokens/s, which is far more card than a 1B needs. Llama 3.2 Community License.

§02·GPUs that run this model
1 total

benchmarked·~ runs via recipe (not benchmarked)· untested·doesn't fit