self-hosted/ai
§01·model · /models

gpt-oss 20B

llmactiveApache-2.0

20B open-weight LLM by OpenAI (gpt-oss) for local reasoning and agentic use, Apache-2.0. Twenty-three cards here carry a published recipe, down to a 12 GB floor. Measured at 298.2 tok/s generation on an RTX 5090, and 6,364 prefill tok/s on a 16 GB RTX 4080 SUPER: for agent loops the prefill number is the one that decides whether a long tool history feels instant or sluggish, since the prompt is re-read far more often than the answer is written.

§02·GPUs that run this model
23 total

benchmarked·~ runs via recipe (not benchmarked)· untested·doesn't fit