self-hosted/ai
§01·model · /models

Qwen3.5 35B

multimodalactiveApache-2.0

35B Mixture-of-Experts LLM by Alibaba — Qwen3.5-35B-A3B, with only about 3B parameters active per token. That is the entire point of the design, and it shows in the figures recorded here: 111.2 tokens/s on a 3090 and 165.2 on a 5090 at MXFP4, against 34.7 and 61.1 for the dense Gemma4 31B on the very same two cards. Coverage is those two cards, from a 24 GB floor. Apache-2.0.

§02·same family
1 other model

Other models grouped with Qwen3.5 35B. Sizes and modalities can differ, and nothing here says whether one fits your GPU — open a model for its own compatibility table.

§03·GPUs that run this model
2 total

benchmarked·~ runs via recipe (not benchmarked)· untested·doesn't fit