self-hosted/ai
§01·model · /models

Agents-A1 35B-A3B

multimodalactiveApache-2.0

A 35B Mixture-of-Experts agentic model from InternScience, trained for long-horizon search, engineering, scientific research, instruction following and tool calling. Apache-2.0, 262K context. The weights carry a vision tower (333 of 31,666 tensors) and InternScience ships a first-party mmproj, so image input works on mainline llama.cpp through the qwen3vl_merger projector - although the model card documents and benchmarks only the text-agent side. First-party GGUF and FP8 builds are published alongside the safetensors. At Q4_K_M the language weights are 19.71 GiB plus a 0.84 GiB projector, putting the practical floor at a 24 GB card; mlx-community publishes 3-bit through bf16 conversions for Apple silicon.

Download· 3 variants
§02·GPUs that run this model
1 total
GPUVRAMSeriesBest speedMin VRAMWorksBenchmarksRecipe
RTX 309024GB30~0recipecheck ↗

benchmarked·~ runs via recipe (not benchmarked)· untested·doesn't fit