self-hosted/ai
§01·model · /models

Agents-A1 4B

multimodalactiveApache-2.0

The dense 4B member of InternScience's Agents-A1 agentic line, released 2026-07-14 as the local-assistant tier. Like its 35B MoE sibling it is a vision-language model - 297 of its 723 tensors are visual - and InternScience ships a first-party mmproj that mainline llama.cpp loads through the qwen3vl_merger projector. Apache-2.0, 262K context, with first-party Q4_K_M, Q8_0, F16 GGUF and FP8 builds. At Q4_K_M the language weights are 2.52 GiB plus a 0.63 GiB projector, so it fits every consumer GPU this site covers, 8 GB cards included. The card documents vLLM and SGLang only; the vendor does not state the base model.

Download· 3 variants
§02·GPUs that run this model
2 total
GPUVRAMSeriesBest speedMin VRAMWorksBenchmarksRecipe
Apple M2 Max64GBapple~0recipecheck ↗
RTX 3060 Ti8GB30~0recipecheck ↗

benchmarked·~ runs via recipe (not benchmarked)· untested·doesn't fit