self-hosted/ai
§01·model · /models

Agents-A1 4B

multimodalactiveApache-2.0

The dense 4B member of InternScience's Agents-A1 agentic line, released 2026-07-14 as the local-assistant tier. Like its 35B MoE sibling it is a vision-language model - 297 of its 723 tensors are visual - and InternScience ships a first-party mmproj that mainline llama.cpp loads through the qwen3vl_merger projector. Apache-2.0, 262K context, with first-party Q4_K_M, Q8_0, F16 GGUF and FP8 builds.

At Q4_K_M the language weights are 2.52 GiB plus a 0.63 GiB projector, so it fits every consumer GPU this site covers, 8 GB cards included. The card documents vLLM and SGLang only; the vendor does not state the base model.

Download· 3 variants
§02·same family
1 other model

Other models grouped with Agents-A1 4B. Sizes and modalities can differ, and nothing here says whether one fits your GPU — open a model for its own compatibility table.

§03·GPUs that run this model
2 total

benchmarked·~ runs via recipe (not benchmarked)· untested·doesn't fit