§01·model · /models
Agents-A1 4B
multimodalactiveApache-2.0
The dense 4B member of InternScience's Agents-A1 agentic line, released 2026-07-14 as the local-assistant tier. Like its 35B MoE sibling it is a vision-language model - 297 of its 723 tensors are visual - and InternScience ships a first-party mmproj that mainline llama.cpp loads through the qwen3vl_merger projector. Apache-2.0, 262K context, with first-party Q4_K_M, Q8_0, F16 GGUF and FP8 builds. At Q4_K_M the language weights are 2.52 GiB plus a 0.63 GiB projector, so it fits every consumer GPU this site covers, 8 GB cards included. The card documents vLLM and SGLang only; the vendor does not state the base model.
Download· 3 variants
§02·GPUs that run this model
2 total| GPU | VRAM | Series | Best speed | Min VRAM | Works | Benchmarks | Recipe | |
|---|---|---|---|---|---|---|---|---|
| Apple M2 Max | 64GB | apple | ~ | 0 | recipe | check ↗ | ||
| RTX 3060 Ti | 8GB | 30 | ~ | 0 | recipe | check ↗ |
✓ benchmarked·~ runs via recipe (not benchmarked)·— untested·✕doesn't fit