self-hosted/ai
§01·model · /models

KAT-Coder V2.5 Dev

llmactiveapache-2.0

KAT-Coder-V2.5-Dev is Kwaipilot's open-weight release of KAT-Coder V2.5, an agentic-coding Mixture-of-Experts model. The model card states 35B total parameters with 3B activated, built on the Qwen3.6-35B-A3B base, a 262,144-token context window, and English + Chinese support. This release ships language-model weights only and is text-only: the vision components are not included (verified — the safetensors index contains 31,333 tensors and zero vision tensors). Kwaipilot reports SWE-bench Verified 69.40, SWE-bench Multilingual 63.00, SWE-bench Pro 45.96 and Terminal-Bench 2.1 41.02. These are vendor-reported figures and are not independently reproduced on this site. The card documents SGLang (>=0.5.10), vLLM (>=0.19.0), KTransformers and Hugging Face Transformers. There is no first-party GGUF and no Ollama library entry; the local single-GPU path comes from community quantizations (bartowski, mradermacher, mudler, owao) built on llama.cpp b10087. Measured quant sizes: Q4_K_M 19.92 GiB, Q3_K_M 15.11 GiB, Q2_K 11.75 GiB, IQ2_XXS 9.11 GiB — so a 24 GB card runs Q4_K_M, a 16 GB card Q3, and a 12 GB card Q2/IQ2, with context sized to the remaining headroom. Community MLX 4-bit and 6-bit builds cover Apple Silicon. Full weights are 64.56 GiB in bf16 (~34.7B parameters). Apache-2.0, ungated.

§02·GPUs that run this model
2 total
GPUVRAMSeriesBest speedMin VRAMWorksBenchmarksRecipe
Apple M2 Max64GBapple~0recipecheck ↗
RTX 309024GB30~0recipecheck ↗

benchmarked·~ runs via recipe (not benchmarked)· untested·doesn't fit