What You'll Build
A local OpenAI-compatible server running Qwen3.8-27B — Qwen's 27B vision-language model — on a single RTX 3090 Ti, answering both text and image prompts at a 131,072-token context window. The install is llama.cpp with a 4-bit GGUF plus its separate vision projector.
Hardware data: RTX 3090 Ti (24GB VRAM) · ~21.1 GiB derived working set at 128K context · See benchmark data
⚠️ Known issue: a 4-bit K or V cache (
q4_1, or a K type that differs from the V type) silently moves attention to the CPU on default CUDA builds, collapsing prefill by roughly 20×. Useq8_0for both, as the command below does. Details in Troubleshooting.
Requirements
| Component | Minimum | This recipe |
|---|---|---|
| GPU | 24GB VRAM | RTX 3090 Ti (24GB) — derived budget below; no benchmark exists for this pair (/contribute) |
| RAM | 16GB | — |
| Storage | 17.40 GB | 16,464,440,224 B weights + 931,146,432 B projector, per the HF tree API |
| Software | llama.cpp with CUDA; qwen35 + qwen3vl_merger support, shipping since ~b8001 | — |
Why a 27B model leaves room for 128K of context
Qwen3.8-27B is a hybrid: the model card gives its stack as "Hidden Layout: 16 × (3 × (Gated DeltaNet → FFN) → 1 × (Gated Attention → FFN))". Only 16 of the 64 blocks are full-attention layers; the other 48 are Gated DeltaNet linear-attention layers that carry a fixed-size recurrent state instead of a growing KV cache. llama.cpp reproduces exactly that split — src/models/qwen35.cpp marks a layer recurrent when (i + 1) % full_attention_interval != 0, which for full_attention_interval = 4 leaves attention at blocks 3, 7, … 63.
That is the whole reason this fits. With head_count_kv = 4 and key_length = value_length = 256 (read from the GGUF header), one token costs 2 × 4 × 256 × 2 bytes = 4 KiB per attention layer, so 64 KiB per token at f16 across 16 layers — a quarter of what a conventional 64-layer 27B would charge. A community reader reached the same conclusion on the model's discussions tab, noting "KV usage is much lower than a conventional 64-layer full-attention 27B model".
The 48 recurrent layers cost a constant instead. llama-hparams.cpp sizes them at (d_conv − 1) × (d_inner + 2 × n_group × d_state) for the convolution state and d_state × d_inner for the recurrent state — 30,720 + 786,432 elements per layer, held in F32 regardless of --cache-type-k/v, and multiplied by the number of parallel sequences. At --parallel 1 that is 156,893,184 B total, about 0.15 GiB.
Derived budget at -c 131072 with q8_0 K and V:
| Component | Bytes | GiB |
|---|---|---|
Weights, UD-Q4_K_M (unsloth) | 16,464,440,224 | 15.334 |
| Vision projector, BF16 | 931,146,432 | 0.867 |
| KV cache, q8_0, 16 layers @ 131072 | 4,563,402,752 | 4.250 |
| Recurrent state, F32 × 1 sequence | 156,893,184 | 0.146 |
| f16 conversion buffer, K+V, one layer @ 131072 | 536,870,912 | 0.500 |
| Total | 22,652,753,504 | 21.097 |
The last row is the one a quantised cache hides, and it is what decides how far the context can go. llama.cpp reads a quantised KV in place only in its vector FlashAttention kernel, which serves batches of one or two tokens — decoding. Prefill takes the MMA kernel, which converts K and V to f16 first: ggml_cuda_flash_attn_ext_get_alloc_size sets need_f16_K and need_f16_V for BEST_FATTN_KERNEL_MMA_F16 and BEST_FATTN_KERNEL_TILE alike, and the helper in fattn-common.cuh sizes the buffer at ggml_nelements(K)*ggml_type_size(GGML_TYPE_F16) — by the whole window, not the batch, so a smaller -ub will not shrink it. Two things keep it affordable: the graph allocator reuses one block across the 16 attention layers, so the row above is a floor rather than a ceiling, and it is reserved at startup rather than grown mid-session, because llama_context::sched_reserve builds its worst-case prompt-processing graph against memory->init_full(). Compare it against llama.cpp's own estimate with --fit-print on.
That leaves 2.903 GiB of the card's 24 GiB for the rest of the graph, the CUDA context and the vision graph, none of which is derived here. This is a derived envelope from cited file sizes and the runtime's own allocation formulas, not a measured peak — see /check/qwen3-8-27b/rtx-3090-ti for measured data as it lands. Two practical deductions come out of that 2.903 GiB rather than out of the components: cudaDeviceProp::totalGlobalMem sits slightly under the nominal 24,576 MiB, and the display driver reserves more on top if a monitor is attached. Check yours with nvidia-smi --query-gpu=memory.total,memory.used --format=csv before committing to a context size.
Three ways to spend the budget on context, and they do not rank the way the flag names suggest. The same weights with an unquantized f16 KV at 131072 need 24.347 GiB and do not fit; drop to -c 65536 and f16 KV fits at 20.347 GiB. The model's full native window — the card advertises "Context Length: 262,144 natively and extensible up to 1,000,000 tokens." — needs 25.847 GiB at q8_0 and is out of reach. But q4_0 brings it back inside the card: because only 16 layers hold a growing cache, halving the per-token cost very nearly pays for doubling the tokens (262,144 tokens at q4_0 cost 4,831,838,208 B against 4,563,402,752 for 131,072 at q8_0 — 1.059×), and -c 262144 --cache-type-k q4_0 --cache-type-v q4_0 totals 21.847 GiB, leaving 2.153 GiB. q4_0 is the one 4-bit setting with a FlashAttention kernel compiled by default, so it does not hit the trap in Troubleshooting. What it costs is cache precision, which this page does not measure — the arithmetic says the configuration fits, not that it answers as well.
Every filename here carries the UD- prefix for a reason. On 2026-08-19 unsloth deleted the plain Q3_K_M, Q3_K_S, Q4_K_M, Q4_K_S, Q5_K_M, Q5_K_S, Q6_K, IQ4_NL and IQ4_XS builds from this repository and kept only the Dynamic ones, so a command naming one of those now 404s. Re-enumerated against the repository's file tree on 2026-09-12; the header facts this page relies on (block_count = 65, nextn_predict_layers = 1, head_count_kv = 4, key_length = value_length = 256, context_length = 262144) were re-read from UD-Q4_K_M's own GGUF header on the same date.
Which GGUF, and why the RTX 3090 recipe no longer diverges
The sibling recipe for the RTX 3090 used to lead unsloth's IQ4_NL build, for a reason specific to that card: the only public prefill figures for this model on any 24 GB card were measured on that exact file, in llama.cpp issue #27109, on an RTX 3090. Leading with the file someone had actually measured was the right call there. That reason expired on 2026-08-19, when unsloth withdrew IQ4_NL along with every other non-Dynamic build. The measurement still stands, but no file you can download today is the one it was taken on, so it no longer selects a build. Both 24 GB pages now lead UD-Q4_K_M. There is no measurement on an RTX 3090 Ti at all — see Results for the space searched — so here the choice was always ordinary arithmetic.
On that arithmetic UD-Q4_K_M wins. Three surviving builds fit the same 131,072-token window on this card: UD-Q4_K_M at 21.097 GiB, UD-Q4_K_S (15,358,213,024 B) at 20.067 GiB and UD-IQ4_XS (14,252,845,984 B) at 19.037 GiB, each with the same projector, KV and recurrent state. The steps down buy 1.030 and 2.060 GiB of headroom against progressively fewer bits per weight, and on a card with 2.9 GiB to spare the least lossy build is the better default. Nothing else separates them here, and the runtime is unusually clear about it. ggml_cuda_should_use_mmq in ggml/src/ggml-cuda/mmq.cu has exactly one type-dependent gate — a switch of MMQ-supported types that lists GGML_TYPE_Q4_K and GGML_TYPE_IQ4_XS alike, each with its own mul_mat_q_case instantiation in the same file — and then, past a shared-memory guard that is not type-dependent either, returns true outright at if (turing_mma_available(cc)). That predicate (common.cuh) holds for any NVIDIA card at Turing or above, and this one is compute capability 8.6, so the decision never reaches the batch-size heuristic below it. Both builds take the same kernel path on this card.
For corroboration that this configuration is one people actually run rather than one this page derived: a reader publishing serving configs in discussion #34 runs unsloth's Qwen3.8-27B-Q4_K_M.gguf — the build withdrawn since, whose Dynamic successor this page leads — under llama-server at --ctx-size 131072 with --cache-type-k q8_0 --cache-type-v q8_0, --parallel 1 and -b 2048 -ub 512 — on an RTX 4090, a different 24 GB card, and without a projector, so it is corroboration of the settings and not a measurement of this card or of the vision path.
If 2.903 GiB proves too tight once your driver's reservation is accounted for, UD-Q4_K_S is the documented step down and buys back 1.030 GiB with no other change to the command. Do not climb instead: UD-Q5_K_M (19,771,509,664 B) totals 24.177 GiB in the same configuration, which is over the card outright.
What the card change does and does not affect. Relative to the RTX 3090 the sibling recipe targets, nothing in the budget above moves: both are 24 GiB parts. The memory subsystem is the same shape and slightly quicker — NVIDIA's Ada architecture whitepaper Table 1 lists the RTX 3090 Ti at 384-bit, 21 Gbps GDDR6X and 1008 GB/sec, where the GA102 whitepaper gives the RTX 3090 the same 384-bit interface at 19.5 Gbps and 936 GB/sec. That is about 7.7% more bandwidth on an otherwise identical GA102 memory path, with the same 6144 KB of L2 and the same compute capability 8.6, so every architecture-dependent branch in the CUDA backend resolves identically for the two cards. It is also not a licence to scale the RTX 3090's published numbers by 7.7% and print the result — see Results.
Installation
1. Get a CUDA llama.cpp build
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
cmake -S . -B build -DGGML_CUDA=ON
cmake --build build --config Release -j
On build numbers. The two things this recipe structurally needs are the qwen35 architecture in src/llama-arch.cpp and the qwen3vl_merger projector in tools/mtmd/clip-impl.h. Both are present at least as far back as b8001 — thousands of builds before this model was released — so the support floor is old and any remotely current build clears it. What actually matters is the opposite end: the model is days old and several bugs against it are still open, so prefer a recent build and pin the one you tested. If you intend to use a 4-bit KV cache, q4_0/q4_0 is in the default compiled set and needs no extra flag at all; only the other 4-bit types and mismatched K/V pairs do, via -DGGML_CUDA_FA_QUANTS=all — a build flag, not a build number, so no version bump substitutes for it, and prebuilt container images do not carry it. See Troubleshooting.
2. Download the weights and the vision projector
Vision is a separate file. The main GGUF alone gives you a text-only model.
pip install -U "huggingface_hub[cli]"
hf download unsloth/Qwen3.8-27B-GGUF \
Qwen3.8-27B-UD-Q4_K_M.gguf mmproj-BF16.gguf \
--local-dir ./qwen3.8-27b
mmproj-BF16.gguf declares clip.projector_type = qwen3vl_merger, which llama.cpp maps to its PROJECTOR_TYPE_QWEN3VL handler — the projector is structurally loadable by this runtime, not merely published alongside it.
Running
./build/bin/llama-server \
-m ./qwen3.8-27b/Qwen3.8-27B-UD-Q4_K_M.gguf \
--mmproj ./qwen3.8-27b/mmproj-BF16.gguf \
-ngl 99 -c 131072 --parallel 1 \
--cache-type-k q8_0 --cache-type-v q8_0 \
--flash-attn on -b 2048 -ub 512 \
--jinja --reasoning-format deepseek \
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 \
--presence-penalty 0.0 --repeat-penalty 1.0 \
--host 127.0.0.1 --port 8080
The server listens on http://127.0.0.1:8080 with an OpenAI-compatible API and a built-in web UI that accepts image uploads. On load it logs loaded multimodal model once the projector is accepted.
The sampler values above are the model card's own recommended settings for thinking mode.
--parallel 1 is load-bearing — do not trim it. The 48 linear-attention layers hold their recurrent state per sequence, and llama.cpp sizes that allocation as max(1, n_seq_max), so the cost scales with the slot count. It is easy to assume the flag is redundant, because common/common.h declares n_parallel = 1 — but that is not what llama-server runs. common/arg.cpp overrides it to -1 ("auto by default") for the server example before your arguments are parsed, and tools/server/server.cpp then resolves a negative value to n_parallel = 4 with kv_unified = true. Leaving the flag off therefore allocates the recurrent state four times over — 627,572,736 B instead of 156,893,184 B, an extra 0.438 GiB off headroom you do not have much of — for slots you did not ask for. (llama-cli and llama-mtmd-cli never take that branch, so they genuinely do default to one sequence.) The KV cache itself is unchanged either way, because the auto path also turns on a unified cache, which keeps the full context available to a single sequence rather than dividing it.
Thinking is on by default, at the highest effort setting. The card documents reasoning_effort levels of xhigh, medium and low with xhigh as the default, which is why short prompts can produce long reasoning traces. Pass --reasoning-effort low (or medium) to trade depth for latency, and --reasoning-budget 0 to disable thinking entirely. --reasoning-effort needs a build of b10434 or newer — the flag does not exist before that, and on an older build the effort level is ignored even when sent over HTTP. Troubleshooting has the exact boundary and the workaround for older builds.
Images
curl -s http://127.0.0.1:8080/v1/chat/completions -H "Content-Type: application/json" -d '{
"messages": [{"role":"user","content":[
{"type":"text","text":"What is in this image?"},
{"type":"image_url","image_url":{"url":"data:image/jpeg;base64,'"$(base64 -w0 photo.jpg)"'"}}
]}]
}'
If grounding or OCR accuracy looks poor, llama.cpp itself suggests raising the image token floor with --image-min-tokens 1024; it prints that hint on load for Qwen-VL models (example log).
Video, if you need it
The model card describes Qwen3.8-27B as "a native vision-language model that understands images and videos", and llama.cpp does have a video path — but it is a build-time option, not a given. mtmd-helper.h documents the feature as "video input helpers (requires ffmpeg/ffprobe installed on the system)", gated on the MTMD_VIDEO compile flag (default ON, force-disabled when LLAMA_SUBPROCESS is off). With such a build, llama-mtmd-cli exposes a /video <path> command and llama-server accepts an input_video content part. Confirm your binary reports video support before planning around it — a build without ffmpeg on PATH will refuse the input.
Results
- Speed: omitted, because nothing has been measured on this card. The absence is a searched one, not an assumed one. All 98 discussions on the model's HF repo were enumerated and fetched individually: no thread names this card at all. The closest are one publishing an RTX 4090 serving config, which carries no timings, and two describing a 2× RTX 3090 vLLM deployment of a different quant. GitHub's issue search returns 24 issues mentioning Qwen3.8 in the llama.cpp repository at the time of writing, and pairing that term with this card's name returns no real hit — the single result is an unrelated 2023 issue that merely happens to be numbered #3090 — while the same index returns 127 hits for the card's name across that repository overall. It also indexes comment bodies and not just opening posts — issue #27023 carries
Qwen3.8in neither its title nor its body, only in a reply, and the search returns it anyway — so an empty result is a real negative rather than a thread the index could not see. The backend has no benchmark for this pair either. If you measure your own, please send it via /contribute so it lands on /check/qwen3-8-27b/rtx-3090-ti. - Do not scale the RTX 3090's published figures to this card. Prefill numbers for this model do exist on that card — quoted in Troubleshooting below, because they are the evidence for the KV-cache trap — but they were taken on a different build of the weights, under a different cache configuration, at a prompt length this recipe is not built around. A 7.7% bandwidth advantage does not convert any of that into a figure for the RTX 3090 Ti.
- VRAM usage: ~21.1 GiB derived working set at 128K context (table above), leaving ~2.9 GiB on a 24 GB card. Derived from cited file sizes and llama.cpp's allocation formulas, not measured.
- Quality notes: this is a thinking model with
xhighreasoning effort by default; expect long traces on short questions. The card's non-thinking preset istemperature=0.7,top_p=0.80,top_k=20,presence_penalty=1.5.
For the full benchmark data, see /check/qwen3-8-27b/rtx-3090-ti.
Troubleshooting
Prefill collapses to a few dozen tokens/s after switching to a 4-bit KV cache
This is the most expensive trap on any CUDA build of this model, and the mechanism changed on 2026-09-09. Only some K/V type pairs get a compiled FlashAttention kernel. The default set is four, and llama.cpp's build documentation states it outright: GGML_CUDA_FA_QUANTS defaults to q4_0-q4_0;q8_0-q8_0;f16-f16;bf16-bf16. Note what that does and does not include — q4_0-q4_0 is compiled; every mixed pair is not.
Before PR #28079 the switch was GGML_CUDA_FA_ALL_QUANTS, and an unsupported or mismatched pair returned BEST_FATTN_KERNEL_NONE from ggml/src/ggml-cuda/fattn.cu, at which point the graph scheduler quietly relocated the attention op to the CPU backend with no error printed. After that PR it stays on the GPU and says so: the docs promise that "Combinations that were not compiled fall back to f16-f16 kernel with a warning.", and ggml_cuda_flash_attn_ext_vec logs once that it is "converting K and V to f16 instead (slow)".
It was reported on an RTX 3090 in issue #27109, where prefill on 4–10K prompts fell from 991–1276 t/s at q8_0/q8_0 to 34–106 t/s with K=q4_1, V=q8_0, and the reporter traced it to the fallback: "Prefill then runs on CPU at ~34 t/s." Generation speed was unaffected, which is what makes it hard to spot. Read those numbers as the size of the cliff, not as throughput to expect here — they are another card's, on the IQ4_NL build unsloth has since withdrawn, at a prompt length this recipe does not target.
Two fixes. Either keep K and V the same type and inside the default four — q8_0/q8_0 as the command above does, or q4_0/q4_0, which is also compiled — or rebuild with the extra kernels:
cmake -S . -B build -DGGML_CUDA=ON -DGGML_CUDA_FA_QUANTS=all
cmake --build build --config Release -j
Note that the issue's title condemns q4_0 and its own thread walks that back: the same reporter records that K=q8_0, V=q8_0 and K=q4_0, V=q4_0 are "supported, stays on GPU", and both of the thread's q4_0 datapoints are mixed K/V configurations. The rule that survives is match the two types, not "avoid 4-bit". A community pull request, #27140, proposes making all the small KV quants fast without the compile flag; as of writing it is open and unmerged, its author discloses the kernel code was AI-written, and it was tested only on the reporter's own hardware. Nothing in this recipe depends on it.
Ollama's default tag turns on speculative decoding
qwen3.8:27b and qwen3.8:27b-mtp-q4_K_M resolve to byte-identical manifests on the Ollama registry — same 16,810,714,464 B weights blob, same 931,146,016 B projector, and the same 114-byte params blob, which reads {"draft_num_predict":4,…}. The only no-speculation tag is qwen3.8:27b-q4_K_M, whose 92-byte params blob is identical except that the draft_num_predict key is absent. So ollama run qwen3.8:27b is running multi-token prediction whether or not you asked for it. Both tags fit this card, but Ollama will not give you 131,072 tokens of context by default — raise num_ctx explicitly.
Whether speculation helps is not settled, and no datapoint below is on this card. A reader benchmarking on an Apple M4 Pro in discussion #80 concluded "So under Ollama, speculation repays its own overhead in the best case and never more." — 11.78 tok/s with speculation off, 11.40 when drafts are accepted, 5.14 when they are rejected. Separately, llama.cpp issue #26750 reports draft-mtp acceptance falling to 35–41% on CUDA against ~92% on Vulkan — but its CUDA measurements are on a workstation Blackwell card and its target is a previous-generation 9B model, and a follow-up in the same thread records 84% CUDA acceptance at build b10261, placing the regression in a narrow and much newer window. Benchmark qwen3.8:27b against qwen3.8:27b-q4_K_M on your own machine before keeping it.
Multi-token prediction: which file, and whether to bother
The two GGUF publishers package the MTP head differently, and the header settles it. unsloth's build reports block_count = 65 with nextn_predict_layers = 1 — the head is inside the main file, so --spec-type draft-mtp needs no extra download. ggml-org's Qwen3.8-27B-Q4_K_M.gguf reports block_count = 64 and no such key; that repo ships the head separately as mtp-Qwen3.8-27B-Q4_0.gguf. Do not mix a target from one publisher with a drafter from the other. Note also that enabling speculation adds a 17th attention layer's worth of KV, since llama.cpp gives the MTP block a plain attention cache rather than the hybrid one — on a budget with 2.9 GiB spare that is worth measuring before you commit to it.
Reasoning effort does nothing on a build older than b10434
Issue #27023 is worth reading before you try to shorten those traces, because this model is in it. A reader running Qwen3.8-27B-Q6_K.gguf with --mmproj mmproj-F16.gguf reports "Same problem here with Qwen3.8-27B." and pins the behaviour down precisely: a top-level reasoning_effort field sent over HTTP has no effect, while the same value passed inside chat_template_kwargs does work, with /apply-template confirming that the injected system instruction changes from xhigh to low. A project contributor replies "@kidultff this is fixed on master."
"Fixed on master" is a statement made on a date, not a version, so here is the version — read out of the source at each tag rather than taken on trust. The fix is commit 7e4c0a9, "chat : pass reasoning_effort to template", merging PR #26941. The symbol it introduces, caps_apply_reasoning_effort in common/jinja/caps.cpp, is absent at tag b10433 and present at b10434 — both files fetched whole at HTTP 200, so that boundary is measured rather than inferred. The same commit is what adds the --reasoning-effort flag to common/arg.cpp in the first place: at b10433 that flag does not exist, at b10434 it does. And before it, the server honoured only reasoning_effort: none — the line the commit deletes from the server's own README says so outright: "Other values (e.g., low, max) have no effect on reasoning."
So on b10434 or newer, --reasoning-effort low behaves as the Running section describes. On anything older the flag is not recognised at all, and the workaround is the one from the thread — either --chat-template-kwargs '{"reasoning_effort": "low"}' at startup, or the same key per request:
"chat_template_kwargs": {"reasoning_effort": "low"}
--reasoning-budget 0 is unaffected either way; that flag predates the fix. Finally, read the issue's title — which sounds like a blanket "reasoning effort is broken" — against its own thread: the original reporter's two models cannot exercise the feature at all, as the first responder points out (Gemma 4 has no effort levels, Muse Glimmer uses reasoning_strength) and the reporter then concedes. The evidence that bears on this page is the later Qwen3.8-27B report, not the opening one.
The server freezes, or asserts on startup, on a multi-GPU box
Two open reports pair this model with --split-mode tensor: #27122 records reproducible CUDA lockups when tensor split is combined with --spec-type draft-mtp (the reporter notes --split-mode layer avoids it), and #27116 records a startup assertion with tensor split and an iq4_nl KV cache. Neither applies to a single RTX 3090 Ti — --split-mode is a multi-GPU setting — and they are listed only so you can rule them out if you later add a second card.
An image request crashes the server
There is one open report of silent crashes on image input, issue #27124, but it is on a Vulkan build on an AMD Ryzen AI MAX+ 395 APU, not a CUDA build on an NVIDIA card, so its failure mode should not be assumed to transfer here. If you do hit a crash on CUDA, note that the vision path for this architecture is weeks old: pin your build, and report it upstream with the build number.
Where the model came from, if you are hunting for a different quant
Qwen does not publish a GGUF of this model. The Qwen organisation's Hugging Face account carries exactly four Qwen3.8 repositories — Qwen3.8-27B, Qwen3.8-27B-FP8, Qwen3.8-2.4T-A95B and Qwen3.8-2.4T-A95B-FP8 — and none is a GGUF conversion, even though the same organisation ships first-party GGUFs for many older models. Every GGUF referenced here is a community conversion. Note also that a nominal quant name is not a size: Q4_K_M ranges from 16.81 GB (lmstudio-community) to 18.97 GB (ggml-org) across publishers, so check the byte count of the specific file you are downloading rather than trusting the label. Quantised non-GGUF builds have started to appear too — unsloth published an NVFP4 conversion on 2026-08-13 — but it is a safetensors build for vLLM-class runtimes rather than a GGUF, and the only published serving config for it names an NVIDIA DGX Spark under vLLM (discussion #34), whose author reaches for the GGUF when serving a consumer card. Nothing among the sources checked here establishes a path for that build on this card.