Qwen3.8 27B on RTX 5060 Ti
Yes — Qwen3.8 27B runs on the RTX 5060 Ti (16 GB). Fastest community-measured generation speed: 28.49 tokens/s.
Qwen3.8-27B on RTX 5060 Ti: a vision-capable 27B in 16 GB on a 128-bit bus
| Task | Quant | Speed | VRAM | Works | Confidence | Source | Verified |
|---|---|---|---|---|---|---|---|
| llm | UD-Q3_K_XL | 579.2prefill tokens/s | ✓ | huggingface.co· first-party measurement | 2026-09-14 | ||
| llm | UD-Q3_K_XL | 28.49tokens/s | ✓ | huggingface.co· first-party measurement | 2026-09-14 |
- llmUD-Q3_K_XL579.2 prefill tokens/s
Prompt processing in the llama-server run: 1,043 prompt tokens including one 1024x1024 image, vision projector loaded, 32,768 context, q8_0 K/V, flash attention, --parallel 1. llama-bench text-only prefill of 512 tokens at depth: 956 tok/s empty, 817 at 16K, 715 at 32K, 569 at 65,536 context. llama.cpp b10951, unsloth UD-Q3_K_XL revision 4ca72078. Measured by the site operator on his own card, Windows 11, driver 591.86, display attached. Unreplicated: one rig, one operator.
recorded from huggingface.co ↗ · verified 2026-09-14
- llmUD-Q3_K_XL28.49 tokens/s
llama-server with the recipe's flags: mmproj-F16 vision projector loaded, 32,768 context, q8_0 K/V, flash attention, --parallel 1, -ngl 99; one 1024x1024 image, 1,043 prompt tokens, 151 generated. llama.cpp b10951 (official Windows CUDA 13.3 build), unsloth UD-Q3_K_XL at revision 4ca72078 (13,146,393,504 B). Whole-card peak 15,712 of 16,311 MiB (599 MiB free) with the desktop and a resident ComfyUI CUDA context included; shared GPU memory rose 132 MiB, so it runs in VRAM. llama-bench text-only decode, 128 tokens at depth: 28.73 tok/s empty, 26.38 at 16K, 24.24 at 32K, 20.85 at 65,536 context (peak 15,783 MiB, fits). Measured by the site operator on his own card, Windows 11, driver 591.86, display attached. Unreplicated: one rig, one operator.
recorded from huggingface.co ↗ · verified 2026-09-14
- multimodalintermediate48GB+
Qwen3.8-27B on Apple M4 Max: 4-bit MLX Vision-Language at the Full 262K Context
- multimodalintermediate48GB+
Qwen3.8-27B on Apple M3 Max: 4-bit MLX Vision-Language at the Full 262K Context
- multimodaladvanced16GB+
Qwen3.8-27B on RX 7800 XT: a vision-capable 27B inside 16 GB on ROCm
- multimodalintermediate24GB+
Qwen3.8-27B on RX 7900 XTX: 128K-context vision chat on ROCm with llama.cpp
Can you run Qwen3.8 27B on RTX 5060 Ti?
Yes — Qwen3.8 27B runs on the RTX 5060 Ti (16 GB). Fastest community-measured generation speed: 28.49 tokens/s.
Which quantizations have been tested for Qwen3.8 27B on RTX 5060 Ti?
UD-Q3_K_XL — measured in community benchmarks.
How fast is Qwen3.8 27B on RTX 5060 Ti?
Up to 28.49 tokens/s (llm), the fastest community-measured generation speed.
Are there step-by-step instructions for Qwen3.8 27B on RTX 5060 Ti?
Yes — a step-by-step recipe documents Qwen3.8 27B on the RTX 5060 Ti, linked at the top of this page.