self-hosted/ai
§01·compatibility · /check

Qwen3.8 27B on RTX 5060 Ti

Yes — Qwen3.8 27B runs on the RTX 5060 Ti (16 GB). Fastest community-measured generation speed: 28.49 tokens/s.

✓ runsmultimodalactive50 series16GB VRAM
step-by-step recipe for this pair

Qwen3.8-27B on RTX 5060 Ti: a vision-capable 27B in 16 GB on a 128-bit bus

multimodaladvanced16GB+
model
name
Qwen3.8 27B
slug
qwen3-8-27b
vertical
multimodal
status
active
open detail ↗
gpu
name
RTX 5060 Ti
slug
rtx-5060-ti
vram
16 GB
series
50
open detail ↗
§02·benchmarks
TaskQuantSpeedVRAMWorksConfidenceSourceVerified
llmUD-Q3_K_XL579.2prefill tokens/s✓huggingface.co· first-party measurement2026-09-14
llmUD-Q3_K_XL28.49tokens/s✓huggingface.co· first-party measurement2026-09-14
§03·how it was measured
  • llmUD-Q3_K_XL579.2 prefill tokens/s

    Prompt processing in the llama-server run: 1,043 prompt tokens including one 1024x1024 image, vision projector loaded, 32,768 context, q8_0 K/V, flash attention, --parallel 1. llama-bench text-only prefill of 512 tokens at depth: 956 tok/s empty, 817 at 16K, 715 at 32K, 569 at 65,536 context. llama.cpp b10951, unsloth UD-Q3_K_XL revision 4ca72078. Measured by the site operator on his own card, Windows 11, driver 591.86, display attached. Unreplicated: one rig, one operator.

    recorded from huggingface.co ↗ · verified 2026-09-14

  • llmUD-Q3_K_XL28.49 tokens/s

    llama-server with the recipe's flags: mmproj-F16 vision projector loaded, 32,768 context, q8_0 K/V, flash attention, --parallel 1, -ngl 99; one 1024x1024 image, 1,043 prompt tokens, 151 generated. llama.cpp b10951 (official Windows CUDA 13.3 build), unsloth UD-Q3_K_XL at revision 4ca72078 (13,146,393,504 B). Whole-card peak 15,712 of 16,311 MiB (599 MiB free) with the desktop and a resident ComfyUI CUDA context included; shared GPU memory rose 132 MiB, so it runs in VRAM. llama-bench text-only decode, 128 tokens at depth: 28.73 tok/s empty, 26.38 at 16K, 24.24 at 32K, 20.85 at 65,536 context (peak 15,783 MiB, fits). Measured by the site operator on his own card, Windows 11, driver 591.86, display attached. Unreplicated: one rig, one operator.

    recorded from huggingface.co ↗ · verified 2026-09-14

§04·more Qwen3.8 27B recipes
§05·common questions
Can you run Qwen3.8 27B on RTX 5060 Ti?

Yes — Qwen3.8 27B runs on the RTX 5060 Ti (16 GB). Fastest community-measured generation speed: 28.49 tokens/s.

Which quantizations have been tested for Qwen3.8 27B on RTX 5060 Ti?

UD-Q3_K_XL — measured in community benchmarks.

How fast is Qwen3.8 27B on RTX 5060 Ti?

Up to 28.49 tokens/s (llm), the fastest community-measured generation speed.

Are there step-by-step instructions for Qwen3.8 27B on RTX 5060 Ti?

Yes — a step-by-step recipe documents Qwen3.8 27B on the RTX 5060 Ti, linked at the top of this page.