Qwen-Image-2.1 on RTX 5060 Ti
Yes — Qwen-Image-2.1 runs on the RTX 5060 Ti (16 GB). Fastest community-measured result: 41.92 s.
Qwen-Image-2.1 on RTX 5060 Ti: the int8 ComfyUI template, 2K output and editing in 16 GB
| Task | Quant | Speed | VRAM | Works | Confidence | Source | Verified |
|---|---|---|---|---|---|---|---|
| t2i | bf16 DiT + int8_convrot encoder | 41.92s | ✓ | huggingface.co· first-party measurement | 2026-09-22 | ||
| t2i | int8_convrot (DiT + encoder) | 21.39s | ✓ | huggingface.co· first-party measurement | 2026-09-22 |
- t2ibf16 DiT + int8_convrot encoder41.92 s
Same session and template with the diffusion model swapped for qwen_image_2.1_bf16 (13.25 GiB): 1024x1024, 25 steps, 1.38 s/it, cold start. 2048x2048: 208.45 s (7.43 s/it). On this card and build the int8 template is 2.0x faster per image at 1024x1024 (21.39 s) and 1.7x at 2048x2048 (124.08 s), with a near-identical image at the same seed. On the other install (post-tag core, cu128 PyTorch, int8 on the eager path) bf16 took 46.17 s and was the faster of the two. Whole-card peak 15,657 MiB at 1024x1024 (the dynamic VRAM loader fills the card). Shared GPU memory rose 1.2-4.3 GiB with dedicated memory never sampled full: cause not established (the model drive was auto-detected as fast, a mode that offloads over unpinned RAM); not attributed to fallback. ComfyUI 0.37.0 portable, PyTorch 2.13.0+cu130, Windows 11, driver 591.86, display attached. Unreplicated: one rig, one operator.
recorded from huggingface.co ↗ · verified 2026-09-22
- t2iint8_convrot (DiT + encoder)21.39 s
ComfyUI 0.37.0 Windows portable, PyTorch 2.13.0+cu130 (comfy-kitchen CUDA backend enabled). Comfy-Org 'Qwen Image 2.1: Text to Image' template unchanged: qwen_image_2.1_int8_convrot + qwen3vl_8b_int8_convrot + qwen_image_2.1_vae_bf16 (Comfy-Org/Qwen-Image-2.1 @ ace0edeb), 25 steps, cfg 1, euler, simple. Cold start (models unloaded first), so the time includes loading the weights; sampling 0.61 s/it; the repeat with a new seed took 20.62 s. 2048x2048: 124.08 s (4.35 s/it). On a different install, a Comfy Desktop core from post-tag master on a cu128 PyTorch (ComfyUI's cu130 warning; int8 on the eager path; Python 3.12), the same two runs took 69.41 s and 297.92 s. Whole-card peak 15,560-15,844 of 16,311 MiB at 1024x1024: the dynamic VRAM loader fills the card, not a requirement. Dynamic VRAM and pinned memory on (defaults). No OOM. Windows 11, driver 591.86, display attached. Unreplicated: one rig, one operator.
recorded from huggingface.co ↗ · verified 2026-09-22
- imageintermediate12GB+
Qwen-Image-2.1 on RTX 5070: int8 template in 12 GB, 2K, editing, and what native NVFP4 changes
- imageintermediate12GB+
Qwen-Image-2.1 on RTX 4070 Ti: int8 ComfyUI in 12 GB, with third-party 4 MP and edit timings
- imageintermediate12GB+
Qwen-Image-2.1 on the 12 GB RTX 4070 SUPER: what fits, int8 ComfyUI setup, 2K and editing
- imageintermediate12GB+
Qwen-Image-2.1 on any RTX 4070 board: int8 ComfyUI install in 12 GB, 2K and image editing
Can you run Qwen-Image-2.1 on RTX 5060 Ti?
Yes — Qwen-Image-2.1 runs on the RTX 5060 Ti (16 GB). Fastest community-measured result: 41.92 s.
Which quantizations have been tested for Qwen-Image-2.1 on RTX 5060 Ti?
bf16 DiT + int8_convrot encoder, int8_convrot (DiT + encoder) — measured in community benchmarks.
How fast is Qwen-Image-2.1 on RTX 5060 Ti?
Up to 41.92 s (t2i), the fastest community-measured result.
Are there step-by-step instructions for Qwen-Image-2.1 on RTX 5060 Ti?
Yes — a step-by-step recipe documents Qwen-Image-2.1 on the RTX 5060 Ti, linked at the top of this page.