self-hosted/ai
§01·recipe · video

Wan 2.2 TI2V-5B on RTX 5060 Ti: 720p Text/Image-to-Video in ComfyUI

videointermediate8GB+ VRAM
May 19, 2026

This intermediate recipe sets up Wan 2.2 TI2V-5B on the RTX 5060 Ti, needing about 8 GB of VRAM.

models
tools
prerequisites
  • NVIDIA RTX 5060 Ti (16GB VRAM) or any 8GB+ card with ComfyUI native offloading
  • Python 3.10+ (3.12 recommended for the Blackwell 50-series)
  • ComfyUI installed and updated to a build that ships the Wan 2.2 templates
  • 32GB+ system RAM recommended (offloading is RAM-heavy)

What You'll Build

A local ComfyUI pipeline that turns a text prompt (or a starting image) into a 5-second 720p video using the Wan 2.2 TI2V-5B model — the only Wan 2.2 variant that fits a consumer 16GB card. The recipe walks through both the official native workflow (FP16 safetensors with built-in offloading) and the QuantStack Q8 GGUF path for tighter VRAM.

Hardware data: RTX 5060 Ti (16GB VRAM) · measured: 5 s of 1280×704 video in 541.8 s on FP16, whole-card peak 15,766 of 16,311 MiB · See benchmark data

ℹ️ The timings and memory readings on this page are measured. On 2026-09-14 the official template ran six times on the site operator's own RTX 5060 Ti 16GB — four on the FP16 path, two on the Q8 GGUF path, at 1024×576 and 1280×704 — in ComfyUI 0.32.0 with no attention accelerator installed. All six completed. The models were unloaded before every run, so the times include loading them. Readings are whole-card nvidia-smi samples with the Windows desktop included. One rig, one operator: unreplicated. The raw session is published here, and the figures are filed at /check/wan-2-2/rtx-5060-ti.

Why TI2V-5B and not the 14B variants? The official Wan-Video/Wan2.2 repo states T2V-A14B, I2V-A14B, S2V-14B and Animate-14B all need "at least 80GB VRAM" on a single GPU. Only TI2V-5B (5B dense, not MoE) is documented as a single-consumer-GPU target.

Requirements

ComponentMinimumTested
GPU8GB VRAM (per ComfyUI native offloading note)RTX 5060 Ti (16GB) — measured, six runs, zero OOM
RAM16GB32GB (31.1 GiB usable) — the community recommends 64GB for Q8 + Sage Attention
Storage~13GB (FP16 5B weights + VAE + text encoder) or ~6GB (Q8 GGUF + VAE + text encoder)—
SoftwareComfyUI (recent build), Python 3.10+, PyTorch ≥ 2.4ComfyUI 0.32.0, PyTorch 2.12.0.dev cu128, Python 3.12

Installation

1. Install / update ComfyUI

Use a build new enough to expose the Wan 2.2 templates under Workflow → Browse Templates → Video → "Wan2.2 5B video generation". See the ComfyUI Wan 2.2 tutorial for the menu path.

2. Download model files (official FP16 path)

Per the ComfyUI native workflow docs, place these files in ComfyUI/models/:

ComfyUI/models/diffusion_models/wan2.2_ti2v_5B_fp16.safetensors
ComfyUI/models/text_encoders/umt5_xxl_fp8_e4m3fn_scaled.safetensors
ComfyUI/models/vae/wan2.2_vae.safetensors

Or grab the raw weights from the official HF repo using the Wan-AI install guide:

pip install "huggingface_hub[cli]"
huggingface-cli download Wan-AI/Wan2.2-TI2V-5B --local-dir ./Wan2.2-TI2V-5B

3. (Alternative) Install ComfyUI-GGUF and a Q8 quant

For lower peak VRAM on a 16GB card, use the community Q8 quant from QuantStack/Wan2.2-TI2V-5B-GGUF (the repo confirms it is "a direct conversion of Wan-AI/Wan2.2-TI2V-5B").

Install city96/ComfyUI-GGUF:

git clone https://github.com/city96/ComfyUI-GGUF ComfyUI/custom_nodes/ComfyUI-GGUF
pip install --upgrade gguf

Download the Q8_0 file (5.4 GB) and place it in the unet folder:

huggingface-cli download QuantStack/Wan2.2-TI2V-5B-GGUF \
  Wan2.2-TI2V-5B-Q8_0.gguf \
  --local-dir ./ComfyUI/models/unet

In the official template, swap the Load Diffusion Model node for Unet Loader (GGUF) (under the bootleg category) and point it at the .gguf file.

4. (Optional) Install Sage Attention 2.2

Community-reported in the Wan2.2-Animate-14B discussion #4 as what that user ran for a ~7-minute, 5-second clip at 1024×574 on a 5060 Ti. It is not required: every measured run on this page used no attention accelerator at all. The two setups differ in more than Sage, so do not read the timings against each other as a verdict on it. Prebuilt Windows wheels: woct0rdho/SageAttention v2.2 release.

Running

After loading the Wan2.2 5B video generation template, enter a prompt in the positive-prompt node and queue. The Wan22ImageToVideoLatent node exposes resolution and frame count.

If you prefer the CLI path from the official repo:

git clone https://github.com/Wan-Video/Wan2.2.git
cd Wan2.2
pip install -r requirements.txt
python generate.py --task ti2v-5B --size 1280*704 \
  --ckpt_dir ./Wan2.2-TI2V-5B \
  --offload_model True --convert_model_dtype --t5_cpu \
  --prompt "a panda playing guitar by a lake at sunset"

The CLI command above is the exact invocation the Wan 2.2 README documents for TI2V-5B on a 24GB+ card; on a 16GB 5060 Ti the ComfyUI route with offloading (or the GGUF route) is the more reliable option.

Results

Measured on the site operator's RTX 5060 Ti on 2026-09-14 — ComfyUI 0.32.0, the official template's settings (20 steps, uni_pc, CFG 5, shift 8, 24 fps), no attention accelerator, Windows 11 with a monitor on the card:

weightssize × framestimewhole-card peak
FP161024×576 × 121 (5 s)305.1 s15,813 MiB
FP161280×704 × 121 (5 s, template default)541.8 s15,766 MiB
FP161280×704 × 121, repeat542.5 s15,795 MiB
FP161280×704 × 145 (6 s)696.7 s14,860 MiB
Q8_0 GGUF1024×576 × 121 (5 s)310.8 s14,661 MiB
Q8_0 GGUF1280×704 × 121 (5 s)545.1 s14,635 MiB

Every run started cold — models unloaded first — so each time includes loading them from disk. The card reports 16,311 MiB in total. The 6-second run peaked lower than the 5-second one at the same size; nothing in the session tests why, so do not plan around it.

  • Speed: about 9 minutes for the template's default 5-second 1280×704 clip (541.8 s), on either path, and about 5 minutes at 1024×576 (305.1 s FP16, 310.8 s Q8). For comparison, Ricardo130's report in HF discussion #4 puts a 1024×574 5-second clip at ~7 minutes with Q8 GGUF + Sage Attention 2.2 on the same card, on a different install. Raw session: on Hugging Face.
  • VRAM usage: TI2V-5B is documented as fitting "well on 8GB vram with the ComfyUI native offloading" per the official ComfyUI tutorial. On the 16GB card, measured, the default FP16 clip peaked at 15,766 MiB of 16,311 — 545 MiB free with the desktop included, which fits but is not much headroom — and the Q8 GGUF path peaked 1,131 MiB lower at the same size (1,152 MiB lower at 1024×576). Windows' shared GPU memory rose about 12.4 GiB during the FP16 runs — about the ceiling ComfyUI sets for pinned host memory on Windows, 40% of system RAM (MAX_PINNED_MEMORY = ram * 0.40 in comfy/model_management.py), which is 12.4 GiB on the measured 32 GB machine — and 7.3 GiB during the Q8 runs. Dedicated memory was never sampled within 256 MiB of full, so none of that rise is attributed to VRAM spilling over. Live data: /check/wan-2-2/rtx-5060-ti
  • Quality notes: TI2V-5B is the only Wan 2.2 variant the official repo documents as runnable on a single consumer GPU. The 14B-class siblings (T2V-A14B, I2V-A14B, Animate-14B, S2V-14B) are documented as needing 80GB+ and are out of scope for this card unquantized. The TI2V-5B output is 720p (1280×704 or 704×1280) at 24 fps for 5 seconds.

For the full benchmark data, see /check/wan-2-2/rtx-5060-ti.

Troubleshooting

"Incompatibility error with the VGA 5xxx series" on first install

Reported by the original poster in discussion #4. The Blackwell 50-series needs a recent CUDA + PyTorch stack. The working combo reported in that thread is PyTorch 2.9 with CUDA 12.9 on Python 3.12. Older torch wheels (without sm_120 kernels) silently fall back to CPU or fail at model load.

Out of memory on 16GB even with FP16

Make sure ComfyUI's native offloading is active (it is by default in recent builds — the official tutorial explicitly relies on it for the 8GB minimum claim). If the FP16 path still OOMs at 720p, drop to the QuantStack Q8_0 GGUF (5.4 GB on disk) via the Unet Loader (GGUF) node — peak VRAM drops by about 1.1 GiB (measured at 1280×704 and 1024×576) and quality loss is minimal at Q8. It does not make generation faster: the measured Q8 runs took 3–6 s longer than FP16 at the same size.

Want the 14B variants?

Per the official README, the 14B variants require 80GB+ single-GPU VRAM. Community GGUF quants of the 14B Wan variants exist but are out of scope for this recipe — file a request on /contribute if you want a 14B-quantized recipe added once a stable workflow lands.

share this recipe
Share
common questions
How much VRAM does Wan 2.2 TI2V-5B need?

About 8 GB — the minimum this recipe targets.

Which GPUs is Wan 2.2 TI2V-5B tested on?

RTX 5060 Ti (16 GB).

How hard is this setup?

Intermediate — follow the steps above.

next