LTX-2.5
Lightricks' LTX-2.5 is a 22B audio-video DiT that generates synchronised video and audio from a text, image, or video input. Its headline change over LTX-2.3 is native multishot generation: several connected shots produced in a single pass, holding character identity, environment, lighting, voice and visual style across the cuts, where earlier versions of the line produced one continuous shot.
The release ships split rather than as a monolith - one safetensors file per component. That means a 22B transformer in either a full (dev) or distilled form, a custom Gemma 4 12B text encoder, separate video and audio VAEs, latent upscalers, and an optional duration head that infers a clip's length from the prompt. A new diffusion video decoder replaces the older VAE reconstruction stage.
It runs locally through the ltx-pipelines CLI (CUDA 12.7 or newer), through ComfyUI, or through Diffusers. Released under the LTX-2 Community License, which is free for commercial use below $10M annual revenue. The HuggingFace repository is gated - accept the licence terms before downloading.
✓ benchmarked·~ runs via recipe (not benchmarked)·— untested·✕doesn't fit