
LTX 2.5 is an open-weight video model from Lightricks that turns a single image or text prompt into a synchronized audio-video clip. It targets film studios, robotics teams, and solo creators who want fast, brand-controlled generation they can run on their own GPUs instead of renting a closed API.
Core Features
- 22-billion-parameter dual-stream diffusion transformer that generates video and audio in one forward pass, with no separate text-to-speech step or post-sync.
- Native multishot generation keeps a character, environment, lighting, and voice consistent across connected wide, medium, and close-up shots in a single output.
- Automatic duration prediction plus a rebuilt diffusion video decoder that cuts artifacts in high-motion footage and renders text and faces more cleanly.
- 4K HDR and RAW workflows, day-one ComfyUI templates for text-to-video and image-to-video, and quantized int8 or NVFP4 checkpoints on Hugging Face.
- A distilled fast model that runs on consumer Nvidia RTX cards with as little as 16GB of VRAM.
Best For
- Film and agency teams that need consistent multi-shot scenes without stitching separate generations by hand.
- Robotics and physical-AI groups fine-tuning the dedicated checkpoint on non-cinematic data.
- Solo creators who want watermark-free output and no per-generation fees under $10M in revenue.
Pricing
LTX 2.5 is free for commercial use for any organization under $10 million in annual revenue under the LTX-2.x Community License, with fine-tuning rights and no mandatory branding. Larger companies negotiate a paid license. The managed API is metered per second: about $0.09 at 720p rising to $0.37 at 4K, so a 10-second 720p clip costs roughly $0.90.
Our Take
Best for teams that want open weights they can self-host and fine-tune. The trade-off is the headline 6.8-second speed needs two Nvidia GB200 superchips, so realistic local runs are slower and need a 40GB-plus GPU at FP8. Browse more AI Video tools.










