LTX 2.5 — Text/Image to Video

Quick answer: LTX 2.5 is Lightricks' text/image-to-video model, generating video with synchronized audio in a single run. This CosyFlow runs the official two-stage distilled pipeline — a base pass plus a 2× spatial upscale — and needs a CUDA GPU with 32GB+ VRAM, or CosyCloud if your hardware doesn't meet that.

Before you run this

This is a heavier template than most of the CosyFlow library. LTX 2.5 needs a CUDA GPU with 32GB+ VRAM and about 100GB of free disk space for the model files, downloaded once on first use. That's meaningfully more than the image-generation CosyFlows in this library. If your machine doesn't meet that, running this locally will fail at either the download step or the first generation — use CosyCloud instead rather than troubleshooting a local run that hardware can't support.

Windows desktop or CosyCloud only — this template runs on Promptus's ComfyUI environment, which requires the desktop app.

What this template does

This is LTX's own official "start here" workflow for LTX 2.5 — text-to-video, with an optional still image as your opening frame. It runs in two stages: a base-resolution pass, a 2× spatial upscale, and a short refine pass, and it generates the video and its audio together in one run — you don't need a separate audio step.

If you need something this template doesn't do, there are other LTX 2.5 CosyFlows for those cases: audio-to-video (matching video to an existing audio track), text-to-audio only, video-to-video with a reference clip, inpainting/outpainting on existing footage, and image-to-video with motion paths you draw yourself. Search "LTX 2.5" in CosyFlows to see the full set.

How to use it

  1. Open the LTX 2.5 CosyFlow. On first use, Promptus will need to download the required model files — this only happens once.
  2. Write your prompt. Be specific about what happens in the shot, not just the mood. Describe, in order: the subject, then what it does (the sequence of movement), then the camera, then lighting and atmosphere, then any sound. A prompt like "a woman walks through a neon-lit night market, steam rising from food stalls, camera tracking beside her, market chatter and sizzling in the background" gives the model something concrete to follow. A vague prompt like "cool cyberpunk city vibe" gives it much less to work with.
  3. Optional: add a first-frame image. If you want the video to open on a specific image rather than something purely generated from text, add it here. The model will animate forward from that frame.
  4. Set your duration. Clip length has to land on 1 + a multiple of 8 frames — Promptus converts your seconds input into frames automatically, so the actual output may land a touch off whatever number you typed. That's expected, not an error.
  5. Run it. The base pass generates first, then the upscale and refine pass runs automatically after — you'll see two stages complete, not one.
  6. Review before you commit to more. Check composition, motion, and how well the audio lines up before generating variations at higher settings. A quick low-effort pass to confirm the idea works is cheaper than repeatedly running the full two-stage pipeline on a prompt that isn't there yet.

Tips for better results

  • Order matters in your prompt: subject → action/sequence → camera → lighting/atmosphere → sound. The model follows structured instructions more reliably than a mood-only description.
  • If dialogue matters, keep individual lines short and give each one room — don't cram multiple lines of speech into a short clip.
  • This is the distilled model variant — built for fast iteration, not maximum fidelity. If you need the highest possible quality and have the time/hardware budget for it, that's a different (non-distilled) LTX 2.5 pipeline, not this template.
  • For the full prompting guide, see LTX's own writeup: ltx.io/blog/prompting-guide-for-ltx-2.

If something goes wrong

  • Fails at the download step: check free disk space (100GB) before retrying — a partial download from a full disk can look like a different error than what it actually is.
  • Out-of-memory during generation: your GPU doesn't meet the 32GB VRAM requirement for local execution — switch to CosyCloud rather than trying to force a local run down to a lower resolution, since this template isn't tuned for lower-VRAM hardware the way some image CosyFlows are.
  • Audio and video feel out of sync: confirm you're on the two-stage template and not the single-stage variant — the single-stage version trades some quality for speed and can be less precise on sync-sensitive prompts.
  • Output looks soft/low-detail: that's expected from the distilled model at speed-optimized settings. If detail matters more than speed for this particular generation, that's a signal to use the non-distilled pipeline instead of pushing this template past what it's built for.

Frequently asked questions

Does LTX 2.5 generate audio automatically?
Yes — video and audio are generated together in one run, so you don't need a separate audio step.

How much VRAM do I need to run LTX 2.5 locally?
At least 32GB, plus about 100GB of free disk space for the model files (downloaded once on first use). Below that, use CosyCloud instead of trying to force a local run.

Can I start the video from my own image?
Yes — add an optional first-frame image and the model animates forward from it.

Why did my output come out slightly shorter or longer than the duration I set?
Clip length has to land on 1 + a multiple of 8 frames, so Promptus rounds your seconds input to the nearest valid frame count. That's expected, not an error.

Is this the highest-quality LTX 2.5 pipeline?
No — this is the distilled variant, built for fast iteration. The non-distilled pipeline is a separate template for maximum fidelity.

Reviews

What our community is saying about us

ai art illustrator
Jose Romero
Illustrator
ai art generator
Dmitry Selivanov
Art director
ai art brand
Sadiya Abdullah
Brand agency
ai video generator
Lupe Rodriguez
Product designer
ai art graphic designer
Dianne Russell
Graphic Designer
ai game designer
Marie McKinney
Game designer
ai image generator

Share your compute

Join our distributed GPU compute network. Help us make AI accessible, scalable
and secure for designer, developers and start-ups.

promptus ai video generator
FAQ Section
Frequently Asked Questions
What is Promptus?

Promptus enable users to generate customized, AI videos, high-resolution images, AI characters, music, audio, chat and more with ease using the latest's AI models.

Anyone can use Promptus AI—designers, marketers, hobbyists, or beginners—to generate AI art, photos, and videos quickly and easily.

What is an AI workflow?

An AI workflow is a structured process that turns inputs into repeatable outputs using models, prompts, and logic.

Can I create realistic photos with AI?
What is the GPU compute?
Does Promptus work with ComfyUI workflows?
Start running your first workflow
Go from idea to production-ready output in minutes.
Try Promptus for free ➜