Quick answer: Stable Video Diffusion is Stability AI's latent diffusion model fine-tuned for image-to-video generation, producing 14–25 frames at up to 576×1024 resolution while preserving frame-to-frame consistency. Frame rate is customizable between 3 and 30 fps, with a typical clip rendering in under 2 minutes — though newer local models like Wan and LTX-Video have largely superseded it for current use.
Latent diffusion finetuned for image-to-video: generates 14–25 frames at up to 576×1024, preserving content consistency frame-to-frame. Stability AI’s own release notes the frame rate is customizable between 3 and 30 fps, with a typical clip rendering in under 2 minutes.
Join our distributed GPU compute network. Help us make AI accessible, scalable
and secure for designer, developers and start-ups.