
Quick answer: Wan 2.1 is Alibaba's free, open-source (Apache 2.0) AI video generator. It topped the VBench leaderboard at its February 2025 launch, beating Sora and Hunyuan Video at the time - but the field has moved fast since, and several newer models (including Wan's own 2.2 and 2.5 releases) have since surpassed it. It's still a solid, genuinely free option, just no longer the current state of the art.
Wan 2.1 by Alibaba was one of the most impressive AI video generators available at its release - and unlike commercial alternatives such as Google's VEO, it's completely free and open source under the Apache 2.0 license, so it can be downloaded and run unlimited times without usage fees. For those who prefer not to handle local installation, Promptus offers a no-code, browser-based way to run Wan 2.1 and other video models through an integrated ComfyUI interface.
The model was notable for handling complex movement and maintaining consistency across frames, including challenging scenarios like group activity, fight scenes, and animated content - areas where many models of that era struggled with physics and motion coherence.
At its February 2025 launch, Wan 2.1 scored 86.22% on VBench, topping the leaderboard ahead of Sora and Hunyuan Video. That result was real and independently benchmarked. It is not still the current top result: newer releases (including Wan's own 2.2 and 2.5 models, along with other newer commercial and open models) have since posted higher scores. If you want the current state of the art rather than this specific release, see Promptus's guides for the newer Wan versions.
Several platforms offer Wan 2.1 access without local installation: Alibaba's own Qwen platform, Hugging Face's free testing space (wait times can be long during peak usage), and Promptus, which pairs it with a full ComfyUI workflow and no setup requirement.
Local installation requires a CUDA GPU with at least 8GB VRAM (quantized versions can run lower). The installation uses ComfyUI. Required downloads: a text encoder (MT5, fp16 or fp8), a VAE file for video processing, the video model itself (14B parameters for 720p, or 1.3B for 480p), and a CLIP Vision model for image-to-video. Each component goes into its designated ComfyUI model folder.
The text-to-video workflow accepts positive and negative prompts, with settings for video dimensions, frame count, and generation parameters. The 14B model supports up to 720p. Seed, step count, and CFG scale control quality and prompt adherence - higher step counts improve quality at the cost of generation time.
Upload a starting frame and optionally describe the desired motion. The model preserves facial features and background detail while adding realistic movement - this works particularly well for portrait animation and scene transitions.
Generation time depends on hardware. An RTX 5000 with 16GB VRAM typically generates a 30-step video in 6-10 minutes. fp8-quantized models trade a small amount of quality for meaningfully faster generation, which is worth it on lower-VRAM setups.
New to this? Start with Promptus's browser-based version to get a feel for the workflow before attempting a local install - the pre-built CosyFlow handles the ComfyUI setup for you.