Ready to create some AI magic? Describe your scene, pick your settings, and watch as AI brings your vision to life in stunning video.
WAN 3.0 is Alibaba's newest video generation model on Zyka. Unified text-to-video and image-to-video with native audio, up to 30 seconds, and 480p–1080p output.
Start from a text prompt alone (T2V), or upload a start frame — optionally with an end frame for interpolation (I2V).
Pick resolution (480p–1080p), duration (2–30s), aspect ratio, and whether to include generated audio.
WAN 3.0 produces smoother motion and stronger scene coherence than earlier WAN generations.
WAN 3.0 is Alibaba's latest-generation video model on Zyka, succeeding WAN 2.7 with improved motion smoothness, scene fidelity, and visual coherence.
Unlike WAN 2.6's separate T2V and I2V models, WAN 3.0 is a single unified model — one form handles text-only, image animation, and start/end frame interpolation.
Native audio generation is built in: enable sound generation to produce synchronized audio alongside your video without uploading a separate audio file.
WAN 3.0 supports longer clips (up to 30s vs 15s), 480p resolution, adaptive aspect ratio, and native audio generation. It uses fal's alibaba/wan-3.0 endpoints.
Yes. Upload a start frame to animate it, or add both start and end frames for interpolation — all in the same model.
WAN 3.0 generates audio natively via a toggle. Custom audio upload is not supported on this model — use WAN 2.6 or 2.7 for uploaded audio workflows.
480p, 720p, and 1080p. Pricing scales with resolution per second of output.
Billing is per second of video at fal rates plus 20%: 6 credits/s at 480p ($0.06/s), 12 credits/s at 720p ($0.12/s), and 24 credits/s at 1080p ($0.24/s). T2V, I2V, and reference-to-video use the same rate. Example: a 5s clip at 720p costs 60 credits.