Try these ideas:
🎬 One video · $2.99 · no sign-up — we’ll email you the link. Already have credits? Sign in
Wan 3.0 by Alibaba unifies text, image, video, audio, keyframes and document references in a single model. Generate 480p, 720p or 1080p videos up to 30 seconds with native audio — text-to-video and image-to-video with first/last frame control, live on VO3 AI.
Try Wan 3.0 text-to-video and image-to-video directly. Select Wan 3.0 from the model dropdown to experience Alibaba's latest omni-reference video technology.
Generate AI Model videos with AI
Try these ideas:
🎬 One video · $2.99 · no sign-up — we’ll email you the link. Already have credits? Sign in
Wan 3.0 is Alibaba's most advanced AI video generation model — a major leap from Wan 2.7. Instead of separate endpoints for each task, Wan 3.0 handles text-to-video, image-to-video, and first/last-frame animation through a single unified model, and layers in omni-reference understanding: images, videos, audio, keyframes, documents, and webpages can all shape the scene, motion, sound, visual identity, and narrative direction of the output.
Wan 3.0 is available on VO3 AI through the KIE API infrastructure. It outputs 480p, 720p, or 1080p video with durations up to 30 seconds and native audio built in. On the create page you can use Wan 3.0 today for text-to-video and image-to-video with first and last frame control.
Generate 480p/720p/1080p videos with native audio from
a text prompt — English or Chinese
Animate a single image with
first-frame control and adaptive aspect ratio
Upload a start and end frame —
Wan 3.0 generates smooth motion between them
Up to 10 reference images
to lock character and style consistency
Up to 5 reference videos to carry
motion and look into a new scene
Up to 5 audio references to drive rhythm, voice and timing
Reference a document or webpage
to shape the narrative direction
Prompts up to 20,000 characters in English and Chinese
Three resolution tiers for speed-first
drafts up to cinematic 1080p output
Longer-form generation with duration-based
pricing for cost control
Built-in audio synthesis that
matches generated visuals automatically
Adaptive output plus 16:9, 9:16, 1:1, 4:3 and 3:4 framing
One endpoint for text, image,
keyframe and reference generation
Omni-reference keeps character identity
stable across the scene
Image-to-video uses the same rates · Credits start from $2.99 · View all plans
Omni-reference generation, 480p/720p/1080p output, native audio. Alibaba's most capable AI video model, available on VO3 AI.
Wan 3.0 AI video generator | Alibaba Wan 3.0 model | Wan 3.0 text to video | Wan 3.0 image to video | Wan 3.0 omni-reference | Wan 3.0 first last frame | Wan 3.0 native audio | AI video generator 2026 | VO3 AI Wan 3.0