| Model status | Available on MindVideo AI | Available on MindVideo AI |
| Core inputs | Text, images, videos, audio | Text, images, videos, audio |
| Supported workflows | Text to video, image to video, video to video, video editing, and reference-based generation | Text to video, image to video, video to video, video editing, and reference-based generation |
| Video duration | Up to 30 seconds in a single generation | 4 to 15 seconds |
| Reference materials | Up to 30 images, 10 video clips, and 10 audio clips | Up to 9 reference images, 3 reference videos, and 3 reference audio files |
| Multi-round extension | Supports more coherent video extension for longer sequences | Supports video extension |
| Audio generation | Improved native audio-video generation | Supports native audio generation |
| Audio and lip sync | Improved coordination between dialogue, sound, motion, and scene timing | Supports audio-visual sync and lip sync |
| Visual consistency | Better continuity for characters, products, scenes, and visual style across longer sequences | Improved subject consistency, with room for improvement in multi-subject scenes |
| Targeted editing | More targeted changes with prompts, references, and localized editing | Prompt-based editing control of clips, characters, actions, and storylines |
| 3D workflow | Clay renders and 3D references to guide composition, movement, camera direction, and lighting | Not mentioned |