Text to video
Describe the subject, motion, camera, lighting, and sound in one connected prompt.
Create short AI video from a text prompt or a source image. MiniMax H3 Max supports text-to-video and image-to-video, 480P or 768P output, 5–15 second clips, and a published benchmark reports a 5-second render in under 3 seconds.
* This published benchmark is for a 5-second H3 Max clip; actual time can vary by queue and settings.
Create 5–15 second AI videos from a prompt or a source image, then preview and download the completed MP4.
You can describe both the visual scene and the audio/dialogue in this single prompt.
Your generated video will appear here
MiniMax H3 Max is a post-trained MiniMax H3 variant built for fast, prompt-led video creation. It tunes the base model toward stronger prompt adherence and visual aesthetics, so the details you specify—subject, action, camera, lighting, pace, and sound direction—have a clearer role in the finished shot.
Speed is the practical difference. A published benchmark reports that a five-second H3 Max clip renders in under three seconds. That makes it useful for rapid creative iteration: write a focused prompt, inspect motion and framing, then make the next pass more specific instead of waiting through a long render cycle.
Use text-to-video when the whole scene starts as a written brief. Switch to image-to-video when a product photo, character still, illustration, or keyframe should anchor the opening frame; an optional end image can guide where the shot finishes. The current workspace offers 480P or 768P output and 5–15 second clips.
Describe the subject, motion, camera, lighting, and sound in one connected prompt.
Choose 480P for faster drafts or 768P when the shot needs more visual detail.
Generate short clips in whole-second durations. A published benchmark reports a 5-second H3 Max clip renders in under 3 seconds.
Text-to-video supports 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16 compositions.
Image-to-video can use a source image and an optional end image to guide the transition.
Build one focused shot at a time: connect the subject, action, camera movement, lighting, pacing, and sound in a clear instruction.
Animate a pack shot, poster, interface, or product image with controlled camera motion and a concise sound cue.
Use a still character frame as the visual anchor, then describe expression, gesture, timing, and the scene’s sound.
Turn a launch concept into a 5–15 second vertical, square, widescreen, or cinematic social-ready draft.
Play two completed H3 Max examples: one generated from text and one guided by an input image.
A text prompt directs subject, scene, camera movement, atmosphere, and sound direction.
A source image anchors the opening frame while the prompt describes expression, dialogue, and scene motion.
Choose H3 Max for fast text or image generation with post-training for prompt adherence and aesthetics. Choose MiniMax H3 when the project needs 2K output or the broader reference and editing workflow.
| Dimension | MiniMax H3 Max | MiniMax H3 |
|---|---|---|
| Positioning | Post-trained H3 for prompt adherence, aesthetics, and throughput | Open-weights general-purpose multimodal video model |
| Resolution | 480P or 768P | 2K at 24 fps |
| Speed | Published benchmark: a 5-second clip renders in under 3 seconds | Use H3 when 2K or the broader multimodal workflow matters more than speed |
| Generation inputs | Text-to-video and image-to-video | Text, images, video, and audio in one context |
| Endpoint scope | Text-to-video; image-to-video with optional end frame | Text-to-video, image-to-video, reference-to-video, and editing |
| Best fit | Fast drafts, prompt-led shots, and image animation | 2K delivery, multi-reference direction, and footage edits |
Choose a source, describe one connected shot, then generate and refine after you can see the motion and timing.
Start with a text prompt for text-to-video, or upload a source image for image-to-video.
State what is on screen, what changes, how the camera moves, and what the scene should sound like.
Select duration and resolution. Review the MP4, then make the next prompt or source image more specific.
Clear answers about the H3 Max workflow, supported output settings, and the best way to use prompts and source images.
It is a post-trained MiniMax H3 variant for text-to-video and image-to-video, tuned for prompt adherence and visual aesthetics.
The workspace offers 480P or 768P output, 5–15 second duration, and six text-to-video aspect ratios from 21:9 through 9:16.
Yes. Image-to-video uses your uploaded image as the starting frame. You can also add an optional end image to guide the final frame.
Start with the visible subject and action, then add setting, camera movement, lighting, pacing, and sound. One connected shot is easier to evaluate and refine.
When you provide a source image, image-to-video follows that image’s output aspect ratio.
Completed MP4 videos remain in the workspace and in My Works, where you can preview or download them.
Write one clear shot prompt, choose a source image if needed, and create a real H3 Max video in the workspace.