A microscopic cybernetic hummingbird with translucent crystal wings bursts through a suspended wall of rain at night. The camera races beside it at wing level as molten-gold circuitry ignites across every feather; it corkscrews around a lightning bolt, scattering water into a radiant shockwave, then stops inches from camera with wings blazing. Photoreal cinematic macro, stable anatomy, physically convincing motion, aggressive speed ramp, deep thunder, crystalline wing chimes, rushing rain, no text, subtitles, logos, brands, or watermark.
MiniMax H3 Reference to Video API
MiniMax H3 creates fixed 2K video with prompt-directed action, cinematic framing, native audio, and a reference workflow for images, motion clips, and sound.
Use text for rapid concept development, or coordinate uploaded media with explicit @ mentions when identity, movement, timing, or audio direction must stay anchored.
Output resolution
Final charge is the rate multiplied by output duration.
- 2K
- $0.13/s
Playground Examples
Try production-style prompt starters.
Preserve the masked courier, scarlet ceramic armor, six-legged chrome panther, canyon, and color palette from @Image1. The panther launches into a sprint as the courier vaults onto its back; both race straight toward the luminous sandstorm while the camera drops to ground level and accelerates beside them. The storm wall splits around a black rock arch in a violent burst of salt and light. Stable identities and anatomy, realistic weight, cinematic impact, roaring wind, metal footfalls, no text, subtitles, logos, or watermark.
Preserve the percussionist, obsidian drum ring, volcanic arena, and symmetrical composition from @Image1. The mallets strike once in perfect unison; a visible circular pressure wave tears across the ash, every giant drum answers in sequence, lava fissures flare, and lightning crashes into the center as the camera punches forward through flying embers. Monumental synchronized bass, sharp stone impacts, thunder and volcanic rumble, stable scene geometry, no text, subtitles, logos, or watermark.
Similar Models
Compare adjacent capabilities before switching models.
Seedance 2.0
ByteDanceA multimodal video generation model for text prompts, first-frame animation, and reference-guided production clips.
Gemini Omni
GoogleCreate 4 to 10-second videos from text, up to seven reference images, or one short reference video with synchronized sound.
Kling 3.0
KuaishouVideo generation with 3-15 second duration control, multi-shot storytelling, native audio, and start/end frames.
Model overview
MiniMax H3 2K multimodal video generator and API
MiniMax H3 brings fixed 2K output, native audio, and prompt or reference-guided generation into one production workflow. Build 5-15 second cinematic scenes, product films, character moments, campaign assets, and social cuts, then automate repeated creation through the asynchronous sjolt task API.

Fixed 2K output for production-ready detail
Compose wide establishing shots, product close-ups, character performances, and vertical social scenes with consistent 2K output. Six fixed aspect ratios cover cinematic, landscape, square, and portrait delivery, while auto can follow source framing on the reference route.
- Supports Reference to Video and reference-asset generation workflows.
- Fits product assets, ad previews, visual direction, and social content tests.
- Adjust configured inputs, media, and switches before generation.

Multimodal references with explicit direction
Guide subjects, products, styling, composition, movement, editing rhythm, or soundtrack with up to 9 images, 3 MP4 clips, and 3 audio tracks. Mention each asset in the prompt so its role is unambiguous. Audio must accompany an image or video reference rather than serve as the only source.

Native audio and asynchronous delivery
Direct dialogue, ambience, effects, and music alongside visible action so the audiovisual concept develops as one scene. Submit a task, poll its identifier or receive a webhook, and collect the generated MP4 from the completed result for editing, publishing, or downstream automation.
- Compare available models from Google, MiniMax, OpenAI, ByteDance, Kuaishou, SJolt AI in one place.
- Copy the request body directly to the server to reduce frontend/backend parameter drift.
- Failure states, retries, result preview, and review checkpoints stay in the same workflow.
Model characteristics
Designed for detailed 2K production with coordinated visual, motion, and sound direction.
Prompt-led video
Direct subjects, action, environment, camera movement, lighting, dialogue, ambience, and effects from one written production brief.
Multimodal reference control
Combine images, short MP4 clips, and audio tracks to anchor identity, visual language, motion, pacing, and soundtrack direction.
Flexible 5-15 second formats
Choose 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 for planned delivery, or use automatic framing with references.
API-ready workflow
Use separate text and reference routes with task polling and webhook delivery for prototypes, creative tools, or repeatable content pipelines.
FAQ
These are the first questions to answer when evaluating this model.
What is MiniMax H3?
MiniMax H3 is a video generation model for fixed 2K clips with native audio. It supports prompt-only creation and multimodal reference direction through separate sjolt API routes.
Which MiniMax H3 workflows are available?
Use text-to-video for a prompt-only request. Use reference-to-video when images, videos, or audio should guide identity, style, motion, pacing, composition, or sound.
Which reference files can I upload?
Upload up to 9 PNG, JPG, JPEG, or WebP images, 3 MP4 videos, and 3 MP3, WAV, M4A, AAC, or OGG audio tracks.
What are the reference duration limits?
Each video or audio item must be 2-15 seconds. Combined video length cannot exceed 15 seconds, and combined audio length cannot exceed 15 seconds. Audio cannot be the only reference.
Does MiniMax H3 create sound?
Yes. Describe dialogue, ambience, sound effects, or music in the prompt and align each cue with the corresponding subject, action, or moment.
Which formats and durations are available?
Text generation supports 21:9, 16:9, 4:3, 1:1, 3:4, 9:16. Reference generation also supports automatic framing. Output duration ranges from 5 to 15 seconds.
How does the MiniMax H3 API workflow work?
Submit the selected route with a required prompt and its supported fields. Poll the returned task identifier or use a webhook, then read the generated MP4 URL from a successful result.