Preserve the masked courier, scarlet ceramic armor, six-legged chrome panther, canyon, and color palette from @Image1. The panther launches into a sprint as the courier vaults onto its back; both race straight toward the luminous sandstorm while the camera drops to ground level and accelerates beside them. Stable identities and anatomy, realistic weight, cinematic impact, no text, subtitles, logos, or watermark.
MiniMax H3 Dev Reference to Video API
MiniMax H3 Dev creates directed 4-15 second video from a prompt alone or from image references with optional audio guidance.
Use text-to-video for prompt-led scenes, or reference-to-video to anchor identity, product details, styling, or composition with up to five images.
Output resolution
Final usage is the selected resolution rate multiplied by output duration.
- 768p
- $0.021/s
Playground Examples
Try production-style prompt starters.
Preserve the percussionist, obsidian drum ring, volcanic arena, and symmetrical composition from @Image1. Time the central strike and the expanding pressure wave to @Audio1; the surrounding drums answer in sequence as lava fissures flare and the camera punches forward through flying embers. Stable scene geometry, no text, subtitles, logos, or watermark.
Similar Models
Compare adjacent capabilities before switching models.
Seedance 2.5
ByteDanceCreate videos up to 30 seconds, combine as many as 50 multimodal references, and refine specific moments with frame-level control.
Seedance 2.0
ByteDanceA multimodal video generation model for text prompts, first-frame animation, and reference-guided production clips.
Gemini Omni
GoogleCreate 4 to 10-second videos from text, up to seven reference images, or one short reference video with synchronized sound.
Model overview
MiniMax H3 Dev text and reference video generator and API
MiniMax H3 Dev is a separate self-host video model for product motion, character scenes, visual development, social clips, and audiovisual concepts. Start from a written prompt or combine image references with optional audio direction before asynchronous delivery through the sjolt task API.

Text-led and image-anchored generation
Use text-to-video when the scene should begin from a written brief. Switch to reference-to-video and add up to 5 images to establish the subject, product, materials, palette, styling, environment, or composition.
- Supports Reference to Video and reference-asset generation workflows.
- Fits product assets, ad previews, visual direction, and social content tests.
- Adjust configured inputs, media, and switches before generation.

Prompted action and camera direction
Describe the movement as a sequence: opening state, subject action, camera path, environmental response, and ending frame. Concrete motion verbs and a restrained shot plan help keep the result readable across cinematic, square, and vertical formats.

Optional audio-guided timing
Add up to three audio references and mention them beside the beats, transitions, impacts, dialogue, ambience, or musical cues they should guide. Submit once, then poll the task identifier or receive a webhook when the generated video is ready.
- Compare available models from Black Forest Labs, Google, MiniMax, OpenAI, ByteDance, Kuaishou, SJolt AI in one place.
- Copy the request body directly to the server to reduce frontend/backend parameter drift.
- Failure states, retries, result preview, and review checkpoints stay in the same workflow.
Model characteristics
Built for focused prompt-led or image-guided experiments with explicit visual and audio direction.
Two focused workflows
Use text-to-video without reference files, or use reference-to-video when concrete visual anchors should guide the result.
Optional sound direction
Use up to 3 audio tracks to guide rhythm, dialogue, ambience, effects, or musical timing.
Flexible delivery formats
Create 4-15 second clips at 768p in 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 for cinematic, landscape, square, or portrait use.
Asynchronous API workflow
Submit a text-led or reference-guided task, keep the returned identifier, and use polling or a webhook to collect the finished video.
FAQ
These are the first questions to answer when evaluating this model.
What is MiniMax H3 Dev?
MiniMax H3 Dev is a separate self-host video model on sjolt. It supports prompt-only creation and image-reference direction through two dedicated routes.
Which workflows are available?
Use text-to-video without reference files. Use reference-to-video with one or more images and optional audio tracks.
How many references can I use?
Reference-to-video requires 1-5 images and can include up to 3 optional audio tracks.
How should I write the prompt?
Identify the role of each image or audio reference, then describe subject motion, camera movement, lighting, environmental reactions, sound cues, and the intended ending state.
Which formats and durations are available?
Choose 768p, 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, and a whole-number duration from 4 to 15 seconds.
How does the API workflow work?
Submit the selected route with its required prompt and duration plus optional resolution and framing. Add image URLs only on reference-to-video, poll the returned task identifier or use a webhook, then read the generated video URL from a successful result.