MiniMax H3 Dev Reference to Video API

minimax / H3 Dev

MiniMax H3 Dev creates directed 4-15 second video from a prompt alone or from image references with optional audio guidance.

Use text-to-video for prompt-led scenes, or reference-to-video to anchor identity, product details, styling, or composition with up to five images.

Pricing

Output resolution

Final usage is the selected resolution rate multiplied by output duration.

768p
$0.021/s
Input

Prompt, assets, and output parameters form one generation request.

Describe the result and mention uploaded references as @Image or @Audio. Use @ to reference uploaded files.
Input media

Upload the media inputs configured for this model.

Waiting for assetsUploaded files will appear here for preview and removal.
Waiting for assetsUploaded files will appear here for preview and removal.
ResolutionGenerate 768p output.
Aspect ratioChoose a cinematic, landscape, square, or vertical frame.
Video durationChoose a fixed 4-15 second output.
5s

Playground Examples

Try production-style prompt starters.

Preserve the masked courier, scarlet ceramic armor, six-legged chrome panther, canyon, and color palette from @Image1. The panther launches into a sprint as the courier vaults onto its back; both race straight toward the luminous sandstorm while the camera drops to ground level and accelerates beside them. Stable identities and anatomy, realistic weight, cinematic impact, no text, subtitles, logos, or watermark.

Preserve the percussionist, obsidian drum ring, volcanic arena, and symmetrical composition from @Image1. Time the central strike and the expanding pressure wave to @Audio1; the surrounding drums answer in sequence as lava fissures flare and the camera punches forward through flying embers. Stable scene geometry, no text, subtitles, logos, or watermark.

Model overview

MiniMax H3 Dev text and reference video generator and API

MiniMax H3 Dev is a separate self-host video model for product motion, character scenes, visual development, social clips, and audiovisual concepts. Start from a written prompt or combine image references with optional audio direction before asynchronous delivery through the sjolt task API.

MiniMax H3 Dev generation example
01 · Video generation

Text-led and image-anchored generation

Use text-to-video when the scene should begin from a written brief. Switch to reference-to-video and add up to 5 images to establish the subject, product, materials, palette, styling, environment, or composition.

  • Supports Reference to Video and reference-asset generation workflows.
  • Fits product assets, ad previews, visual direction, and social content tests.
  • Adjust configured inputs, media, and switches before generation.
Reference asset control example
02 · Asset control

Prompted action and camera direction

Describe the movement as a sequence: opening state, subject action, camera path, environmental response, and ending frame. Concrete motion verbs and a restrained shot plan help keep the result readable across cinematic, square, and vertical formats.

Reference imagesReference audio
Unified model API and result management example
03 · API integration

Optional audio-guided timing

Add up to three audio references and mention them beside the beats, transitions, impacts, dialogue, ambience, or musical cues they should guide. Submit once, then poll the task identifier or receive a webhook when the generated video is ready.

  • Compare available models from Black Forest Labs, Google, MiniMax, OpenAI, ByteDance, Kuaishou, SJolt AI in one place.
  • Copy the request body directly to the server to reduce frontend/backend parameter drift.
  • Failure states, retries, result preview, and review checkpoints stay in the same workflow.

Model characteristics

Built for focused prompt-led or image-guided experiments with explicit visual and audio direction.

01

Two focused workflows

Use text-to-video without reference files, or use reference-to-video when concrete visual anchors should guide the result.

02

Optional sound direction

Use up to 3 audio tracks to guide rhythm, dialogue, ambience, effects, or musical timing.

03

Flexible delivery formats

Create 4-15 second clips at 768p in 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 for cinematic, landscape, square, or portrait use.

04

Asynchronous API workflow

Submit a text-led or reference-guided task, keep the returned identifier, and use polling or a webhook to collect the finished video.

FAQ

These are the first questions to answer when evaluating this model.

What is MiniMax H3 Dev?

MiniMax H3 Dev is a separate self-host video model on sjolt. It supports prompt-only creation and image-reference direction through two dedicated routes.

Which workflows are available?

Use text-to-video without reference files. Use reference-to-video with one or more images and optional audio tracks.

How many references can I use?

Reference-to-video requires 1-5 images and can include up to 3 optional audio tracks.

How should I write the prompt?

Identify the role of each image or audio reference, then describe subject motion, camera movement, lighting, environmental reactions, sound cues, and the intended ending state.

Which formats and durations are available?

Choose 768p, 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, and a whole-number duration from 4 to 15 seconds.

How does the API workflow work?

Submit the selected route with its required prompt and duration plus optional resolution and framing. Add image URLs only on reference-to-video, poll the returned task identifier or use a webhook, then read the generated video URL from a successful result.