MiniMax H3 Dev Image to video API

minimax / H3 Dev

MiniMax H3 Dev creates directed 4-15 second video from text, first and last frames, or image references with optional audio guidance.

Use image-to-video to animate a required first frame toward an optional last frame, or choose text and reference workflows for broader direction.

Pricing

Output resolution

Final usage is the selected resolution rate multiplied by output duration.

768p
$0.021/s
Input

Prompt, assets, and output parameters form one generation request.

Describe the motion and transition from the first frame toward the optional last frame in up to 7,000 characters.
Input media

Upload the media inputs configured for this model.

ResolutionGenerate 768p output.
Video durationChoose a fixed 4-15 second output.
5s

Playground Examples

Try production-style prompt starters.

Preserve the masked courier, scarlet ceramic armor, six-legged chrome panther, canyon, and color palette from @Image1. The panther launches into a sprint as the courier vaults onto its back; both race straight toward the luminous sandstorm while the camera drops to ground level and accelerates beside them. Stable identities and anatomy, realistic weight, cinematic impact, no text, subtitles, logos, or watermark.

Preserve the percussionist, obsidian drum ring, volcanic arena, and symmetrical composition from @Image1. Time the central strike and the expanding pressure wave to @Audio1; the surrounding drums answer in sequence as lava fissures flare and the camera punches forward through flying embers. Stable scene geometry, no text, subtitles, logos, or watermark.

Model overview

MiniMax H3 Dev text, image, and reference video generator and API

MiniMax H3 Dev is a 768p video model for product motion, character scenes, visual development, social clips, and audiovisual concepts. Start from a written prompt, animate a first frame toward an optional last frame, or combine image references with optional audio direction.

MiniMax H3 Dev generation example
01 · Video generation

Text-led and image-anchored generation

Use text-to-video when the scene should begin from a written brief. Use image-to-video for a controlled first-frame animation, or switch to reference-to-video and add up to 5 images to establish the subject, product, materials, palette, styling, environment, or composition.

  • Supports Image to video and reference-asset generation workflows.
  • Fits product assets, ad previews, visual direction, and social content tests.
  • Adjust configured inputs, media, and switches before generation.
Reference asset control example
02 · Asset control

First and last frame direction

Image-to-video requires one opening image and accepts an optional closing image. The prompt directs subject action, camera motion, environmental response, and the transition between the two compositions while framing follows the first frame.

First frameLast frame
Unified model API and result management example
03 · API integration

Optional audio-guided timing

Add up to three audio references and mention them beside the beats, transitions, impacts, dialogue, ambience, or musical cues they should guide. Submit once, then poll the task identifier or receive a webhook when the generated video is ready.

  • Compare available models from Black Forest Labs, Google, MiniMax, OpenAI, ByteDance, Kuaishou, SJolt AI, Suno in one place.
  • Copy the request body directly to the server to reduce frontend/backend parameter drift.
  • Failure states, retries, result preview, and review checkpoints stay in the same workflow.

Model characteristics

Built for focused prompt-led, frame-guided, or reference-guided experiments with explicit visual and audio direction.

01

Three focused workflows

Use text-to-video without source files, image-to-video with first and optional last frames, or reference-to-video when broader visual anchors should guide the result.

02

First and last frames

Anchor the opening composition with one required image and optionally guide the exact closing composition with a second image. Output framing follows the first frame.

03

Optional sound direction

Use up to 3 audio tracks to guide rhythm, dialogue, ambience, effects, or musical timing. Each track must be 2-15 seconds, with at most 15 seconds combined.

04

Flexible delivery formats

Create 4-15 second clips at 768p. Text and reference workflows support 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, while image-to-video follows the first frame.

05

Asynchronous API workflow

Submit a text-led or reference-guided task, keep the returned identifier, and use polling or a webhook to collect the finished video.

FAQ

These are the first questions to answer when evaluating this model.

What is MiniMax H3 Dev?

MiniMax H3 Dev is a 768p video model on sjolt. It supports prompt-only creation, first-and-last-frame animation, and image-reference direction through three dedicated routes.

Which workflows are available?

Use text-to-video without source files, image-to-video with a required first frame and optional last frame, or reference-to-video with one or more images and optional audio tracks.

Can I control the first and last frames?

Yes. Image-to-video requires image_url for the first frame and accepts end_image_url for an optional final frame. Its output aspect ratio follows the first frame.

How many references can I use?

Reference-to-video requires 1-5 images and can include up to 3 optional audio tracks.

How should I write the prompt?

Identify the role of each image or audio reference, then describe subject motion, camera movement, lighting, environmental reactions, sound cues, and the intended ending state.

Which formats and durations are available?

Choose 768p and a whole-number duration from 4 to 15 seconds. Text and reference workflows support 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, while image-to-video follows the first frame.

How does the API workflow work?

Submit the selected route with its required prompt and duration plus optional resolution. Add image_url and end_image_url only on image-to-video, or image_urls and audio_urls only on reference-to-video, then poll the returned task identifier or use a webhook.