Grok Imagine Video 1.5 Image to video API

xai / Video 1.5

Create short video with synchronized audio from text, animate one source image, or guide a new scene with several visual references.

Choose a dedicated route for prompt-only creation, single-image animation, or broader visual reference direction.

Pricing

Output duration

Final usage is the per-second rate multiplied by the selected output duration.

All variants
$0.0225/s
Input

Prompt, assets, and output parameters form one generation request.

Describe how the supplied image should move, how the camera should behave, and which sounds should accompany the scene in up to 4,096 characters.
Input media

Upload the media inputs configured for this model.

ResolutionGenerate 720p output.
DurationChoose a whole-number duration from 1 to 15 seconds.
8s

Model overview

Grok Imagine Video 1.5 text, image, and reference video API

Grok Imagine Video 1.5 on sjolt supports three focused workflows for 1-15 second clips. Start with text, animate one source image, or guide a scene with 2-7 images while directing visuals and synchronized sound in one prompt.

01 · Video generation

Prompt-led motion and camera direction

Describe subjects, action, environment, shot order, camera movement, lighting, dialogue, ambience, and sound effects. Text-to-video keeps framing independently selectable, while image-led generation uses the supplied visual anchors to stabilize the scene.

  • Supports Image to video and reference-asset generation workflows.
  • Fits product assets, ad previews, visual direction, and social content tests.
  • Adjust configured inputs, media, and switches before generation.
02 · Asset control

Single-image and multi-reference workflows

Use image-to-video when one source image should define the subject and output framing. Use reference-to-video with 2-7 images when identity, wardrobe, product details, palette, styling, or composition should be combined across several visual anchors.

Source image
03 · API integration

Native audio and asynchronous delivery

Direct dialogue, ambience, effects, and music alongside the visual action. Choose 720p, select a whole-number duration, submit the matching task route, then poll its identifier or receive a webhook when the MP4 is ready.

  • Compare available models from Alibaba, Black Forest Labs, Google, MiniMax, OpenAI, ByteDance, Kuaishou, SJolt AI, Suno, xAI in one place.
  • Copy the request body directly to the server to reduce frontend/backend parameter drift.
  • Failure states, retries, result preview, and review checkpoints stay in the same workflow.

Model characteristics

Built for prompt-only scenes, source-image animation, and multi-reference video with coordinated sound.

01

Three dedicated routes

Use text-to-video without media, image-to-video with exactly one source image, or reference-to-video with two to seven visual references.

02

Up to seven references

Combine as many as 7 JPG, JPEG, PNG, or WebP images and identify their roles with @Image mentions in the Playground prompt.

03

Flexible short-form output

Create 1-15 second clips at 720p with automatic, landscape, portrait, square, 3:2, or 2:3 framing where the route supports it.

04

Synchronized sound

Describe speech, ambience, movement, effects, and musical cues in the same prompt so audio direction stays tied to visible events.

FAQ

These are the first questions to answer when evaluating this model.

Which Grok Imagine Video 1.5 routes are available?

Use text-to-video for a prompt-only scene, image-to-video for exactly one source image, or reference-to-video for two to seven visual references.

What is the difference between image-to-video and reference-to-video?

Image-to-video treats exactly one image as the source composition and follows its framing. Reference-to-video accepts two to seven images as broader identity, style, subject, or composition guidance and supports explicit framing.

Which image files are supported?

The Playground accepts JPG, JPEG, PNG, and WebP images up to 20 MB each. Image-to-video accepts exactly one image, while reference-to-video requires two to seven.

Which output settings can I select?

Choose 720p and any whole-number duration from 1 to 15 seconds. Text and reference routes also expose auto, 1:1, 16:9, 9:16, 3:2, 2:3 framing.

Does the model generate audio?

Yes. Include dialogue, ambience, sound effects, movement sounds, or music in the prompt and align each cue with the visible action.

How does task delivery work?

Submit one of the three routes, store the returned task identifier, then poll the shared status endpoint or provide a top-level webhook to receive the completed video URL.