Seedance 2.5 Preview Image to video API

bytedance / seedance-2.5-preview

Generate up to 30 seconds from text, first/last frames, or as many as 50 multimodal references.

Start from text, animate a first frame with an optional ending frame, or combine images, video, and audio in the reference workflow.

Pricing

Output resolution

Final usage is the selected resolution rate multiplied by output duration.

480p
$0.15/s
720p
$0.30/s
Input

Prompt, assets, and output parameters form one generation request.

Optionally describe the motion, camera movement, scene changes, and sound you want.
Input media

Upload the media inputs configured for this model.

Aspect ratioChoose a landscape, portrait, square, or cinematic output frame.
ResolutionChoose 480p for drafts or 720p for higher-detail output.
DurationChoose an output from 4 to 30 seconds.
5s

Model overview

Seedance 2.5 API (Preview) for Multimodal Video Generation

Seedance 2.5 Preview on sjolt exposes text-to-video, image-to-video, and reference-to-video routes for 4 to 30-second output. Start from text, animate a first frame with an optional last frame, or guide a generation with up to 30 images, 10 MP4 clips, and 10 audio references through the asynchronous task API.

01 · Video generation

Up to 30 seconds in one generation

Generate any whole-second duration from 4 through 30. Use the longer timeline to build a complete scene with an opening, development, camera progression, synchronized sound, and a deliberate ending instead of stitching together several short clips.

  • Supports Image to video and reference-asset generation workflows.
  • Fits product assets, ad previews, visual direction, and social content tests.
  • Adjust configured inputs, media, and switches before generation.
02 · Asset control

Up to 50 multimodal references

Combine up to 30 images, 10 MP4 clips, and 10 audio files in one reference-guided request. Assign each source a role with @ mentions to preserve characters and products, borrow camera movement, match visual style, and coordinate dialogue, music, or ambience.

First frameLast frame
03 · API integration

Frame-level editing control

Describe changes at specific moments and anchor them with reference frames, clips, and audio cues. Target a character, action, prop, background, or sound beat while helping the surrounding identity, motion, lighting, and audiovisual continuity remain coherent.

  • Compare available models from Google, MiniMax, OpenAI, ByteDance, Kuaishou, SJolt AI in one place.
  • Copy the request body directly to the server to reduce frontend/backend parameter drift.
  • Failure states, retries, result preview, and review checkpoints stay in the same workflow.

Model characteristics

Three production-focused advances for longer, more controllable video creation.

01

Thirty-second storytelling

Plan a complete narrative or visual arc with enough time for scene development, measured camera movement, transitions, and a resolved ending.

02

50 reference assets

Bring together 30 images, 10 videos, and 10 audio files for identity, style, motion, camera, dialogue, and sound guidance.

03

Precise local refinement

Use time-ordered instructions and multimodal anchors to focus changes on a particular moment or visual element while maintaining continuity around it.

FAQ

These are the first questions to answer when evaluating this model.

What is Seedance 2.5 Preview?

Seedance 2.5 Preview is a ByteDance video generation model with separate routes for prompt-only generation, first/last-frame animation, and multimodal reference guidance.

Which input files can I upload?

Upload up to 30 PNG, JPG, JPEG, or WebP images, 10 MP4 videos, and 10 MP3, WAV, M4A, AAC, or OGG audio files.

Are reference files required?

The text-to-video route does not accept reference files. Image-to-video requires a first frame and accepts an optional last frame. Reference-to-video requires at least one image, video, or audio URL.

Which output settings are available?

Choose a whole-second duration from 4 to 30, 480p or 720p output, and 16:9, 9:16, 1:1, 3:4, 4:3, 21:9 framing.

How should I write a multimodal prompt?

State what must remain stable, assign each uploaded reference a role, then describe the action, camera movement, scene changes, lighting, and sound in chronological order.

How does the API workflow work?

Submit the text-only, image-guided, or reference-guided request to its dedicated endpoint. Poll the returned task ID, then read the generated MP4 URL from the successful task result.