Wan 3.0 Reference to Video API

alibaba / 3.0

Wan 3.0 creates directed 2-30 second video from coordinated image, motion, and audio references with adaptive framing and native sound.

Upload one or more reference assets, mention their roles in the prompt, then choose the output length, frame, and resolution.

Pricing

Output resolution

Final usage is the selected resolution rate multiplied by output duration.

480p
$0.04/s
720p
$0.08/s
1080p
$0.16/s
Input

Prompt, assets, and output parameters form one generation request.

Optionally describe the result and mention uploaded references as @Image, @Video, or @Audio. Maximum 20,000 characters. Use @ to reference uploaded files.
Input media

Upload the media inputs configured for this model.

Waiting for assetsUploaded files will appear here for preview and removal.
Waiting for assetsUploaded files will appear here for preview and removal.
Waiting for assetsUploaded files will appear here for preview and removal.
Aspect ratioUse Auto to follow the references, or choose a fixed landscape, square, or portrait frame.
ResolutionChoose 480p, 720p, or 1080p output.
Video durationChoose a whole-number output length from 2 to 30 seconds.
5s

Model overview

Wan 3.0 multimodal reference video generator and API

Wan 3.0 on sjolt provides a focused reference-to-video workflow for character continuity, product motion, visual development, campaign scenes, music-led sequences, and short-form storytelling. Coordinate images, MP4 clips, and audio tracks, then retrieve the completed MP4 through the shared asynchronous task API.

Wan 3.0 generation example
01 · Video generation

Coordinate image, motion, and audio references

Combine up to 10 images, 5 MP4 clips, and 5 audio tracks in one request. Assign every asset a clear role with @ mentions so subjects, visual language, camera movement, timing, dialogue, ambience, or music remain easy to interpret.

  • Supports Reference to Video and reference-asset generation workflows.
  • Fits product assets, ad previews, visual direction, and social content tests.
  • Adjust configured inputs, media, and switches before generation.
Reference asset control example
02 · Asset control

Direct motion and sound as one scene

Describe the opening state, subject action, camera path, environmental response, sound cues, and intended ending. Use concise temporal language to align visible events with dialogue, effects, ambience, or musical beats.

Reference imagesReference videosReference audio
Unified model API and result management example
03 · API integration

Flexible framing and asynchronous delivery

Use adaptive framing or choose 16:9, 9:16, 1:1, 4:3, 3:4, select 480p, 720p, or 1080p, and create a whole-number clip from 2 to 30 seconds. Submit once, then poll the task identifier or receive a webhook when the video is ready.

  • Compare available models from Alibaba, Black Forest Labs, Google, MiniMax, OpenAI, ByteDance, Kuaishou, SJolt AI, Suno in one place.
  • Copy the request body directly to the server to reduce frontend/backend parameter drift.
  • Failure states, retries, result preview, and review checkpoints stay in the same workflow.

Model characteristics

Built for reference-led video direction across visual identity, movement, timing, and sound.

01

Multimodal reference control

Use images for identity and composition, videos for motion and camera rhythm, and audio for dialogue, effects, ambience, or music direction.

02

Optional prompt direction

Generate from references alone or add a written production brief with explicit @ mentions for each uploaded asset.

03

Flexible delivery formats

Choose adaptive framing or 16:9, 9:16, 1:1, 4:3, 3:4, output from 480p through 1080p, and a whole-number duration from 2 to 30 seconds.

04

Unified task workflow

Use one provider-qualified create route, the shared task-status endpoint, and optional webhook delivery for repeatable production pipelines.

FAQ

These are the first questions to answer when evaluating this model.

Which Wan 3.0 workflow is available on sjolt?

The initial release exposes reference-to-video with image, video, and audio guidance. Dedicated text-to-video and image-to-video variants can be added under the same Wan 3.0 model namespace later.

Which reference files can I upload?

Upload up to 10 PNG, JPG, JPEG, or WebP images, 5 MP4 clips, and 5 MP3, WAV, M4A, AAC, or OGG audio tracks. At least one reference asset is required.

What are the media duration limits?

Reference videos may total at most 15 seconds, reference audio may total at most 15 seconds, and reference-video length plus output length cannot exceed 30 seconds.

Is a prompt required?

No. References can drive the request alone. When adding instructions, mention each asset as @Image, @Video, or @Audio and describe its intended role.

Which output settings are available?

Choose adaptive framing or 16:9, 9:16, 1:1, 4:3, 3:4, 480p, 720p, or 1080p, and a whole-number duration from 2 to 30 seconds.

How does the API workflow work?

Submit the provider-qualified reference route, keep the returned task identifier, then poll the shared task endpoint or receive a webhook. Successful tasks return the generated MP4 URL.