Gemini Omni Reference to Video API

google / Omni

Gemini Omni creates short-form video with prompt-directed motion, reference guidance, landscape or portrait framing, and synchronized audio.

Start from text alone or add image and video references to preserve a subject, visual language, composition, or motion direction.

Pricing

Per-second comparison

Final usage is the sjolt rate multiplied by the selected output duration.

sjolt
$0.05/s
Input

Prompt, assets, and output parameters form one generation request.

Describe the scene, action, camera movement, style, and sound you want. Use @ to reference uploaded files.
Input media

Upload the media inputs configured for this model.

Waiting for assetsUploaded files will appear here for preview and removal.
Waiting for assetsUploaded files will appear here for preview and removal.
Aspect ratioChoose landscape or portrait output.
DurationChoose a 4, 6, 8, or 10-second output.

Playground Examples

Try production-style prompt starters.

Preserve the supplied performer and facial identity through a rapid studio fashion montage. Shift the wardrobe and background between leopard print under hard black light, black lace against electric blue, cream fabric against pink, and warm plaid against soft gray. Use a close fisheye camera, confident direct-to-camera poses, stable anatomy, clean transitions, and synchronized electronic ambience. No text, subtitles, logos, brands, or watermark.

Two martial artists duel on a barren moonlit plain beneath a black sky. One wears flowing white robes with long silver hair; the other wears layered charcoal robes. Track their parries, spinning kicks, leaps, and close-range hand strikes in a continuous medium-wide shot with realistic cloth and hair motion. Audio: fabric snaps, foot impacts, controlled breaths, and low wind. No text, subtitles, logos, brands, or watermark.

Two longtime friends in tailored navy and cream jackets meet for lunch on an elegant coastal restaurant terrace. Begin as they greet beside a white table, cut to the navy-jacketed man sitting down, show a top view of tomato soup being served, then alternate natural conversational close-ups with the ocean behind them. Bright Mediterranean daylight, realistic gestures, coherent table details, and synchronized dialogue and dining ambience. No text, subtitles, logos, brands, or watermark.

Model overview

Gemini Omni multimodal video generator and API

Gemini Omni on sjolt provides text-to-video and reference-to-video workflows for 4, 6, 8, 10-second clips. Generate in 16:9 or 9:16, guide a scene with up to 7 images or one short MP4, then use the same request shape through the asynchronous task API.

01 · Video generation

Prompt-led scenes with coordinated performance

Describe performers, staging, wardrobe, instruments or props, camera movement, lighting, and sound in one production brief. Gemini Omni can coordinate several subjects through a shared scene while maintaining the main action. Write the prompt in shot order and tie each sound cue to the visible performer or event.

  • Supports Reference to Video and reference-asset generation workflows.
  • Fits product assets, ad previews, visual direction, and social content tests.
  • Adjust configured inputs, media, and switches before generation.
02 · Asset control

Image and video reference guidance

Add up to 7 PNG, JPG, JPEG, or WebP references for subject, style, and composition. You can also supply one MP4 up to 10 seconds for motion or scene guidance. Use clear @Image and @Video mentions in the Playground prompt when several references serve different roles.

Reference imagesReference video
03 · API integration

Native sound, flexible framing, and API delivery

Direct ambience, effects, and music alongside the visual action, then choose 16:9 for landscape delivery or 9:16 for vertical social formats. Select 4, 6, 8, 10 seconds, submit the matching API request, poll its task ID, and retrieve the MP4 URL after success.

  • Compare available models from Google, OpenAI, ByteDance, Kuaishou, SJolt AI in one place.
  • Copy the request body directly to the server to reduce frontend/backend parameter drift.
  • Failure states, retries, result preview, and review checkpoints stay in the same workflow.

Model characteristics

Built for prompt-led and reference-guided short video with coordinated visuals, motion, and sound.

01

Text to video

Direct subjects, action, environment, composition, camera, lighting, and sound from one prompt.

02

Multimodal references

Use up to 7 images and one short MP4 to guide identity, style, motion, or scene structure.

03

Flexible short-form output

Choose 4, 6, 8, 10 seconds with 16:9 or 9:16 framing.

04

Native audio and delivery

Prompt synchronized ambience and effects, then retrieve the generated MP4 from the task result.

FAQ

These are the first questions to answer when evaluating this model.

What is Gemini Omni?

Gemini Omni is Google's multimodal creation model for video generation and reference-guided scene transformation. sjolt exposes text and reference routes through one asynchronous task API.

Which Gemini Omni routes are available?

Use text-to-video for a prompt-only request or reference-to-video when images or an MP4 should guide the result.

Which reference files can I upload?

The reference route accepts up to 7 PNG, JPG, JPEG, or WebP images and one MP4 up to 10 seconds.

Does Gemini Omni generate sound?

Yes. Include desired ambience, sound effects, or music in the prompt and align each sound cue with the corresponding visual action.

How should I write a reference-guided prompt?

Identify the role of each reference, list the visual traits that must remain stable, then describe the new action, environment, camera movement, and sound in sequence.

Which output settings are available?

Choose 16:9 or 9:16 and an output duration of 4, 6, 8, 10 seconds.

How does the Gemini Omni API workflow work?

Submit the selected route with a prompt and optional supported references. Poll the returned task ID, then read the generated MP4 URL from a successful response.