Gemini Omni 1.1 Flash Text to video API

google / 1.1 Flash

Gemini Omni 1.1 Flash turns production prompts into short landscape or portrait clips at a selected output resolution.

Use text alone, or add optional image and video references to guide subjects, composition, visual language, or motion.

Pricing

Output resolution

Final usage is the selected resolution rate multiplied by output duration.

360p
$0.03/s
720p
$0.10/s
1080p
$0.15/s
4K
$0.30/s
Input

Prompt, assets, and output parameters form one generation request.

Describe the scene, action, camera direction, visual treatment, and sound cues.
Aspect ratioChoose landscape or portrait framing.
DurationChoose a 4, 6, 8, or 10-second output.
ResolutionChoose 360p, 720p, 1080p, or 4K output.

Model overview

Gemini Omni 1.1 Flash video generator and API

Gemini Omni 1.1 Flash on sjolt provides text-to-video and reference-to-video workflows for 4, 6, 8, 10-second clips. Choose 16:9 or 9:16 framing, select one of four output resolutions, and use the same inputs through the asynchronous task API.

Gemini Omni 1.1 Flash generation example
01 · Video generation

Four output resolution choices

Select 360p, 720p, 1080p, or 4K for each generation. API requests use 360p, 720p, 1080p, and 4k, while the interface presents the highest tier as 4K.

  • Supports Text to video and reference-asset generation workflows.
  • Fits product assets, ad previews, visual direction, and social content tests.
  • Adjust configured inputs, media, and switches before generation.
Reference asset control example
02 · Asset control

Multimodal reference guidance

Add up to 7 PNG, JPG, JPEG, or WebP images and optionally one MP4 up to 10 seconds. References can guide subject appearance, composition, visual treatment, or motion while the prompt defines the intended scene.

Prompt-led generation
Unified model API and result management example
03 · API integration

Short-video API workflow

Submit a text or reference request for 4, 6, 8, 10 seconds, keep the returned task ID, and query task status until the generated video URL is available. An optional callback URL can receive the terminal task result.

  • Compare available models from Alibaba, Black Forest Labs, Google, MiniMax, OpenAI, ByteDance, Kuaishou, SJolt AI, Suno, xAI in one place.
  • Copy the request body directly to the server to reduce frontend/backend parameter drift.
  • Failure states, retries, result preview, and review checkpoints stay in the same workflow.

Model characteristics

Designed for prompt-led and reference-guided short video with explicit output controls.

01

Text-led generation

Direct the scene, action, camera, visual treatment, and sound cues from one required prompt.

02

Optional references

Guide the result with up to 7 images and one MP4 no longer than 10 seconds.

03

Output controls

Choose 16:9 or 9:16 framing, 4, 6, 8, 10 seconds, and one of four resolutions.

04

Asynchronous delivery

Submit a task, poll its status, and retrieve generated video URLs from the successful result.

FAQ

These are the first questions to answer when evaluating this model.

What is Gemini Omni 1.1 Flash?

Gemini Omni 1.1 Flash is a Google video model available through separate text-to-video and reference-to-video task routes on sjolt.

Which routes are available for Gemini Omni 1.1 Flash?

Use text-to-video for prompt-only generation or reference-to-video when optional images or a short MP4 should guide the result.

Are references required on the reference route?

No. The reference route accepts a prompt without media, matching the optional-reference workflow in the Playground.

Which reference files can I use?

Use up to 7 PNG, JPG, JPEG, or WebP images and one MP4 no longer than 10 seconds.

Which output resolutions are supported?

The available output resolutions are 360p, 720p, 1080p, and 4K. API requests use `4k` for the 4K option.

Which durations and aspect ratios are supported?

Choose 4, 6, 8, 10 seconds and 16:9 or 9:16 framing.

How does the Gemini Omni 1.1 Flash API workflow work?

Submit a request with a prompt and supported settings, poll the returned task ID, then read generated video URLs from a successful result.