Winged riders race across towering sunlit clouds as the camera tracks beside them, cinematic fantasy realism.
Kling 2.5 Turbo Text to video API
Generate a short video from text alone or animate a supplied opening frame with optional end-frame guidance.
Choose text-to-video for prompt-led composition, or image-to-video when a source frame should anchor the subject and framing.
Pro pricing
All listed rates are for Pro mode. Standard is currently unavailable.
- Pro price
- $0.042/s
Playground Examples
Try production-style prompt starters.
Similar Models
Compare adjacent capabilities before switching models.
Seedance 2.5
ByteDanceCreate videos up to 30 seconds, combine as many as 50 multimodal references, and refine specific moments with frame-level control.
Seedance 2.0
ByteDanceA multimodal video generation model for text prompts, first-frame animation, and reference-guided production clips.
Gemini Omni
GoogleCreate 4 to 10-second videos from text, up to seven reference images, or one short reference video with synchronized sound.
Model overview
Kling 2.5 Turbo text and image video generation
Kling 2.5 Turbo on sjolt provides separate text-to-video and image-to-video routes for 5 or 10-second Pro clips, with optional final-frame control for image animation.

Prompt-led text-to-video
Describe the subject, action, environment, camera movement, and style, then choose a landscape, portrait, or square composition.
- Supports Text to video and reference-asset generation workflows.
- Fits product assets, ad previews, visual direction, and social content tests.
- Adjust configured inputs, media, and switches before generation.

Start and end frame guidance
Image-to-video anchors the opening composition to one required image and can use a second image to guide the closing frame.

Pro mode
Generate in Pro mode with 5 or 10-second duration choices. Standard is visible in the controls but remains unavailable.
- Compare available models from Black Forest Labs, Google, MiniMax, OpenAI, ByteDance, Kuaishou, SJolt AI in one place.
- Copy the request body directly to the server to reduce frontend/backend parameter drift.
- Failure states, retries, result preview, and review checkpoints stay in the same workflow.
Model characteristics
Built for concise video creation with direct inputs.
Text to video
Generate a complete scene from prompt direction with landscape, portrait, or square framing.
Image to video
Animate one required opening image while the prompt directs subject, scene, and camera motion.
End-frame guidance
Add an optional closing image when the clip should transition toward a specific final composition.
Two durations
Choose a 5-second clip for a compact motion beat or 10 seconds for a longer action and camera progression.
FAQ
These are the first questions to answer when evaluating this model.
Which Kling 2.5 Turbo routes are available?
sjolt exposes separate text-to-video and image-to-video routes. Both currently generate in Pro mode.
Which generation modes are supported?
Pro is currently available. Standard appears in the selector but is disabled.
Which durations are supported?
Choose either 5 or 10 seconds.
Can I control the first and last frame?
Yes. Image-to-video requires image_url for the opening frame and accepts optional end_image_url for the closing frame.
Which image formats can I upload?
The Playground accepts one JPG, JPEG, PNG, or WebP start frame and one optional end frame.