DeepSeek V4 Pro Text generation API

deepseek / deepseek-v4-pro

A DeepSeek model for text generation and reasoning tasks that need more capability.

Use deepseek/deepseek-v4-pro through the sjolt LLM API.

Model IDdeepseek/deepseek-v4-pro
Pricing

Token rates

Input
$1.056 / 1M tokens
Cache read
$0.0352 / 1M tokens
Output
$3.168 / 1M tokens
DeepSeek V4 Prodeepseek/deepseek-v4-pro

Start a conversation

Write a message below to begin.

Model overview

DeepSeek V4 Pro for demanding production reasoning

Take on complex reasoning, production agentic systems, and coding work that benefits from a higher-capability model and deliberate analysis.

DeepSeek V4 Pro text generation model cover
01 · Text generation

Engineer deeper agentic systems

Give production agents room to plan, call tools, inspect results, and revise complex implementations across longer chains of work.

  • Send one required user prompt with optional system instructions.
  • Review the exact non-streaming OpenAI Chat request before running it.
  • Inspect assistant text, finish reason, response ID, and token usage together.
DeepSeek V4 Pro long-context text generation model cover
02 · Request controls

Reason across extensive context

Bring large codebases, technical records, requirements, and prior decisions into a 1M-token context window for connected analysis.

User promptSystem promptNon-streaming response
DeepSeek V4 Pro API compatibility model cover
03 · API integration

Fit established API stacks

Integrate with OpenAI Chat Completions, OpenAI Responses, or Anthropic Messages while keeping one sjolt endpoint.

  • Compare enabled DeepSeek text models in one place.
  • Copy the request preview to keep application requests aligned with the Playground.
  • Handle assistant responses and provider availability states in the same workflow.

Model characteristics

Better for creative validation than isolated one-off generation.

01

Workload focus

Complex reasoning, production agentic systems, and demanding coding tasks.

02

Context window

Up to 1,000,000 input tokens.

03

Maximum API output

Up to 384,000 tokens.

04

Reasoning modes

Disable thinking when it is not needed, or select low, high, or max effort.

FAQ

These are the first questions to answer when evaluating this model.

When should I choose DeepSeek V4 Pro?

Choose it for complex reasoning, production agentic workflows, or coding tasks where additional model capability is more important than the lighter Flash profile.

Does it support production agent workflows?

Yes. It is positioned for multi-step agents that plan, use tools, evaluate results, and refine substantial coding or operational work.

What are the context and output limits?

The API supports up to 1M context tokens and up to 384K output tokens.

How do I control thinking?

Thinking can be disabled. When enabled, set effort to low, high, or max to match the task.