Language model workflows
Claude Haiku 5.5: API Pricing, Migration, and When to Use It
Haiku 5.5 is a released Anthropic model, but it is not currently in the SJolt catalog. Here is how to evaluate its pricing and request changes without confusing provider access with SJolt support.
Claude Haiku 5.5 launched on October 7, 2026. Anthropic positions it for frequent, narrowly scoped work such as extraction, classification, and routing. The direct Claude API identifier is claude-haiku-5-5. The shorter search phrase “Haiku 5.5” refers to the same model.
As checked on October 10, SJolt does not list Haiku 5.5. Use Anthropic’s documented access paths to evaluate that model. SJolt does list Claude Sonnet 5.5 and Opus 5.5, which are separate models with their own contracts and prices.
Start with the model and the access path
The model reference lists a 1M-token context window, up to 128K output tokens, and text-and-image input with text output. These are provider limits. They do not mean the model generates image files or that a particular account has enough throughput for a large batch.
Choose one representative job before connecting a production queue. For a creative application, a suitable first job is extracting the subject, framing, required text, and missing facts from a customer brief. Keep generation of the finished image as a later step so you can inspect the interpretation first.
Budget around the 100K prompt boundary
The following direct Claude Platform rates were checked on October 10, 2026. They are USD per million tokens and are not SJolt prices. Prompt length selects the pricing tier; do not budget a long request using the headline starting rate.
| Prompt length | Input / million tokens | Output / million tokens |
|---|---|---|
| Up to 100,000 tokens | $0.10 | $0.50 |
| Over 100,000 tokens | $0.50 | $2.50 |
For illustration, 20,000 uncached input tokens and 2,000 billable output tokens cost $0.003 at the short-prompt rates: $0.002 for input plus $0.001 for output. This arithmetic excludes tools, caching, discounts, retries, and other charges. It is a budget example, not a measured workload.
Save token usage per completed job and include unsuccessful attempts in your totals. Dividing total spend by accepted results is more informative than comparing the price of a single request. Also track the time a person spends correcting extracted facts; a cheap classification that causes the wrong generation can dominate the later image budget.
Audit old requests before changing the model name
Anthropic’s migration guide calls out changes beyond the model ID. Replace manual thinking budgets with adaptive thinking; omit temperature, top_p, and top_k; and remove a final assistant prefill. Select response blocks by type instead of assuming the first block is answer text.
{
"model": "claude-haiku-5-5",
"max_tokens": 4096,
"thinking": {
"type": "adaptive"
},
"output_config": {
"effort": "medium"
},
"messages": [
{
"role": "user",
"content": "Extract subject, framing, required_text, and missing_facts from this brief: A square studio image of our blue mug. Capacity is unspecified. Return JSON only."
}
]
}This body illustrates the direct Claude Messages API contract; it is not an SJolt Haiku endpoint. A prompt asking for JSON does not enforce a schema. Parse the returned text and validate keys and values before acting on it.
Recount tokens with the new model: Anthropic estimates roughly 30% more tokens for the same text than Haiku 4.5, with variation by content. Review truncation and refusal handling as well. Keep signed thinking blocks unchanged when continuing tool conversations and consult the guide before replaying histories across accounts.
Use a routing test with an explicit uncertainty path
Build a small set of briefs that includes complete requests, contradictory instructions, missing references, and unsupported output requests. Write the expected route before evaluating responses. A model should be allowed to return “needs clarification”; forcing every case into a generation route hides uncertainty.
| Brief condition | Expected application behavior |
|---|---|
| Clear new image request | Extract a complete image brief for review. |
| Edit request without a source image | Ask for the missing reference before generation. |
| Product capacity not provided | Preserve it as unknown; do not invent a label. |
| Conflicting square and widescreen instructions | Return the conflict for resolution. |
Treat these as application acceptance rules, not claims about observed model performance. Measure route accuracy separately from schema validity. A perfectly valid JSON object can still select the wrong workflow. Retain the input, answer, expected label, and review note for every failure so that a prompt change can be checked against the same cases.
Haiku or Sonnet depends on the work you keep
Anthropic describes Haiku as a fit for bounded tasks and larger models as better suited to more complex coding work. Use that positioning to form a hypothesis, then test your actual briefs. A difficult multi-document planning task may deserve a stronger model even if a simple field extraction does not.
For an SJolt-only workflow today, evaluate a listed language model for the briefing step. Send the approved prompt separately to Nano Banana 2.1 for a raster image. The language response and the asynchronous media task have different lifecycles; store both records and do not interpret a useful brief as a finished image.
No model requests or independent speed tests were run for this guide. The decision criteria above are a proposed evaluation method; the availability and migration details come from the linked primary documentation.
Sources & further reading
Take the next idea into production.
Explore the models, test a workflow in the playground, and use the same request in your application.