How to Use the Claude Haiku 5.5 API ?

Claude Haiku 5.5 API guide: first call with claude-haiku-5-5 in curl, Python and TypeScript, plus effort, thinking, caching, batch and refusals.

INEZA Felin-Michel

INEZA Felin-Michel

8 October 2026

How to Use the Claude Haiku 5.5 API ?

Apidog for Enterprise

On-Premises Deploy

SSO & RBAC

SOC 2 Compliant

Explore Apidog Enterprise

To use the Claude Haiku 5.5 API, send a POST request to https://api.anthropic.com/v1/messages with "model": "claude-haiku-5-5", your key in the x-api-key header, and anthropic-version: 2023-06-01. It costs $0.10/$0.50 per million input/output tokens for prompts up to 100K tokens ($0.50/$2.50 above that), reads up to 1M tokens of context, writes up to 128K, and defaults to medium effort with adaptive thinking on.

Anthropic released Haiku 5.5 on October 7, 2026, and it’s the first Haiku with effort levels (what is Claude Haiku 5.5 covers specs and positioning). This guide covers a first call in curl, Python and TypeScript, then effort, thinking, caching, batch, refusals and agent toolsets. You can save and assert every request below in Apidog.

button

Claude Haiku 5.5 API at a glance

Parameter Haiku 5.5 behavior
Model ID claude-haiku-5-5 (Bedrock: anthropic.claude-haiku-5-5); no separate alias
Price per MTok, prompts up to 100K tokens $0.10 input, $0.50 output, $0.01 cache reads
Price per MTok, prompts over 100K tokens $0.50 input, $2.50 output, $0.05 cache reads
Context / max output 1M / 128K; 300K on Batch with the output-300k-2026-03-24 beta header
output_config.effort low, medium (default), high, xhigh, max
thinking adaptive by default; disabled only at high effort or below
thinking.display Empty thinking field by default; summarized returns readable text
temperature, top_p, top_k Non-default values return 400
Assistant prefill Returns 400, even with thinking off
Minimum cacheable prompt 512 tokens (4,096 on Haiku 4.5)

Sources: the Haiku 5.5 model page and the Claude API pricing docs.

Your first Claude Haiku 5.5 API call

Create a key in the Claude Console (the Anthropic API key guide walks through it) and export it as ANTHROPIC_API_KEY. Never paste the key into code. Then send this:

curl https://api.anthropic.com/v1/messages \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "claude-haiku-5-5",
    "max_tokens": 4096,
    "output_config": {"effort": "medium"},
    "thinking": {"type": "adaptive", "display": "summarized"},
    "messages": [{"role": "user", "content": "Classify this ticket as billing, bug, or feature request: The export button times out on large projects."}]
  }'

The Python SDK picks up ANTHROPIC_API_KEY from the environment:

import anthropic

client = anthropic.Anthropic()
response = client.messages.create(
    model="claude-haiku-5-5",
    max_tokens=4096,
    output_config={"effort": "medium"},
    thinking={"type": "adaptive", "display": "summarized"},
    messages=[{"role": "user", "content": "Classify this ticket as billing, bug, or feature request: The export button times out on large projects."}],
)

for block in response.content:
    if block.type == "thinking":
        print("[thinking]", block.thinking)
    elif block.type == "text":
        print(block.text)
print(response.stop_reason, response.usage)

TypeScript follows the same shape:

import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic();
const response = await client.messages.create({
  model: "claude-haiku-5-5",
  max_tokens: 4096,
  output_config: { effort: "medium" },
  thinking: { type: "adaptive", display: "summarized" },
  messages: [
    { role: "user", content: "Classify this ticket as billing, bug, or feature request: The export button times out on large projects." },
  ],
});

for (const block of response.content) {
  if (block.type === "text") console.log(block.text);
}
console.log(response.stop_reason, response.usage);

Three habits keep this code working. Select content blocks by type, because a response can open with a thinking block and content[0].text breaks. Leave headroom in max_tokens, because thinking tokens count toward it. And keep the request body clean: no temperature, top_p, top_k, budget_tokens or assistant prefill. Each one is a 400 on this model. If you’re moving older code, the Haiku 5.5 vs Haiku 4.5 guide lists every breaking change with before/after JSON.

Pick an effort level

Effort, set in output_config.effort, is the main dial for quality, latency and cost. The prompting guide gives these starting points:

The cost curve is steep. Here are Anthropic’s own OSWorld 2.1 (offline subset) runs from the launch charts, with partial-credit score and cost per attempt:

Effort Score Cost per attempt
low 42.0% $0.0695
medium 53.3% $0.1257
high 61.3% $0.1827
xhigh 67.6% $0.2792
max 72.4% $0.6111

Going from xhigh to max more than doubles the cost for under five points. The Haiku 5.5 benchmarks breakdown has the other per-effort charts.

One quirk: at xhigh in multi-turn chats, the model sometimes writes its whole answer in its thinking and ends the turn with no visible text. Check for an empty reply before showing it to a user.

Control thinking

Adaptive thinking is on by default, and two things changed from Haiku 4.5. First, the default display hides the text. Each thinking block comes back with an empty thinking field and only a signature. Set "display": "summarized" (as in the first call) when you want readable summaries in logs or a UI. To get less thinking, lower the effort; prompting the model to answer directly didn’t stop it in Anthropic’s testing.

Second, you can turn thinking off, but only at high effort or below:

{
  "model": "claude-haiku-5-5",
  "max_tokens": 1024,
  "thinking": {"type": "disabled"},
  "output_config": {"effort": "low"},
  "messages": [{"role": "user", "content": "Extract the invoice number from: INV-2291, due Nov 3."}]
}

The same body at xhigh or max returns a 400. A forced tool_choice (any or a named tool) is accepted, but the response starts with the tool call and carries no thinking block.

For multi-turn and agent loops, pass every thinking block back unchanged and keep history append-only. Changing system, tools or earlier messages before a returned thinking block can return a 400, and thinking blocks work only in the account that produced them (or one linked to it).

Cache prompts and batch jobs

Caching is where Haiku 5.5 gets cheap. For prompts up to 100K tokens, a cache read costs $0.01 per million tokens against $0.10 for fresh input, a 5-minute cache write costs $0.125 and a 1-hour write $0.20. The minimum cacheable prompt is 512 tokens, down from 4,096 on Haiku 4.5, so short system prompts and tool lists now qualify. Mark the stable prefix with cache_control:

{
  "model": "claude-haiku-5-5",
  "max_tokens": 1024,
  "system": [{
    "type": "text",
    "text": "You are a support triage assistant. <long, stable policy text here>",
    "cache_control": {"type": "ephemeral"}
  }],
  "messages": [{"role": "user", "content": "Ticket: refund not received after 10 days."}]
}

Changing top-level effort between requests invalidates the cache; per-message effort (beta header mid-conversation-output-config-2026-07-01, Claude API and Google Cloud) keeps it. The prompt caching docs cover TTLs, and our prompt caching explainer covers the concept.

For work that can wait, the Message Batches API cuts input and output by 50%: $0.05/$0.25 for prompts up to 100K tokens and $0.25/$1.25 above. Batch is also the only route to 300K output tokens, with the output-300k-2026-03-24 beta header.

Watch the 100K line: “a prompt of over 100,000 tokens pays higher prices,” in Anthropic’s words. The Haiku 5.5 pricing guide works through examples on both sides.

Handle stop_reason “refusal”

Haiku 5.5 runs safety classifiers that can decline a request, and it has no server-side fallback. A declined request comes back with stop_reason: "refusal", and the categories are cyber, frontier_llm, bio and general_harms. If you’re moving from Haiku 4.5, these refusals are new. Sending the same request again usually returns another refusal, so don’t retry blindly:

def run(client, messages):
    response = client.messages.create(
        model="claude-haiku-5-5",
        max_tokens=4096,
        messages=messages,
    )
    if response.stop_reason == "refusal":
        details = getattr(response, "stop_details", None)
        category = getattr(details, "category", "unknown")
        log_refusal(category, messages)  # your logging
        return {"status": "refused", "category": category}
    text = "".join(b.text for b in response.content if b.type == "text")
    return {"status": "ok", "text": text}

Branch on stop_reason before you read content, and route refusals to a person or another model in your own code. Teams doing legitimate security or life sciences work blocked by the cyber or bio classifiers can apply to Anthropic’s Cyber Verification Program or Life Sciences Verification Program.

Computer use and browser use

On the Claude API and Google Cloud, Haiku 5.5 supports computer use only through the computer_toolset_20260801 toolset, which needs no beta header; declaring computer_20250124 returns a 400. Browser use goes through browser_toolset_20260801, which Haiku 4.5 doesn’t support. The Python and TypeScript SDKs added beta classes for both on launch day. See the computer use tool docs for the member tools.

Rate limits

Haiku 5.5 has the same rate limits as Haiku 4.5: 1,000 requests, 2M input tokens and 400K output tokens per minute on the Start tier, up to 10,000 requests, 10M input and 2M output on Scale. Priority Tier isn’t supported. For 429 handling, see the rate limit exceeded guide.

Test the Claude Haiku 5.5 API in Apidog

Saved requests make effort comparisons and refusal debugging repeatable. Here’s the setup in Apidog:

  1. Create an environment and add ANTHROPIC_API_KEY as a secret variable. Reference it as {{ANTHROPIC_API_KEY}} in the x-api-key header, next to anthropic-version: 2023-06-01 and content-type: application/json.
  2. Create a POST request to https://api.anthropic.com/v1/messages, paste the first-call body, and save it.
  3. Add assertions: status is 200, $.stop_reason equals end_turn, $.usage.output_tokens is greater than 0, and $.content[*].type contains text. A refusal or an empty xhigh reply now fails the test instead of slipping through.
  4. Duplicate the request four times with low, high, xhigh and max, and run the folder. You get usage for every effort level on your own prompt.
  5. Add the cached-system-prompt variant and assert that $.usage.cache_read_input_tokens is greater than 0 on the second run.

For broader patterns, see testing LLM applications.

FAQ

What is the Claude Haiku 5.5 model ID? claude-haiku-5-5, with no date suffix and no separate alias, on the Claude API, Google Cloud, Microsoft Foundry and Claude Platform on AWS. On Amazon Bedrock it’s anthropic.claude-haiku-5-5.

Is there a free Claude Haiku 5.5 API? There’s no ongoing free tier, but new API users get a small amount of free credit to test the API. Free Claude.ai users can select Haiku 5.5 in chat, but that isn’t an API key. Max and Team plans now include monthly API credits. The free access guide covers what counts and what doesn’t.

Why does my Haiku 4.5 request return 400? Check for budget_tokens, a non-default temperature or top_p, any top_k, an assistant prefill, or the old computer_20250124 tool. Those are the usual causes.

Can I use Haiku 5.5 in Claude Code? Yes, from v2.1.293. On the Anthropic API the haiku alias resolves to Haiku 5.5. See Claude Haiku 5.5 in Claude Code.

Should I use Haiku 5.5 or Sonnet 5.5 for agentic coding? Anthropic says Sonnet 5.5 and Opus 5.5 “remain better choices for complex agentic coding tasks.” Use Haiku 5.5 for narrowly scoped work: classification, summarization, compaction, subagents and browser use.

Next step

Send the first-call request at medium, then rerun it at low and high on a prompt from your own workload and compare usage.output_tokens and answer quality. Download Apidog to keep all three runs with assertions, so the next model release is a one-field change.

button

Explore more

Claude Haiku 5.5 vs GPT-6 Luna

Claude Haiku 5.5 vs GPT-6 Luna

Haiku 5.5 vs GPT-6 Luna: same $0.10/$0.50 price under 100K tokens, different long-prompt tiers, Anthropic's benchmarks, and per-effort cost data.

8 October 2026

Claude Haiku 5.5 Benchmarks

Claude Haiku 5.5 Benchmarks

Claude Haiku 5.5 benchmarks: 72.4% OSWorld, 1620 GDPval-AA, 39.2% Terminal-Bench at max effort. Who ran each test, per-effort costs, and your own eval.

8 October 2026

Claude Haiku 5.5 vs Haiku 4.5: What Changed and the Breaking Changes to Fix First

Claude Haiku 5.5 vs Haiku 4.5: What Changed and the Breaking Changes to Fix First

Haiku 5.5 vs Haiku 4.5: 90% cheaper up to 100K tokens, 1M context, and five breaking changes that return 400s. Before/after JSON fixes inside.

8 October 2026

Practice API Design-first in Apidog

Discover an easier way to build and use APIs

How to Use the Claude Haiku 5.5 API ?