Skip to content
← Blog
7 min read

Claude Opus 5.5: what changed and how to use it on Kumo

Anthropic shipped Claude Opus 5.5 — Fable 5.1's level for $4 and $20 per million tokens. Kumo pricing, the benchmarks, and two checks that prove the original model is answering.

anthropicclaudemodel verificationopus 5.5pricing

On September 22, 2026 Anthropic released Claude Opus 5.5 — a model that holds Claude Fable 5.1's level on most real work while costing 2.5 times less than it and 20% less than the previous Opus 5. It is already available on Kumo: the same API, one balance topped up by card or crypto, and no weekly quotas. And it is the original Opus 5.5, not a cheaper model wearing its name — below are two checks you can run yourself to confirm it.

Context

1M tokens

up to 128K tokens in one reply

Anthropic price

$4 / $20

input / output per 1M tokens

Kumo price

$2.2 / $11

45% off the vendor's list

Terminal-Bench 4.0

66.4%

Opus 5 scores 52.3%

What Claude Opus 5.5 is

Claude Opus 5.5 is the first model of the Claude 5.5 family and, in Anthropic's own words, the strongest-performing model it has ever put through its automated behavioral audit. In Anthropic’s API the identifier is claude-opus-5-5 — not a convenience alias but a pinned snapshot, so a request under that name tomorrow returns the same model it returns today. It is positioned for long agentic sessions: reading a repository, running the code — tests — fix loop for hours, driving a computer, and working through long analytical tasks.

The context window is 1M tokens and a single reply can carry up to 128K. The reliable knowledge cutoff is June 2026. Adaptive thinking is always on and the default effort level is medium: the model decides how long to think and does not spend a reasoning budget on a trivial request. Anthropic commits to keeping it available until at least September 22, 2027.

The practical meaning of this release is the ratio between price and level. Until September 22 the choice between “smarter” and “cheaper” cost a multiple on the invoice: Fable 5.1 asks $10 per million input tokens and $50 per million output tokens. Opus 5.5 covers most of the same class of work for $4 and $20 — and, by Anthropic's measurements, generates output 30% faster while doing it.

What changed since Opus 5

Eight benchmarks from Anthropic’s announcement, ordered by Opus 5.5’s own score. GDPval-AA sits below the plot: it is an Elo rating, not a percentage, and it does not belong on the same scale.
BenchmarkClaude Opus 5.5Claude Opus 5
Terminal-Bench 4.066.4%52.3%
FrontierCode v1.154.4%48.0%
CursorBench 4.057.8%46.6%
OSWorld 2.0 (partial credit)81.8%74.0%
GDPval-AA v2.11846 Elo1708 Elo
Anthropic's own figures from the announcement of September 22, 2026. The methodology and run conditions are on the announcement page.

The jump is largest exactly where the model works on its own: Terminal-Bench and CursorBench do not measure a single answer, they measure carrying a task to the end in a terminal and in an editor. Fourteen points and eleven points are the difference between “the agent finished it” and “the agent got stuck on step three”.

Anthropic calls out two more things. The first is spend: on typical workloads a run costs 40% less than on Opus 5, and not only because of the rate — the model also spends fewer tokens on the same work. The second is how it talks: less jargon, fewer preambles, the important part first. On diff reviews and written reports that shows up more clearly than the percentage gains do.

What exactly got cheaper

The vendor comparing itself with its own previous model. On Kumo, Opus 5.5 sits a further 45% below the figures listed here.

What Claude Opus 5.5 costs

What is billedAnthropicKumo
Input, uncached$4$2.2
Output$20$11
Cache reads$0.2$0.11
Cache writes$5$2.75
USD per 1M tokens, excluding the Batch API and special modes. Kumo's catalogue discount is 40% off the vendor's list, and on Opus 5.5 it goes deeper: 45%. The figures in force are always in the price list on the models page.
The hollow dot is the same model at the Kumo rate. Further left and higher means more work per dollar.

Opus 5.5 is already available on Kumo

The model is wired into the shared catalog and runs off the same key and the same balance as the rest: no separate subscription, no waiting for a quota to refill, no second invoice. The balance is topped up by Russian card, through SBP or in crypto, and spent on any model in the catalog.

On Kumo the model is named anthropic/claude-opus-5-5 — that is the name in the model list. Both protocols work. OpenAI-compatible clients point at https://api.kumorouter.com/v1; clients speaking the Anthropic Messages protocol point at the bare origin https://api.kumorouter.com, because they append /v1/messages themselves. The key is accepted both as Authorization: Bearer and as x-api-key, so a genuine Claude client connects with no shim in between.

One honest limit: every key carries its own spending caps and alerts, and we do not log request bodies — only billing metadata: the model, the token counts and the time.

export KUMO_API_KEY="kumo_sk_…"

# OpenAI-compatible clients
export OPENAI_BASE_URL="https://api.kumorouter.com/v1"
export OPENAI_API_KEY="$KUMO_API_KEY"

# clients speaking the Anthropic Messages protocol
export ANTHROPIC_BASE_URL="https://api.kumorouter.com"
export ANTHROPIC_AUTH_TOKEN="$KUMO_API_KEY"

The model is the original — and you can prove it

The main fear when buying access through an intermediary is that a cheaper model answers under an expensive model's name. The fear is reasonable: you cannot see the difference by eye, and the gap on the invoice between Opus 5.5 and a simpler model can be several-fold.

So we do not ask you to take our word for it. On Kumo the model is never swapped: ask for anthropic/claude-opus-5-5 and anthropic/claude-opus-5-5 answers, and every response names who produced it. Below are two checks you run yourself, with your own key. They complement each other: the first catches a careless substitution, the second catches one that a correct string in the response cannot hide.

An honest caveat up front: neither check is a cryptographic proof, and we are not pretending otherwise. Together they rule out the cheapest ways to substitute a model — and a fake that costs more than the real thing makes no economic sense.

Check 1. The catalog and the model echo in the response

# the model must appear in the catalog under its real identifier
curl -s https://api.kumorouter.com/v1/models \
  -H "Authorization: Bearer $KUMO_API_KEY" | grep -o '"anthropic/claude-opus-5-5"'

# the model field in the response must repeat exactly the name you sent
curl -s https://api.kumorouter.com/v1/chat/completions \
  -H "Authorization: Bearer $KUMO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"anthropic/claude-opus-5-5","messages":[{"role":"user","content":"ping"}],"max_tokens":16}' \
  | python3 -c "import sys,json; d=json.load(sys.stdin); print(d['model'], d['usage'])"

A minute of work, and it rules out the most common case: a catalog where a pretty name is an alias pointing somewhere else. Watch two places — the identifier in the model list and the model field in the response body. Both must read exactly anthropic/claude-opus-5-5.

What this check does not prove: a tidy intermediary can return the right string while still substituting the model. That is why it is the first check rather than the only one.

Check 2. The 1M window and a needle in the haystack

python3 - <<'PY'
import os, httpx

needle = " The pass phrase of the day is kumo-opus-55-million. "
filler = "This is filler text whose only job is to fill the context window. " * 20000
question = " Question: state the pass phrase of the day on one line, with no explanation."
prompt = filler + needle + filler + question

r = httpx.post(
    "https://api.kumorouter.com/v1/chat/completions",
    headers={"Authorization": "Bearer " + os.environ["KUMO_API_KEY"]},
    json={
        "model": "anthropic/claude-opus-5-5",
        "messages": [{"role": "user", "content": prompt}],
        "max_tokens": 64,
    },
    timeout=900,
)
body = r.json()
print(body["usage"]["prompt_tokens"], "input tokens")
print(body["choices"][0]["message"]["content"])
PY

A cheap model cannot pass this one physically. A 200K window will not accept a request of several hundred thousand tokens — an over-context error comes back instead, and the substitution exposes itself. A model with a real million-token window accepts the request and states the phrase hidden exactly in the middle of it.

Watch both numbers: the answer and prompt_tokens. The input counter also shows that the request arrived whole rather than being quietly truncated halfway through.

The check is not free: the request runs to roughly half a million input tokens, which at Kumo's rate is about a dollar. That is cheaper than one day of work on a model you do not trust.

What each check proves, and what it does not

CheckWhat it provesWhat it does not prove
Catalog and model echoThat the route is declared under the real nameNothing about who actually answered
1M window and needleThat the window really is a million tokensDoes not tell Opus 5.5 from Fable 5.1
Request log in the consoleThat you were billed for that exact modelDoes not verify the content of the answer
The checks are arranged so that each one covers the blind spot of the one before it.

How to connect Opus 5.5 in five minutes

  1. Sign up in the console and top up the balance — by card, through SBP or in crypto.
  2. Create a key and give it a spending cap right away: one key per project saves you the argument about where the money went.
  3. Point your client at the gateway — https://api.kumorouter.com/v1 for OpenAI-compatible tools, the bare origin for clients on the Anthropic protocol.
  4. Set the model to anthropic/claude-opus-5-5 and send the first request.
  5. Run checks 1 and 2 from this article before you start spending in earnest.
curl -s https://api.kumorouter.com/v1/chat/completions \
  -H "Authorization: Bearer $KUMO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anthropic/claude-opus-5-5",
    "messages": [
      {"role": "user", "content": "Find every place in this repository where an environment variable is read and put them in a table."}
    ],
    "max_tokens": 2048
  }'

Frequently asked questions

How does Opus 5.5 differ from Fable 5.1, and which should I pick?
Fable 5.1 remains the top model for the heaviest reasoning and the longest agentic sessions. Opus 5.5 covers most of the same work for $4 and $20 against $10 and $50, and answers faster. The sensible order is to start on Opus 5.5 and move up to Fable 5.1 where your own evals show it falling short.
Is this really the same Opus 5.5 as Anthropic's own?
Yes: the route leads to the original model, not to a stand-in. You can confirm it two ways from the section above — the catalog and the identifier echo, and the million-token window with a needle in it. You run both yourself, with your own key, and the console's request log shows what each request was billed as.
Why is it cheaper on Kumo than at the vendor?
The catalogue discount is 40% off the vendor's list on ordinary input and output tokens, and on Opus 5.5 at launch it goes deeper: 45%, which is $2.2 and $11 instead of $4 and $20. A package ladder applies on top for larger volumes. Compare the same model version and the same kind of token: uncached input, cache reads and cache writes are billed at different rates.
Do I need a subscription, and are there weekly quotas?
No. Billing is pay-as-you-go from one balance: top it up, work, top it up again. There are no weekly or daily subscription quotas here, and the spending limits are the ones you set yourself on each key.
When do Sonnet 5.5 and Haiku 5.5 arrive?
Anthropic promises them in the coming weeks with similar gains. We connect new vendor models as they ship — watch the models page and the blog.

Sources

  • Anthropic's announcement: Introducing Claude Opus 5.5 — benchmarks, pricing and safety data.
  • Vendor documentation: Claude models overview — identifiers, context window, knowledge cutoffs and retirement dates.
  • Anthropic pricing: Pricing — cache rates, the Batch API and per-cloud prices.
  • Kumo pricing: the models page — the rates in force for every model in the catalog.

Start building on Kumo today

One base URL, one balance, every model — at an effective rate you can see before you spend a token