Skip to content
← Blog
8 min read

GPT-6 Sol: what changed and how to use it on Kumo

OpenAI shipped GPT-6 Sol — half the price of GPT-5.6 Sol and close to the best Claude models at a fraction of the cost per task. On Kumo it is $1.1 and $5.5 per million tokens. The benchmarks, and two checks that prove the original model is answering.

gpt-6gpt-6 solmodel verificationopenaipricing

On September 22, 2026 OpenAI released GPT-6 Sol — the GPT-6 generation's working model for coding and agents, trained with the same methods as the flagship GPT-6 Astra. On OpenAI's published tests Sol keeps pace with the best Claude models at a fraction of the cost per task, and it costs half of what GPT-5.6 Sol did. It is already available on Kumo at 45% below OpenAI's list: the same API, one balance topped up by card or crypto, and no weekly quotas. And it is the original GPT-6 Sol, not a cheaper model wearing its name — below are two checks you can run yourself to confirm it.

Context

1.05M tokens

up to 128K tokens in one reply

OpenAI price

$2 / $10

input / output per 1M tokens

Kumo price

$1.1 / $5.5

45% off the vendor's list

AutomationBench

33.2%

at 9% of Claude Opus 5's cost per task

What GPT-6 Sol is

Earlier in September OpenAI introduced GPT-6 Astra, its most capable model. Sol is the next step of the same generation: OpenAI trained it with the methods behind Astra and carries its gains in professional work, factuality, coding and computer use into a faster, cheaper model. Astra remains the best model of the line, and OpenAI says plainly to pick it where you want the best result with no compromise.

GPT-6 Sol is the workhorse for complex coding and agentic workflows: reading a repository, the code — tests — fix loop, automating processes across dozens of tools. The OpenAI API identifier is gpt-6-sol. OpenAI shipped GPT-6 Luna alongside it — the generation's most efficient model for high-volume work; it is not available on Kumo yet.

The context window is 1,050,000 tokens with up to 128K tokens in a single reply; it takes text and images in and returns text. The knowledge cutoff is April 20, 2026. Reasoning depth is set with reasoning.effort, from none to max, with medium as the default.

The practical point of the release is the price. OpenAI cut Sol's rates by 50% against GPT-5.6 Sol's promotional pricing and attributes it to caching and inference improvements rather than a promotion: it now costs $2 per million input tokens and $10 per million output tokens.

What Sol does better

Three tests from OpenAI's announcement. OpenAI chose the opponent and the effort level on each test, so every bar names exactly what it is compared with.

The announcement's main point is not a record but the cost of a result. On AutomationBench, where an agent runs business workflows across 47 tools from sales to finance, Sol at xhigh effort scores 33.2% and beats Claude Opus 5 at max while spending 9% of its cost per task. On DeepSWE v1.1 — long tasks in real codebases — Sol at max scores 68.8%, 1.1 points short of Claude Fable 5's best, at roughly 80% lower cost per task. On OSWorld 2.0 computer use Sol and Opus 5 are practically level, 60.5% against 60.3%, again with a cost gap of about 80%. On Agents’ Last Exam, where agents work long professional tasks across 55 sub-industries, Sol at max scores 56.4% — above Opus 5's best at 60% lower cost per task.

OpenAI singles out two more things. Factuality: on an internal set of conversations where users flagged mistakes, Sol makes about half as many errors as GPT-5.6 Sol. Style: the model inherits Astra's clearer way of talking — less jargon, fewer low-value details, slightly shorter answers with the substance intact; it shows most in conversations about code.

And an honest caveat: every comparison is against Claude Opus 5 and Fable 5.1, not Claude Opus 5.5, which shipped the same day. OpenAI published no head-to-head with it, so choose between them on your own tasks.

OpenAI prints a dollar cost per task only for Sol; the other points are derived from its own multiples (3.9×, 8.9×, 11.1×). Fable 5.1's cost is understated: it excludes the Opus 5 fallbacks needed on roughly 40% of tasks.

What exactly got cheaper

OpenAI's list and Kumo's rates before and after, percentages computed from the prices themselves. At OpenAI, GPT-5.6 Sol is compared at its promotional price.

The other half of the saving is the cache. For GPT-6 OpenAI improved prompt caching: cached input tokens cost 10% of the normal rate, default cache hit rates went up, and changing the reasoning effort or the tool set mid-conversation no longer breaks the cache. For agents that resend the same long context over and over, that matters more than the cut in the base rate.

What GPT-6 Sol costs

What is billedOpenAIKumo
Input, uncached$2$1.1
Cache reads$0.2$0.1 ($0.11)
Cache writes$2.5$1.4 ($1.375)
Output$10$5.5
Requests over 272K tokens×2 on input and cache, ×1.5 on outputthe same rates
USD per 1M tokens, standard processing, excluding Batch, Flex and fast mode. The Kumo rate is 45% below the vendor's list; the exact rate is in brackets where the showcase rounds it. The figures in force are always in the price list on the models page.

GPT-6 Sol is already available on Kumo

The model is wired into the shared catalog and runs off the same key and the same balance as the rest: no separate subscription, no waiting for a quota to refill, no second invoice. The balance is topped up by Russian card, through SBP or in crypto, and spent on any model in the catalog, GPT-6 Astra and the Claude models included.

On Kumo the model is named openai/gpt-6-sol. Both OpenAI protocols work at https://api.kumorouter.com/v1: the Responses API (/v1/responses) and Chat Completions (/v1/chat/completions). If your agent needs tools, use the Responses API: this is OpenAI's own restriction — in Chat Completions, GPT-6 function calling works only with reasoning_effort set to none.

Control stays with you: every key carries its own spending caps and alerts, and we do not log request bodies — only billing metadata: the model, the token counts and the time.

export KUMO_API_KEY="kumo_sk_…"

# OpenAI-compatible clients and SDKs
export OPENAI_BASE_URL="https://api.kumorouter.com/v1"
export OPENAI_API_KEY="$KUMO_API_KEY"

The model is the original — and you can prove it

The main fear when buying access through an intermediary is that a cheaper model answers under an expensive model's name. You cannot see it by eye, and the temptation is real: OpenAI itself sells models several times cheaper than Sol, and quietly serving one of them is the most profitable swap there is.

So we do not ask you to take our word for it. On Kumo the model is never swapped: ask for openai/gpt-6-sol and GPT-6 Sol answers, and every response names who produced it. Below are two checks you run yourself, with your own key, in a few minutes. The first catches a careless substitution; the second catches one that no correct string in the response can hide.

An honest caveat up front: neither check is a cryptographic proof, and neither reliably tells apart a model of the same OpenAI generation. Together they close the cheap ways to substitute a model, and at the end we say what to do about the one they do not catch.

Check 1. The catalog and the model echo in the response

# the model must appear in the catalog under its real name
curl -s https://api.kumorouter.com/v1/models \
  -H "Authorization: Bearer $KUMO_API_KEY" | grep -o '"openai/gpt-6-[a-z]*"'

# the model field in the response must name the model you asked for
curl -s https://api.kumorouter.com/v1/responses \
  -H "Authorization: Bearer $KUMO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"openai/gpt-6-sol","input":"ping","max_output_tokens":512}' \
  | python3 -c "import sys,json; d=json.load(sys.stdin); print(d['model'], d['usage'])"

A minute of work, and it rules out the most common case: a catalog where a pretty name is an alias pointing somewhere else. Watch two places — the name in the model list and the model field in the response body. In the list the model is openai/gpt-6-sol, and the model field repeats exactly the name you sent — no suffix, no other name.

What this check does not prove: a tidy intermediary can return the right string while still substituting the model. That is why it is the first check rather than the only one.

Check 2. The 1.05M window and a needle in the haystack

python3 - <<'PY'
import os, httpx

needle = " The pass phrase of the day is kumo-gpt6-million. "
filler = "This is filler text whose only job is to fill the context window. " * 20000
question = " Question: state the pass phrase of the day on one line, with no explanation."

r = httpx.post(
    "https://api.kumorouter.com/v1/responses",
    headers={"Authorization": "Bearer " + os.environ["KUMO_API_KEY"]},
    json={
        "model": "openai/gpt-6-sol",
        "input": filler + needle + filler + question,
        "reasoning": {"effort": "low"},
        "max_output_tokens": 1024,
    },
    timeout=900,
)
body = r.json()
print(body["usage"]["input_tokens"], "input tokens")
print(body["output"][-1]["content"][0]["text"])
PY

A model with a short window cannot pass this one physically. A 128K or 400K window will not accept a request of over half a million tokens — an over-context error comes back instead, and the substitution exposes itself. A model with a real 1.05M window accepts the request and states the phrase hidden exactly in the middle of it.

Watch both numbers: the answer and input_tokens. The input counter also shows that the request arrived whole rather than being quietly truncated halfway through.

The check is not free, but it is cheap: just over half a million input tokens at Kumo's rate cost about 60 cents, because Kumo adds no long-context surcharge. That is cheaper than one day of work on a model you do not trust.

What each check proves, and what it does not

CheckWhat it provesWhat it does not prove
Catalog and model echoThat the route is declared under the real nameNothing about who actually answered
1.05M window and needleThat the window really exceeds a million tokensDoes not tell Sol from other GPT-6 models
Request log in the consoleThat you were billed for that exact modelDoes not verify the content of the answer
The checks are arranged so that each one covers the blind spot of the one before it. None of them reliably tells a model of the same OpenAI generation apart — your own tasks and the request log do that.

How to connect GPT-6 Sol in five minutes

  1. Sign up in the console and top up the balance — by card, through SBP or in crypto.
  2. Create a key and give it a spending cap right away: one key per project saves you the argument about where the money went.
  3. Point your client at the gateway https://api.kumorouter.com/v1 and use the Kumo key in place of an OpenAI one.
  4. Set the model to openai/gpt-6-sol and send the first request; use the Responses API for tools.
  5. Run both checks from this article before you start spending in earnest.
curl -s https://api.kumorouter.com/v1/responses \
  -H "Authorization: Bearer $KUMO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-6-sol",
    "reasoning": {"effort": "high"},
    "input": "Write a Go function that loads configuration from environment variables with checks for required fields, plus tests for it.",
    "max_output_tokens": 4096
  }'

Frequently asked questions

GPT-6 Sol or GPT-6 Astra?
Astra is for the hardest tasks, where the best result matters more than the price. Sol is the main model for coding and agents: on OpenAI's tests it beats Astra at low effort and stays close to the best Claude models, and on Kumo it costs $1.1 and $5.5 against Astra's $6 and $30. The sensible order is to start on Sol and move up to Astra where your own evals show it falling short.
GPT-6 Sol or Claude Opus 5.5?
On the vendors' list prices Sol costs half: $2 and $10 against $4 and $20 per million tokens. On quality there is no head-to-head — OpenAI compared Sol with Claude Opus 5 and Fable 5.1, and Opus 5.5 shipped the same day. Both models run on Kumo off one key, so the honest move is to run the same tasks on both and compare the results and the charges in the request log.
Is this really the same GPT-6 Sol as OpenAI's own?
Yes: the route leads to the original model, not to a stand-in. You can confirm it two ways from the section above — the catalog and the name echo, and the 1.05M window — and the request log in the console shows which model you were billed for. You run both checks yourself, with your own key.
Why is it cheaper on Kumo than at the vendor?
The routing discount on GPT-6 Sol is 45% off OpenAI's list across every kind of token as of publication, and a package ladder applies on larger volumes on top of that. Compare the same kind of token: uncached input, cache reads and cache writes are billed at different rates.
Do I need a subscription, and are there weekly quotas?
No. Billing is pay-as-you-go from one balance: top it up, work, top it up again. There are no weekly or daily subscription quotas here, and the spending limits are the ones you set yourself on each key.
When does GPT-6 Luna arrive?
We open a model to requests only after its route is verified, so Luna is not available on Kumo yet. When it arrives, it will show on the models page and in the blog.
What happens to GPT-5.6 Sol?
OpenAI has not announced its retirement, and it stays in the Kumo catalog. But GPT-6 Sol costs half as much and is stronger on the vendor's tests, so switching is worth it — changing the model name in the request is all it takes.

Sources

Start building on Kumo today

One base URL, one balance, every model — at an effective rate you can see before you spend a token