On September 22, 2026 OpenAI released GPT-6 Sol — the GPT-6 generation's working model for coding and agents, trained with the same methods as the flagship GPT-6 Astra. On OpenAI's published tests Sol keeps pace with the best Claude models at a fraction of the cost per task, and it costs half of what GPT-5.6 Sol did. It is already available on Kumo at 45% below OpenAI's list: the same API, one balance topped up by card or crypto, and no weekly quotas. And it is the original GPT-6 Sol, not a cheaper model wearing its name — below are two checks you can run yourself to confirm it.
Context
1.05M tokens
up to 128K tokens in one reply
OpenAI price
$2 / $10
input / output per 1M tokens
Kumo price
$1.1 / $5.5
45% off the vendor's list
AutomationBench
33.2%
at 9% of Claude Opus 5's cost per task
What GPT-6 Sol is
Earlier in September OpenAI introduced GPT-6 Astra, its most capable model. Sol is the next step of the same generation: OpenAI trained it with the methods behind Astra and carries its gains in professional work, factuality, coding and computer use into a faster, cheaper model. Astra remains the best model of the line, and OpenAI says plainly to pick it where you want the best result with no compromise.
GPT-6 Sol is the workhorse for complex coding and agentic workflows: reading a repository, the code — tests — fix loop, automating processes across dozens of tools. The OpenAI API identifier is gpt-6-sol. OpenAI shipped GPT-6 Luna alongside it — the generation's most efficient model for high-volume work; it is not available on Kumo yet.
The context window is 1,050,000 tokens with up to 128K tokens in a single reply; it takes text and images in and returns text. The knowledge cutoff is April 20, 2026. Reasoning depth is set with reasoning.effort, from none to max, with medium as the default.
The practical point of the release is the price. OpenAI cut Sol's rates by 50% against GPT-5.6 Sol's promotional pricing and attributes it to caching and inference improvements rather than a promotion: it now costs $2 per million input tokens and $10 per million output tokens.
What Sol does better
The announcement's main point is not a record but the cost of a result. On AutomationBench, where an agent runs business workflows across 47 tools from sales to finance, Sol at xhigh effort scores 33.2% and beats Claude Opus 5 at max while spending 9% of its cost per task. On DeepSWE v1.1 — long tasks in real codebases — Sol at max scores 68.8%, 1.1 points short of Claude Fable 5's best, at roughly 80% lower cost per task. On OSWorld 2.0 computer use Sol and Opus 5 are practically level, 60.5% against 60.3%, again with a cost gap of about 80%. On Agents’ Last Exam, where agents work long professional tasks across 55 sub-industries, Sol at max scores 56.4% — above Opus 5's best at 60% lower cost per task.
OpenAI singles out two more things. Factuality: on an internal set of conversations where users flagged mistakes, Sol makes about half as many errors as GPT-5.6 Sol. Style: the model inherits Astra's clearer way of talking — less jargon, fewer low-value details, slightly shorter answers with the substance intact; it shows most in conversations about code.
And an honest caveat: every comparison is against Claude Opus 5 and Fable 5.1, not Claude Opus 5.5, which shipped the same day. OpenAI published no head-to-head with it, so choose between them on your own tasks.
What exactly got cheaper
The other half of the saving is the cache. For GPT-6 OpenAI improved prompt caching: cached input tokens cost 10% of the normal rate, default cache hit rates went up, and changing the reasoning effort or the tool set mid-conversation no longer breaks the cache. For agents that resend the same long context over and over, that matters more than the cut in the base rate.
What GPT-6 Sol costs
| What is billed | OpenAI | Kumo |
|---|---|---|
| Input, uncached | $2 | $1.1 |
| Cache reads | $0.2 | $0.1 ($0.11) |
| Cache writes | $2.5 | $1.4 ($1.375) |
| Output | $10 | $5.5 |
| Requests over 272K tokens | ×2 on input and cache, ×1.5 on output | the same rates |
GPT-6 Sol is already available on Kumo
The model is wired into the shared catalog and runs off the same key and the same balance as the rest: no separate subscription, no waiting for a quota to refill, no second invoice. The balance is topped up by Russian card, through SBP or in crypto, and spent on any model in the catalog, GPT-6 Astra and the Claude models included.
On Kumo the model is named openai/gpt-6-sol. Both OpenAI protocols work at https://api.kumorouter.com/v1: the Responses API (/v1/responses) and Chat Completions (/v1/chat/completions). If your agent needs tools, use the Responses API: this is OpenAI's own restriction — in Chat Completions, GPT-6 function calling works only with reasoning_effort set to none.
Control stays with you: every key carries its own spending caps and alerts, and we do not log request bodies — only billing metadata: the model, the token counts and the time.
export KUMO_API_KEY="kumo_sk_…"
# OpenAI-compatible clients and SDKs
export OPENAI_BASE_URL="https://api.kumorouter.com/v1"
export OPENAI_API_KEY="$KUMO_API_KEY"
The model is the original — and you can prove it
The main fear when buying access through an intermediary is that a cheaper model answers under an expensive model's name. You cannot see it by eye, and the temptation is real: OpenAI itself sells models several times cheaper than Sol, and quietly serving one of them is the most profitable swap there is.
So we do not ask you to take our word for it. On Kumo the model is never swapped: ask for openai/gpt-6-sol and GPT-6 Sol answers, and every response names who produced it. Below are two checks you run yourself, with your own key, in a few minutes. The first catches a careless substitution; the second catches one that no correct string in the response can hide.
An honest caveat up front: neither check is a cryptographic proof, and neither reliably tells apart a model of the same OpenAI generation. Together they close the cheap ways to substitute a model, and at the end we say what to do about the one they do not catch.
Check 1. The catalog and the model echo in the response
# the model must appear in the catalog under its real name
curl -s https://api.kumorouter.com/v1/models \
-H "Authorization: Bearer $KUMO_API_KEY" | grep -o '"openai/gpt-6-[a-z]*"'
# the model field in the response must name the model you asked for
curl -s https://api.kumorouter.com/v1/responses \
-H "Authorization: Bearer $KUMO_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"openai/gpt-6-sol","input":"ping","max_output_tokens":512}' \
| python3 -c "import sys,json; d=json.load(sys.stdin); print(d['model'], d['usage'])"
A minute of work, and it rules out the most common case: a catalog where a pretty name is an alias pointing somewhere else. Watch two places — the name in the model list and the model field in the response body. In the list the model is openai/gpt-6-sol, and the model field repeats exactly the name you sent — no suffix, no other name.
What this check does not prove: a tidy intermediary can return the right string while still substituting the model. That is why it is the first check rather than the only one.
Check 2. The 1.05M window and a needle in the haystack
python3 - <<'PY'
import os, httpx
needle = " The pass phrase of the day is kumo-gpt6-million. "
filler = "This is filler text whose only job is to fill the context window. " * 20000
question = " Question: state the pass phrase of the day on one line, with no explanation."
r = httpx.post(
"https://api.kumorouter.com/v1/responses",
headers={"Authorization": "Bearer " + os.environ["KUMO_API_KEY"]},
json={
"model": "openai/gpt-6-sol",
"input": filler + needle + filler + question,
"reasoning": {"effort": "low"},
"max_output_tokens": 1024,
},
timeout=900,
)
body = r.json()
print(body["usage"]["input_tokens"], "input tokens")
print(body["output"][-1]["content"][0]["text"])
PY
A model with a short window cannot pass this one physically. A 128K or 400K window will not accept a request of over half a million tokens — an over-context error comes back instead, and the substitution exposes itself. A model with a real 1.05M window accepts the request and states the phrase hidden exactly in the middle of it.
Watch both numbers: the answer and input_tokens. The input counter also shows that the request arrived whole rather than being quietly truncated halfway through.
The check is not free, but it is cheap: just over half a million input tokens at Kumo's rate cost about 60 cents, because Kumo adds no long-context surcharge. That is cheaper than one day of work on a model you do not trust.
What each check proves, and what it does not
| Check | What it proves | What it does not prove |
|---|---|---|
| Catalog and model echo | That the route is declared under the real name | Nothing about who actually answered |
| 1.05M window and needle | That the window really exceeds a million tokens | Does not tell Sol from other GPT-6 models |
| Request log in the console | That you were billed for that exact model | Does not verify the content of the answer |
How to connect GPT-6 Sol in five minutes
- Sign up in the console and top up the balance — by card, through SBP or in crypto.
- Create a key and give it a spending cap right away: one key per project saves you the argument about where the money went.
- Point your client at the gateway https://api.kumorouter.com/v1 and use the Kumo key in place of an OpenAI one.
- Set the model to openai/gpt-6-sol and send the first request; use the Responses API for tools.
- Run both checks from this article before you start spending in earnest.
curl -s https://api.kumorouter.com/v1/responses \
-H "Authorization: Bearer $KUMO_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-6-sol",
"reasoning": {"effort": "high"},
"input": "Write a Go function that loads configuration from environment variables with checks for required fields, plus tests for it.",
"max_output_tokens": 4096
}'
Frequently asked questions
- GPT-6 Sol or GPT-6 Astra?
- Astra is for the hardest tasks, where the best result matters more than the price. Sol is the main model for coding and agents: on OpenAI's tests it beats Astra at low effort and stays close to the best Claude models, and on Kumo it costs $1.1 and $5.5 against Astra's $6 and $30. The sensible order is to start on Sol and move up to Astra where your own evals show it falling short.
- GPT-6 Sol or Claude Opus 5.5?
- On the vendors' list prices Sol costs half: $2 and $10 against $4 and $20 per million tokens. On quality there is no head-to-head — OpenAI compared Sol with Claude Opus 5 and Fable 5.1, and Opus 5.5 shipped the same day. Both models run on Kumo off one key, so the honest move is to run the same tasks on both and compare the results and the charges in the request log.
- Is this really the same GPT-6 Sol as OpenAI's own?
- Yes: the route leads to the original model, not to a stand-in. You can confirm it two ways from the section above — the catalog and the name echo, and the 1.05M window — and the request log in the console shows which model you were billed for. You run both checks yourself, with your own key.
- Why is it cheaper on Kumo than at the vendor?
- The routing discount on GPT-6 Sol is 45% off OpenAI's list across every kind of token as of publication, and a package ladder applies on larger volumes on top of that. Compare the same kind of token: uncached input, cache reads and cache writes are billed at different rates.
- Do I need a subscription, and are there weekly quotas?
- No. Billing is pay-as-you-go from one balance: top it up, work, top it up again. There are no weekly or daily subscription quotas here, and the spending limits are the ones you set yourself on each key.
- When does GPT-6 Luna arrive?
- We open a model to requests only after its route is verified, so Luna is not available on Kumo yet. When it arrives, it will show on the models page and in the blog.
- What happens to GPT-5.6 Sol?
- OpenAI has not announced its retirement, and it stays in the Kumo catalog. But GPT-6 Sol costs half as much and is stronger on the vendor's tests, so switching is worth it — changing the model name in the request is all it takes.
Sources
- OpenAI's announcement: Introducing GPT-6 Sol and Luna — benchmarks, pricing, caching and safety data.
- Vendor documentation: the GPT-6 Sol page — context window, output limit, knowledge cutoff, cache and long-context rates.
- Caching: Better prompt caching for GPT-6 — what changed in prompt caching.
- Kumo pricing: the models page — the rates in force for every model in the catalog.