Skip to content
← Blog
8 min read

Grok 4.7: what changed and how to use it on Kumo

xAI shipped Grok 4.7 — nearly twice Grok 4.6's Terminal-Bench score at Grok 4.6's price of $2 and $6 per million tokens. Kumo pricing, the benchmarks, and two checks that prove the original model is answering.

grokgrok 4.7model verificationpricingxai

On September 21, 2026 xAI released Grok 4.7 — a new model for coding and knowledge work that costs exactly what Grok 4.6 costs and runs at the same speed. For the same $2 and $6 per million tokens it solves nearly twice as many Terminal-Bench tasks as its predecessor. It is already available on Kumo: the same API, one balance topped up by card or crypto, and no weekly quotas. And it is the original Grok 4.7, not a cheaper model wearing its name — below are two checks you can run yourself to confirm it.

Context

500K tokens

text and images in

xAI price

$2 / $6

input / output per 1M, under 200K context

Kumo price

$1.1 / $3.3

45% off the vendor's list

Terminal-Bench 4.0

37.6%

Grok 4.6 scores 20.3%

What Grok 4.7 is

Grok 4.7 is, in xAI's own words, its most capable model for coding and knowledge work. It stays on difficult tasks longer, checks its own work more carefully and handles long context better. xAI also trained it to natively understand its own Grok Bot agent harness, which is where the gains on conversational tasks and general knowledge work come from.

The context window is 500K tokens. The model takes text and images in (JPG and PNG up to 20 MiB) and returns text. It supports function calling, structured outputs and reasoning with an adjustable effort level — low, medium, high or xhigh, with high as the default. The knowledge cutoff is May 2026. The vendor's identifier is grok-4.7 and Kumo's catalog lists it as xai/grok-4.7; xAI's grok-4.7-latest alias will move to the next version in time, so pin the fixed name for reproducible runs.

The practical meaning of this release is that the price did not move. A new version usually arrives with a new price list and the upgrade needs a spreadsheet. Not here: xAI serves Grok 4.7 at the same price and speed as Grok 4.6, so switching is a change of model name in the request. On Kumo Grok 4.7 is even cheaper than its predecessor: $1.1 and $3.3 against Grok 4.6's $1.2 and $3.6. For those who care more about latency, xAI also serves a fast variant — twice the output speed at twice the price.

What changed since Grok 4.6

Six benchmarks from xAI’s announcement, ordered by Grok 4.7’s own score. xAI ran all four models on its own harness; Grok 4.7’s DeepSWE score was run at high reasoning effort.
BenchmarkGrok 4.7Grok 4.6
Terminal-Bench 4.037.6%20.3%
EEBench64.0%53.0%
HealthBench Professional56.7%48.5%
CursorBench 4.046.3%40.4%
DeepSWE v1.1 (high effort)71.0%65.2%
Harvey Legal19.6%15.8%
xAI's own figures from the announcement of September 21, 2026. The methodology and run conditions are on the announcement page.

The largest jump is, again, where the model works on its own. Terminal-Bench does not measure a single answer, it measures carrying a task to the end in a terminal, and seventeen points there are the difference between an agent that abandons the task halfway and one that hands it in. All six benchmarks moved up, by between nearly four and seventeen points.

An honest frame: Grok 4.7 is not the new leader in coding. On the same Terminal-Bench xAI measures Claude Fable 5.1 at 57.9%, and on CursorBench at 51.8% against 46.3%. But Fable 5.1's output costs $50 per million tokens against $6 — more than eight times as much. Grok 4.7 caught up with GPT-5.6 Sol on Terminal-Bench and passed it on CursorBench, EEBench and Harvey Legal, while trailing it on DeepSWE and HealthBench.

And a caveat about other vendors' numbers: the Fable 5.1 and GPT-5.6 Sol results here are xAI's runs, not their vendors'. Anthropic publishes 55.8% for the same Fable 5.1 on Terminal-Bench 4.0: different harnesses give different numbers, so compare models inside one table, never across announcements.

One more change is the safety stack. xAI calls Grok 4.7 its most jailbreak-resistant model yet and rebuilt its refusal system from scratch: on HackerBench v0.3 the model lets 3.3% of risky prompts through. If your workload borders on security — penetration testing, malware analysis — run your own prompts on the new model before moving traffic over.

What grew at the same price

Gains in percentage points, not in percent of the old score. xAI’s price and speed are the same as Grok 4.6’s.

What Grok 4.7 costs

What is billedxAIKumo
Input, context under 200K$2$1.1
Input, context 200K and over$4$1.1
Output, context under 200K$6$3.3
Output, context 200K and over$12$3.3
Cache reads, under / over 200K$0.5 / $1$0.3 ($0.275)
USD per 1M tokens. xAI doubles the rate once a request's context reaches 200K tokens; Kumo's price list, as for Grok 4.6, has no separate long-context rate. Kumo's rate for Grok 4.7 is 45% off the vendor's list; the figures in force are always in the price list on the models page.
The hollow dot is the same model at the Kumo rate. Further left and higher means more work per dollar; all three scores are xAI’s runs.

Grok 4.7 is already available on Kumo

The model is wired into the shared catalog and runs off the same key and the same balance as the rest: no separate subscription, no waiting for a quota to refill, no second invoice. The balance is topped up by Russian card, through SBP or in crypto, and spent on any model in the catalog, from Grok to Claude and GPT.

Grok on Kumo speaks the OpenAI-compatible protocols — Chat Completions and Responses. Clients point at https://api.kumorouter.com/v1 and pass the key as Authorization: Bearer. If you already run Grok 4.6, changing the model name to xai/grok-4.7 is all it takes.

One honest limit: every key carries its own spending caps and alerts, and we do not log request bodies — only billing metadata: the model, the token counts and the time.

export KUMO_API_KEY="kumo_sk_…"

# OpenAI-compatible clients: Chat Completions and Responses
export OPENAI_BASE_URL="https://api.kumorouter.com/v1"
export OPENAI_API_KEY="$KUMO_API_KEY"

The model is the original — and you can prove it

The main fear when buying access through an intermediary is that a cheaper model answers under a new model's name. With Grok 4.7 the fear is easy to understand: at xAI the previous version costs the same and has the same window, so a swap to Grok 4.6 would show neither in the vendor's price list nor in the context length — only in the quality of the work.

So we do not ask you to take our word for it. On Kumo the model is never swapped: ask for xai/grok-4.7 and Grok 4.7 answers, and every response names who produced it. Below are two checks you run yourself, with your own key: the first takes a minute and catches a careless substitution, the second leans on something a short-window model cannot fake.

An honest caveat up front: none of these checks is a cryptographic proof, and we are not pretending otherwise. Together they rule out a careless substitution and any model whose window is shorter than 500K tokens; telling Grok 4.7 from Grok 4.6 is best done on your own tasks — the table at the end of this section says what each check proves.

Check 1. The catalog and the model echo in the response

# the model must appear in the catalog under its identifier
curl -s https://api.kumorouter.com/v1/models \
  -H "Authorization: Bearer $KUMO_API_KEY" | grep -o '"id":"xai/grok-4.7"'

# the model field in the response must name Grok 4.7
curl -s https://api.kumorouter.com/v1/chat/completions \
  -H "Authorization: Bearer $KUMO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"xai/grok-4.7","messages":[{"role":"user","content":"ping"}],"max_tokens":16}' \
  | python3 -c "import sys,json; d=json.load(sys.stdin); print(d['model'], d['usage'])"

A minute of work, and it rules out the most common case: a catalog where a pretty name is an alias pointing somewhere else. Watch two places — the identifier in the model list and the model field in the response body. Both must name Grok 4.7, not 4.6 and not someone else's model.

What this check does not prove: a tidy intermediary can return the right string while still substituting the model. That is why it is the first check rather than the only one.

Check 2. The 500K window and a needle in the haystack

python3 - <<'PY'
import json, os, urllib.request

needle = " The pass phrase of the day is kumo-grok-47-window. "
filler = "This is filler text whose only job is to fill the context window. " * 10000
question = " Question: state the pass phrase of the day on one line, with no explanation."
prompt = filler + needle + filler + question

request = urllib.request.Request(
    "https://api.kumorouter.com/v1/chat/completions",
    data=json.dumps({
        "model": "xai/grok-4.7",
        "messages": [{"role": "user", "content": prompt}],
        "max_tokens": 64,
    }).encode(),
    headers={
        "Authorization": "Bearer " + os.environ["KUMO_API_KEY"],
        "Content-Type": "application/json",
    },
)
with urllib.request.urlopen(request, timeout=900) as response:
    body = json.load(response)
print(body["usage"]["prompt_tokens"], "input tokens")
print(body["choices"][0]["message"]["content"])
PY

A model with a short window cannot pass this one physically. A 128K or 200K window will not accept a request of several hundred thousand tokens — an over-context error comes back instead, and the substitution exposes itself. A model with a real 500K window accepts the request and states the phrase hidden exactly in the middle of it.

Watch both numbers: the answer and prompt_tokens. If the input counter is well below what you sent, the request was quietly truncated halfway through. If prompt_tokens goes past 500K, lower the filler multiplier: for the check to be fair the request has to fit the window.

The check is nearly free: about three hundred thousand input tokens at Kumo's rate cost about a third of a dollar. One caveat: Grok 4.6 has the same window, so this check rules out other vendors' models but not the previous version.

What each check proves, and what it does not

CheckWhat it provesWhat it does not prove
Catalog and model echoThat the route is declared under the real nameNothing about who actually answered
500K window and needleThat the window really is 500K tokensDoes not tell Grok 4.7 from Grok 4.6
Request log in the consoleThat you were billed for that exact modelDoes not verify the content of the answer
Each check covers its own blind spot; telling versions inside the family apart takes your own tasks.

How to connect Grok 4.7 in five minutes

  1. Sign up in the console and top up the balance — by card, through SBP or in crypto.
  2. Create a key and give it a spending cap right away: one key per project saves you the argument about where the money went.
  3. Set the gateway address https://api.kumorouter.com/v1 in your OpenAI-compatible client.
  4. Set the model to xai/grok-4.7 and send the first request.
  5. Run both checks from this article before you start spending in earnest.
curl -s https://api.kumorouter.com/v1/chat/completions \
  -H "Authorization: Bearer $KUMO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "xai/grok-4.7",
    "messages": [
      {"role": "user", "content": "Find every place in this repository where an environment variable is read and put them in a table."}
    ],
    "max_tokens": 2048
  }'

Frequently asked questions

Should I move from Grok 4.6 to Grok 4.7?
Yes, if Grok 4.6 served you well: at xAI the price and speed are the same, on Kumo Grok 4.7 is even cheaper, and all six published benchmarks went up. Switching is a change of model name in the request. The one exception is prompts close to the safety policies: the new model's safeguards are stricter, so run those first.
Grok 4.7 or Claude — which should I pick?
By xAI's own numbers, Claude Fable 5.1 handles the heaviest agentic work in the terminal and the editor better, but its output costs more than eight times as much. Grok 4.7 is the workhorse: high-volume tasks, long documents up to 500K tokens, legal and analytical work. The sensible order is to start on the cheaper model and move up where your own evals show it falling short.
Is this really the same Grok 4.7 as xAI's own?
Yes: the route leads to the original model, not to a stand-in. You can confirm it two ways from the section above — the catalog and the identifier echo, and the 500K window — and the request log in the console shows which model you were billed for. You run both checks yourself, with your own key.
Is there a surcharge for context over 200K tokens?
At xAI, yes: from 200K tokens of context the rates double. Kumo's price list has no separate long-context rate for Grok — the price in force is always on the models page.
Why is it cheaper on Kumo than at the vendor?
For Grok 4.7 the discount is 45% off the vendor's list on ordinary input and output tokens, and a package ladder applies on larger volumes. Compare the same model version and the same kind of token: uncached input and cache reads are billed at different rates.
Do I need a subscription, and are there weekly quotas?
No. Billing is pay-as-you-go from one balance: top it up, work, top it up again. There are no weekly or daily subscription quotas here, and the spending limits are the ones you set yourself on each key.
Will there be a fast Grok 4.7?
xAI serves it separately: twice the output speed at twice the price. We connect new vendor models as they ship — watch the models page and the blog.

Sources

  • xAI's announcement: Introducing Grok 4.7 — benchmarks, pricing, the fast variant and safety data.
  • Vendor documentation: Grok 4.7 — identifier, context window, modalities, effort levels and the context-length price tiers.
  • xAI models overview: Models and pricing — the knowledge cutoff and the comparison with Grok 4.6.
  • Kumo pricing: the models page — the rates in force for every model in the catalog.

Start building on Kumo today

One base URL, one balance, every model — at an effective rate you can see before you spend a token