About Kumo
A billing-and-routing layer, built to be boring
Kumo is one gateway in front of the GPT, Claude and Gemini text models your team already uses: one base URL, one key, one balance — at a lower effective rate than list price. No new SDK, no lock-in, no drama. That’s the whole company, on purpose
What we build
One gateway, everything behind it
Four things, done properly — everything else on the site is detail
A drop-in gateway
Point the OpenAI SDK you already use at https://api.kumorouter.com/v1 and keep every other line of your code. Each request is routed to the model’s upstream provider and the answer comes back verbatim
One balance for the team
Top up once and spend it across every supported model. No per-provider cards, seats or minimums — you pay only for tokens you actually use
A lower effective rate
Routing and caching optimizations bring your effective per-token cost under list price — same models, same API, savings shown per call
Metering you can read
Tokens, models and costs per key and per day in the dashboard — the numbers your finance team asks for, without a spreadsheet safari
How we work
Principles over promises
The same rules run through the product, the pricing page and this site
Verbatim, always
We never rewrite, downgrade or substitute a model behind your back. The response you get is the provider’s response, and the serving model is named on every call
Not a data lake
Prompt and completion bodies are processed to serve the request, then dropped — not stored, not logged, never used for training. The details live on the Security page
Only what’s live
If a feature, certification or metric isn’t real yet, the site says so — status snapshots are labelled as snapshots, and roadmap items are labelled roadmap
Small surface, fast answers
A deliberately small product means questions reach a person who can fix things — security and billing questions are usually answered within one business day
At a glance
The short version
| Fact | Detail |
|---|---|
| Product | OpenAI-compatible API gateway for GPT, Claude and Gemini text models |
| API base | https://api.kumorouter.com/v1 — drop-in for the OpenAI SDK |
| Pricing model | Pay-as-you-go token packages; no seats, no minimums |
| Uptime target | 99.9% target; public status monitoring is being connected |
| Data stance | No prompt logging, no training on your data |
| Russia | Works without a VPN or foreign card — pay in rubles by card or via SBP |
| Talk to us | sales@kumorouter.com support@kumorouter.com @kumo_api |
Kumo in Russia
GPT, Claude and Gemini API — in Russia, without a VPN
Kumo is API access to modern language models from Russian networks: no VPN, no foreign card, no separate account with each provider
One OpenAI-compatible endpoint — https://api.kumorouter.com/v1 — gives access to 30 text models: Claude Opus 4.8, Claude Sonnet 5, Claude Sonnet 4.6, Claude Sonnet 4.5, Claude Opus 5, Claude Opus 4.7, Claude Opus 4.6, Claude Opus 4.5, Claude Haiku 4.5, Claude Fable 5.1, Claude Fable 5 and Claude Opus 5.5 by Anthropic, GPT-5.6 Terra, GPT-5.6 Sol, GPT-5.6 Luna, GPT-5.5, GPT-5.3 Codex Spark, Codex Auto Review, GPT-6 Astra and GPT-6 Sol by OpenAI, Gemini 3.8 Flash High, Gemini 3.7 Flash High, Gemini 3.6 Flash High, Gemini 3.1 Flash Lite and Gemini 3 Flash Preview by Google, Grok 4.6, Grok 4.5 and Grok 4.7 by xAI and GLM-5.3 and GLM-5.2 by Zhipu. The same key works in Cursor, Cline, Continue, n8n, Make and any code built on the official OpenAI SDK
Everything runs directly on Russian networks: top up in rubles with a local card or via SBP. No VPN, foreign card or virtual number — sign up with an email, add funds and call GPT, Claude or Gemini over the API the same day
If you are looking to buy ChatGPT API, Claude API or Gemini API access in Russia, Kumo covers it with one key: a single balance across every model, rates 30–50% below provider list prices, and token metering you can actually read in the dashboard
Models available via the API right now
| Model | Provider | Inputper 1M tokens | Outputper 1M tokens | Context |
|---|---|---|---|---|
| Claude Opus 4.8 | anthropic | 3 USD | 15 USD | 1M |
| Claude Sonnet 5 | anthropic | 1.2 USD | 6 USD | 1M |
| Claude Sonnet 4.6 | anthropic | 1.8 USD | 9 USD | 1M |
| Claude Sonnet 4.5 | anthropic | 1.8 USD | 9 USD | 1M |
| Claude Opus 5 | anthropic | 3 USD | 15 USD | 1M |
| Claude Opus 4.7 | anthropic | 3 USD | 15 USD | 1M |
| Claude Opus 4.6 | anthropic | 3 USD | 15 USD | 1M |
| Claude Opus 4.5 | anthropic | 3 USD | 15 USD | 200K |
| Claude Haiku 4.5 | anthropic | 0.6 USD | 3 USD | 200K |
| Claude Fable 5.1 | anthropic | 6 USD | 30 USD | 1M |
| Claude Fable 5 | anthropic | 6 USD | 30 USD | 1M |
| GPT-5.6 Terra | openai | 1.2 USD | 7.2 USD | 1.1M |
| GPT-5.6 Sol | openai | 2.4 USD | 12 USD | 1.1M |
| GPT-5.6 Luna | openai | 0.1 USD | 0.7 USD | 1.1M |
| GPT-5.5 | openai | 3 USD | 18 USD | 1.1M |
| GPT-5.3 Codex Spark | openai | 1.1 USD | 8.4 USD | 128K |
| Codex Auto Review | openai | 3 USD | 18 USD | 1.1M |
| Gemini 3.8 Flash High | 0.5 USD | 2.3 USD | 1M | |
| Gemini 3.7 Flash High | 0.5 USD | 2.3 USD | 1M | |
| Gemini 3.6 Flash High | 0.5 USD | 2.3 USD | 1M | |
| Gemini 3.1 Flash Lite | 0.2 USD | 0.9 USD | 1M | |
| Gemini 3 Flash Preview | 0.3 USD | 1.8 USD | 1M | |
| Grok 4.6 | xai | 1.2 USD | 3.6 USD | 500K |
| Grok 4.5 | xai | 1.2 USD | 3.6 USD | 500K |
| GLM-5.3 | zai | 0.8 USD | 2.6 USD | 1M |
| GLM-5.2 | zai | 0.8 USD | 2.6 USD | 1M |
| GPT-6 Astra | openai | 6 USD | 30 USD | 1.1M |
| Grok 4.7 | xai | 1.1 USD | 3.3 USD | 500K |
| Claude Opus 5.5 | anthropic | 2.2 USD | 11 USD | 1M |
| GPT-6 Sol | openai | 1.1 USD | 5.5 USD | 1.1M |
Kumo rates per 1M tokens (input / output) after the routing discount — a standard saving of 25–30% off list; committed volume adds up to 20 points more, reaching ~50% at the 1B-token tier. The full catalog with cached-input pricing lives on the models page
Want the longer version?
See exactly how a request travels through the gateway — or just ask us directly