Skip to content

About Kumo

A billing-and-routing layer, built to be boring

Kumo is one gateway in front of the GPT, Claude and Gemini text models your team already uses: one base URL, one key, one balance — at a lower effective rate than list price. No new SDK, no lock-in, no drama. That’s the whole company, on purpose

What we build

One gateway, everything behind it

Four things, done properly — everything else on the site is detail

A drop-in gateway

Point the OpenAI SDK you already use at https://api.kumorouter.com/v1 and keep every other line of your code. Each request is routed to the model’s upstream provider and the answer comes back verbatim

One balance for the team

Top up once and spend it across every supported model. No per-provider cards, seats or minimums — you pay only for tokens you actually use

A lower effective rate

Routing and caching optimizations bring your effective per-token cost under list price — same models, same API, savings shown per call

Metering you can read

Tokens, models and costs per key and per day in the dashboard — the numbers your finance team asks for, without a spreadsheet safari

How we work

Principles over promises

The same rules run through the product, the pricing page and this site

Verbatim, always

We never rewrite, downgrade or substitute a model behind your back. The response you get is the provider’s response, and the serving model is named on every call

Not a data lake

Prompt and completion bodies are processed to serve the request, then dropped — not stored, not logged, never used for training. The details live on the Security page

Only what’s live

If a feature, certification or metric isn’t real yet, the site says so — status snapshots are labelled as snapshots, and roadmap items are labelled roadmap

Small surface, fast answers

A deliberately small product means questions reach a person who can fix things — security and billing questions are usually answered within one business day

At a glance

The short version

FactDetail
ProductOpenAI-compatible API gateway for GPT, Claude and Gemini text models
API basehttps://api.kumorouter.com/v1 — drop-in for the OpenAI SDK
Pricing modelPay-as-you-go token packages; no seats, no minimums
Uptime target99.9% target; public status monitoring is being connected
Data stanceNo prompt logging, no training on your data
RussiaWorks without a VPN or foreign card — pay in rubles by card or via SBP
Talk to ussales@kumorouter.com support@kumorouter.com @kumo_api

Kumo in Russia

GPT, Claude and Gemini API — in Russia, without a VPN

Kumo is API access to modern language models from Russian networks: no VPN, no foreign card, no separate account with each provider

One OpenAI-compatible endpoint — https://api.kumorouter.com/v1 — gives access to 30 text models: Claude Opus 4.8, Claude Sonnet 5, Claude Sonnet 4.6, Claude Sonnet 4.5, Claude Opus 5, Claude Opus 4.7, Claude Opus 4.6, Claude Opus 4.5, Claude Haiku 4.5, Claude Fable 5.1, Claude Fable 5 and Claude Opus 5.5 by Anthropic, GPT-5.6 Terra, GPT-5.6 Sol, GPT-5.6 Luna, GPT-5.5, GPT-5.3 Codex Spark, Codex Auto Review, GPT-6 Astra and GPT-6 Sol by OpenAI, Gemini 3.8 Flash High, Gemini 3.7 Flash High, Gemini 3.6 Flash High, Gemini 3.1 Flash Lite and Gemini 3 Flash Preview by Google, Grok 4.6, Grok 4.5 and Grok 4.7 by xAI and GLM-5.3 and GLM-5.2 by Zhipu. The same key works in Cursor, Cline, Continue, n8n, Make and any code built on the official OpenAI SDK

Everything runs directly on Russian networks: top up in rubles with a local card or via SBP. No VPN, foreign card or virtual number — sign up with an email, add funds and call GPT, Claude or Gemini over the API the same day

If you are looking to buy ChatGPT API, Claude API or Gemini API access in Russia, Kumo covers it with one key: a single balance across every model, rates 30–50% below provider list prices, and token metering you can actually read in the dashboard

Models available via the API right now

ModelProviderInputper 1M tokensOutputper 1M tokensContext
Claude Opus 4.8anthropic3 USD15 USD1M
Claude Sonnet 5anthropic1.2 USD6 USD1M
Claude Sonnet 4.6anthropic1.8 USD9 USD1M
Claude Sonnet 4.5anthropic1.8 USD9 USD1M
Claude Opus 5anthropic3 USD15 USD1M
Claude Opus 4.7anthropic3 USD15 USD1M
Claude Opus 4.6anthropic3 USD15 USD1M
Claude Opus 4.5anthropic3 USD15 USD200K
Claude Haiku 4.5anthropic0.6 USD3 USD200K
Claude Fable 5.1anthropic6 USD30 USD1M
Claude Fable 5anthropic6 USD30 USD1M
GPT-5.6 Terraopenai1.2 USD7.2 USD1.1M
GPT-5.6 Solopenai2.4 USD12 USD1.1M
GPT-5.6 Lunaopenai0.1 USD0.7 USD1.1M
GPT-5.5openai3 USD18 USD1.1M
GPT-5.3 Codex Sparkopenai1.1 USD8.4 USD128K
Codex Auto Reviewopenai3 USD18 USD1.1M
Gemini 3.8 Flash Highgoogle0.5 USD2.3 USD1M
Gemini 3.7 Flash Highgoogle0.5 USD2.3 USD1M
Gemini 3.6 Flash Highgoogle0.5 USD2.3 USD1M
Gemini 3.1 Flash Litegoogle0.2 USD0.9 USD1M
Gemini 3 Flash Previewgoogle0.3 USD1.8 USD1M
Grok 4.6xai1.2 USD3.6 USD500K
Grok 4.5xai1.2 USD3.6 USD500K
GLM-5.3zai0.8 USD2.6 USD1M
GLM-5.2zai0.8 USD2.6 USD1M
GPT-6 Astraopenai6 USD30 USD1.1M
Grok 4.7xai1.1 USD3.3 USD500K
Claude Opus 5.5anthropic2.2 USD11 USD1M
GPT-6 Solopenai1.1 USD5.5 USD1.1M

Kumo rates per 1M tokens (input / output) after the routing discount — a standard saving of 25–30% off list; committed volume adds up to 20 points more, reaching ~50% at the 1B-token tier. The full catalog with cached-input pricing lives on the models page

Want the longer version?

See exactly how a request travels through the gateway — or just ask us directly