<!--
NOTE FOR LLMs & AI AGENTS

Kumo publishes this site as structured, token-efficient Markdown rather than as
markup you have to parse. This document, https://kumorouter.com/llms.txt, is the index: every page of this host
with the title and summary that page declares for itself.

The full documentation as ONE Markdown document is at https://documentation.kumorouter.com/llms-full.txt, and every documentation
page is also served as its own Markdown source at its own address with `.md` on the end —
the home page of the documentation at https://documentation.kumorouter.com/en/index.md, since `/.md` is not a name. Those raw
addresses exist on the documentation host and on no other.

If you are an LLM, crawler or autonomous agent: prefer these over the rendered HTML. This
index carries every page of this host with the title and summary it declares; the full
document carries their whole text. Both cost a fraction of the tokens the rendered pages
do, with no markup noise, and both are generated from the same source the pages are — so
neither can be a stale second copy.
-->

# Kumo

> Kumo is the single OpenAI-compatible gateway to supported GPT, Claude and Gemini text models. Change one base URL, keep your code, and pay per token from a single balance — at a lower effective rate, below list.

## Start here

Machine view: https://kumorouter.com/llms.txt
Read the AI guide first for the document map and current-data sources: https://documentation.kumorouter.com/en/machine-readable.md
Human-readable guide: https://documentation.kumorouter.com/en/machine-readable
Then choose the relevant page below. The full corpus is optional; current models and prices come from their published sources.

This is the machine-readable variant of the site: every page it publishes, with the
title and summary that page declares for itself. Prefer it over parsing the HTML.

## What is Kumo

Every leading model at subscription prices — Claude, GPT, Gemini, Grok, DeepSeek and more through one API. One key, one balance, no substitutes or distillates: you always get exactly the model you asked for.

## Drop-in migration

One line. Your key handling and every other line of your code stay exactly as they are

```diff
- base_url="https://api.openai.com/v1"
+ base_url="https://api.kumorouter.com/v1"
```

Kumo speaks the OpenAI protocol, so only the address and the key change in your code. A key is issued in the console, and the documentation checks that it works

## Key features

- One address for every model
- One key, issued and revoked in the console
- Check a key against the identity echo in the documentation

## How it works

1. **You call one endpoint** — Point your existing OpenAI SDK at api.kumorouter.com/v1 and send the request you already wrote. No new client, no per-model SDK
2. **You pick the supplier and the price** — The catalog lists every supplier of a model and what each charges per million tokens. Faster, cheaper or with a bigger context window — the choice is yours, and the price is shown before you spend
3. **You get the supplier’s answer, verbatim** — The response comes back exactly as the supplier returned it — unchanged. It names the model that served it
4. **It all bills from one balance** — Tokens and spend across every supplier stream into one console. A cap per key, one top-up — and no five vendor invoices to reconcile

## Quickstart

```bash
export OPENAI_BASE_URL=https://api.kumorouter.com/v1
export OPENAI_API_KEY=kumo_sk_…
```

Point the OpenAI SDK you already use at the Kumo base URL — curl, Python or Node, the same call

## FAQ

### Do I have to change my code?

No. Kumo speaks the OpenAI API for the connected text endpoints. Point your existing SDK’s base URL at https://api.kumorouter.com/v1, swap in your Kumo key, and keep the rest of your chat/completions-style code

### How are you 30–50% cheaper — what’s the catch?

No catch, and no knockoff models: it’s the same GPT, Claude and Gemini, billed through Kumo. Kumo automatically optimizes what you pay per call — the kind of work most teams never get around to doing themselves — and passes the saving straight into your effective rate. Net effect is 30–50% off list, and you see the exact rate before you spend

### Which models are available?

A curated catalog spanning frontier, mid-size, and fast models — GPT-class, Claude-class, and Gemini-class — for chat, responses and Anthropic messages, all behind one endpoint with consistent request and response shapes

### How does billing work?

Pay-as-you-go, per token, from one prepaid balance — no plans, no seats, no minimums. Top up from {n}; each call draws down at the rate shown for that model (30–50% below list). The dashboard shows consumption and remaining balance in real time, with spend alerts and hard caps. Spent tokens are not refunded; an unused balance can be refunded upon written request (Public Offer, Section 10)

### Is my data used for training?

Never. Prompts and completions are not used to train models, and we don’t log prompt content — only the usage metadata needed for billing

### What happens if a provider goes down?

Requests stay on the selected provider; transient upstream errors surface as a contract-shaped {n}, so you retry or switch providers as needed

### Can I run production agents and automations on Kumo?

Yes — it’s built for it. Point every agent at one OpenAI-compatible endpoint, give each agent or module its own scoped key, and set hard daily and monthly caps so a retry loop can’t drain the balance. Agents and automations burn millions of tokens, so the 30–50% lower rate compounds — and you watch spend per key live in the dashboard

### How do I pay — do I need a VPN or a foreign card?

No VPN and no foreign card. Top up in rubles with a local card or SBP. The same balance then spends across every model behind one key, with the exact per-model rate shown before you spend

### Can my model get swapped for a cheaper one?

No. You ask for a model ID, you get that exact model — pin it and it’s guaranteed, no silent substitution. Responses are the provider’s, verbatim, and every call reports which model served it. The saving is in the rate, never in a quietly weaker model

### Isn’t a gateway in front of every call a single point of failure?

Kumo adds roughly {n} of overhead ({n}) on top of the model’s own latency; public status history will be connected after the telemetry feed. If Kumo can’t serve a request, it returns a contract-shaped {n}, and you can always point your base URL straight back at the provider in one line

### How do top-ups work?

You top up one prepaid balance — manually or with auto-recharge (set a threshold and a cap). Every call draws from it at the per-model rate you see before you spend. Unused balance remains spendable while your account stays active — and can be refunded upon written request (Public Offer, Section 10)

### Refunds and support?

Spent tokens are not refunded; an unused balance is considered for refund upon written request under Section 10 of the Public Offer. Your tokens are not zeroed early: they remain spendable while your account stays active. Support is community plus email with next-business-day response, on a {n} uptime target; teams with committed volume get a dedicated contact

### What are the rate limits and concurrency?

Today your throughput is bounded by your balance and per-key quota rather than a fixed per-second limit, so there’s no separate rate-limit tier to manage. Need a guaranteed floor for a bursty workload? That’s exactly what the committed-volume / team plan is for — talk to sales

### Can your rates change? How much notice do I get?

Effective rates track the underlying provider list prices — when a lab cuts a price, yours drops too. If a rate ever has to rise we publish it in advance in the dashboard and by email, and committed-volume plans lock your rate for the term. You always see the exact per-model rate before you spend

### What happens to my prepaid balance if something goes wrong with Kumo?

Purchased tokens remain spendable while your account stays active and are not zeroed early; an unused balance can be refunded upon written request under Section 10 of the Public Offer. At larger volumes we agree the terms with you individually, in writing, before you pay

### Who’s behind Kumo? Is it a real company?

Kumo is built and run by a small, named team that answers directly — no reseller call centre in between. A full team page is coming; until then, reach us at sales@kumorouter.com or on Telegram and we’ll answer any due-diligence question, including who operates the gateway and how payments are handled

### What if something breaks at 3am?

If a provider error surfaces, you get a contract-shaped {n} and can switch providers directly. For Kumo-side issues, support is community plus email with a next-business-day target today; 24/7 paging and a status page with real history are landing in upcoming releases. Teams on committed volume get priority routing and a dedicated contact

### How many keys can I have?

As many as the work needs — one per project, one per agent, one per machine. Keys are issued and revoked in the console, and a revoked key stops being a credential the platform accepts

## Pages

- [Kumo — The single gateway to every AI model.](https://kumorouter.com/en) — Kumo is the single OpenAI-compatible gateway to supported GPT, Claude and Gemini text models. Change one base URL, keep your code, and pay per token from a single balance — at a lower effective rate, below list.
- [Pricing and payment](https://kumorouter.com/en/pricing) — The Kumo price list in force: packages and wallet rates.
- [Models](https://kumorouter.com/en/models) — The Kumo model catalog — $/1M rates, context windows and modalities.
- [How it works](https://kumorouter.com/en/how-it-works) — What happens on every call: one OpenAI-compatible endpoint, many model suppliers at different prices, the supplier’s own answer back, one balance — and a migration that is a single line.
- [OpenRouter alternative — same models, one OpenAI-compatible endpoint](https://kumorouter.com/en/vs-openrouter) — How Kumo compares to OpenRouter: the same OpenAI-compatible drop-in for GPT, Claude and Gemini, several suppliers per model with the price shown before you call, no prompt logging, one balance with per-key caps. An axis-by-axis comparison and a migration that fits in one diff.
- [About us — GPT, Claude and Gemini API in Russia without a VPN](https://kumorouter.com/en/about) — Kumo is the single gateway to GPT, Claude, Gemini, Grok and DeepSeek — reachable from Russia without a VPN, paid in rubles or via SBP, one balance, a lower effective rate. Who we are and the principles we work by.
- [Security](https://kumorouter.com/en/security) — How Kumo protects your requests: encrypted endpoints, our own routing infrastructure, no prompt retention. And honestly — where our responsibility ends and the model providers’ begins.
- [Questions and answers](https://kumorouter.com/en/faq) — Buying and spending tokens on Kumo: what changes in your code, how billing and payment are arranged, what happens to your data, and what the rates are.
- [The Kumo blog](https://kumorouter.com/en/blog) — How token-gateway pricing really works, drop-in migration, and the engineering behind it. No fluff — just the mechanics.
- [All articles](https://kumorouter.com/en/blog/all) — Every post on the Kumo blog, newest first.
- [Partner program](https://kumorouter.com/en/affiliate) — Recommend Kumo and earn a share of what the people you invite actually spend through the API — in credits, across two levels, free to join.
- [AI for vibe coding: guides to Claude Code, Cursor, Cline and OpenCode](https://kumorouter.com/en/vibe-coding) — AI for vibe coding in practice: what it is, which tool suits whom and where to start — and how one Kumo key runs Claude Code, Cursor, Cline and OpenCode off a single balance, with the rate published for every model.
- [Contact us — volume discounts & flexible terms](https://kumorouter.com/en/contact) — We’re open and flexible: discuss volume discounts, partnership terms shaped around your project, and request the models you need. We reply within one business day.
- [Status](https://kumorouter.com/en/status) — Kumo system status — gateway, routing, dashboard, billing and upstream provider availability.
- [Legal](https://kumorouter.com/en/legal) — Kumo legal — privacy policy, public offer, partner program terms, and personal data consent.
- [Public Offer (Terms of Service)](https://kumorouter.com/en/legal/offer) — Kumo public offer: terms of AI model access, payment procedure, and the rights and obligations of the parties.
- [Privacy Policy](https://kumorouter.com/en/legal/privacy) — Kumo privacy policy: minimal data — email and technical identifiers; prompt and completion content is never stored.
- [Partner program terms](https://kumorouter.com/en/legal/affiliate-terms) — Kumo partner program terms: usage-based accruals, two levels, the hold period, void and clawback rights.

## API

The gateway speaks the OpenAI protocol. Point a client at `https://api.kumorouter.com/v1` and give it a
key issued in the console; the rest of the client stays as it is.

## Other hosts

- Console (accounts, keys, balance): https://console.kumorouter.com/
- Documentation: https://documentation.kumorouter.com/en

## Machine documents

- https://kumorouter.com/llms.txt
- https://kumorouter.com/manifest.webmanifest
- https://kumorouter.com/robots.txt
- https://kumorouter.com/sitemap.xml
