How it works
One gateway. Many suppliers Your pick
Kumo is a single OpenAI-compatible endpoint in front of models from many suppliers. The same model costs a different amount at different suppliers; you see each one’s price and pick the one that fits your task and your budget. One key, one balance
How the service works
What happens on every call
You call one endpoint
Point your existing OpenAI SDK at api.kumorouter.com/v1 and send the request you already wrote. No new client, no per-model SDK
You pick the supplier and the price
The catalog lists every supplier of a model and what each charges per million tokens. Faster, cheaper or with a bigger context window — the choice is yours, and the price is shown before you spend
You get the supplier’s answer, verbatim
The response comes back exactly as the supplier returned it — unchanged. It names the model that served it
It all bills from one balance
Tokens and spend across every supplier stream into one console. A cap per key, one top-up — and no five vendor invoices to reconcile
Drop-in compatible
One line in. First call out
A real, playable terminal — type a command or click a step. The whole migration, the call, and the answer
Security
Read about security before your first call
Kumo is a routing-and-billing layer: your request goes on to the supplier you picked. So do not send sensitive data through Kumo — personal data, secrets, payment details. How we handle what passes through the gateway is described on the security page
FAQ
Questions, answered
The short version of how calls, payment and picking a supplier work
All questions and answers →Spin up your gateway today
One base URL, one balance, your whole team behind it. Free to start