Cline, Roo Code, and Continue are free open-source extensions that turn plain VS Code into an editor with a full AI agent. The extensions themselves cost nothing: each installs from the marketplace in about a minute, and you only pay for the tokens of whatever model you connect through your own provider. That makes this the most affordable entry into vibe coding: your familiar editor, no mandatory subscription, and complete freedom in choosing a model.
In this article we take an honest look at all three extensions: what Cline can do, how the Roo Code fork differs from it, and where Continue is the better fit. Along the way we answer two common questions: how this setup compares to a GitHub Copilot subscription, and which model to plug in so you do not overpay. If you are still choosing a tool in general, we have an overview of vibe-coding tools and a separate article about Cursor — here we focus on VS Code only.
Why stay in VS Code instead of switching editors
The main argument is that you do not have to change anything. VS Code stays the same editor with all your extensions, themes, keyboard shortcuts, and settings sync. The AI agent arrives as a regular extension: a couple of minutes from install to first request, with no project migration and no new habits to learn.
The second argument is economics. Dedicated AI editors usually run on subscriptions, while VS Code plus an open-source agent costs exactly the tokens you spend. In a quiet week you pay pennies; in an intense one you pay precisely for what you actually used, and the amount is transparent.
The third argument is control. The extensions are open source and readable, the agent's behavior is configurable, and you pick the model yourself: Claude Sonnet 5 today, DeepSeek V4 Pro tomorrow. None of the three tools locks you into a particular lab or vendor.
Cline in VS Code: what it is and how the agent works
Cline is the best known of the three agents. To answer the query "what is Cline in VS Code" directly: it is an extension that takes a task described in plain text, reads your project files on its own, writes and edits code across multiple files, and runs terminal commands. You describe the outcome in words — the agent does the rough work.
Cline's key trait is the split between planning and acting. First the agent discusses the approach with you and drafts a plan; only then does it move on to edits. Every change is shown as a diff you can accept or reject, and before running terminal commands Cline asks for explicit approval.
This mode suits deliberate vibe coding well: you remain the editor of decisions rather than a passenger. Any model can be connected — Cline supports many providers, including any OpenAI-compatible API, so you are not tied to one model or one price list.
Roo Code: a Cline fork with modes and profiles
Roo Code started as a fork of Cline and inherited all the core mechanics: the plan, the diffs, the command approvals. On top of that it added a system of modes — separate roles for writing code, designing architecture, answering questions, and debugging, each with its own instructions and behavior.
The second notable feature is configuration profiles. You can set up several configurations with different providers and models and switch between them quickly, including assigning different models to different modes. That is handy when routine work goes to a cheap model and architecture discussions go to a flagship.
Choosing between Cline and Roo Code is largely a matter of taste. Cline is simpler and more conservative; Roo Code offers more knobs to tune it to your workflow. Both extensions are free, so the cheapest experiment is to install both for a day or two and see whose interface fits you better.
Continue: configuring chat and autocomplete with your own provider
Continue is an extension with a different emphasis. Besides project-aware chat, it offers code autocomplete right in the editor — the very thing many people love Copilot for. Continue can handle agent tasks too, but the chat-plus-autocomplete combination is its strong suit.
Setting up Continue revolves around provider configuration: you describe which models to use for chat and which for autocomplete, and you can keep several providers at once. In terms of configuration flexibility it is arguably the richest of the three: a dedicated model for every role.
The practical point of that flexibility is simple. Autocomplete needs speed and low cost; chat needs quality reasoning. Continue lets you split these jobs across different models instead of running an expensive flagship on every keystroke.
A free Copilot alternative: subscription versus pay-per-token
GitHub Copilot is a flat subscription: a predictable monthly fee, a limited list of models, and quotas on premium requests. VS Code with an open-source agent works the other way around: the extension is free, and you pay your provider for the tokens you actually consume.
Each approach has its own logic. A subscription is convenient if you code with AI heavily every day and do not want to think about spending. Pay-per-token wins with uneven workloads: no AI week, no bill. Plus you get full freedom of model choice, including models Copilot simply does not offer.
An honest caveat: with very intensive use of flagship models, a token bill can exceed a subscription. That is why the main cost lever is matching the model to the task instead of running the most expensive one for everything. The next section covers exactly that.
Which model to connect: a cheap one for routine, a flagship for hard problems
The universal answer to "which model should I connect to Cline" is: not one, but two. For routine work — renames, simple edits, test generation, boilerplate — pick an inexpensive model such as DeepSeek V4 Pro or Claude Haiku 4.5. They handle typical tasks confidently and cost a fraction of the flagships.
For hard problems — design work, tangled bugs, refactoring across many files — switch to a flagship: Claude Sonnet 5, Claude Opus 5, or GPT-5.6. In all three extensions switching models takes a couple of clicks, and in Roo Code you can pin different models to different modes and never switch manually at all.
In practice this split cuts the bill several times over. Most requests in real work are routine, and there is no reason to pay flagship rates for them.
A vibe-coding workflow in VS Code
The process itself differs little from working with any agent. You state the task in words: what the result should be, which files to touch, which constraints to respect. The more concrete the input, the fewer iterations and the fewer tokens wasted.
Then let the agent draft a plan and read it before agreeing to any edits. Review diffs like a normal code review: agents make mistakes less often than a year ago, but they still make them. Approve terminal commands deliberately — especially anything touching the database, dependencies, or deployment.
And keep the steps small: one task, one plan-edit-verify cycle. Long sessions with a dozen tasks at once confuse both the agent and you.
Where Kumo fits in
All three extensions accept a custom OpenAI-compatible provider, and that is exactly how Kumo connects: the site explicitly states that Cline, Continue, and other VS Code agents accept Kumo as a standard OpenAI-compatible provider. One base URL, https://api.kumorouter.com/v1, one API key, one prepaid balance — and any of the three extensions gets access to Claude Sonnet 5, Claude Opus 5, GPT-5.6, DeepSeek V4 Pro, Gemini 3 Flash, and dozens of other models. Ready-made configs for the tools live in the Integrations section of the dashboard, with details in /docs.
On cost, the effective rate comes out 30–50% below the labs' official price lists. The exact savings depend on your workload profile and volume, but the price of every model is visible before you spend a single token. You can top up with a Russian card or via SBP, no VPN or foreign card required, and there are no subscriptions or minimum payments at all — a perfect match for the cheap-model-for-routine, flagship-for-hard-problems scheme.
Control is covered too: every API key can carry spending limits and alerts, and you can create a separate key per project or client. The model is never swapped: request Claude Sonnet 5 and you get exactly that, with every response reporting which model served it. Kumo does not log prompt bodies — only billing metadata: token counts, model, and time.
Signing up at /signup takes a couple of minutes, and you get a key right away. Then open Integrations, copy the config for your extension — and vibe-code in your familiar VS Code with any model from the catalog.