Roo Code runs several modes, not one

Per-seat licences for closed coding assistants scale with headcount, and the renewal quote grows every year. Roo Code offers a way out that most tool comparisons miss: it is one VS Code extension with five built-in modes, and each mode can run a different model. The rest of this guide gives each mode a model chosen for that mode's job, and ends with a tool-calling check you can run before any real work starts.

The five built-in modes are Code, Architect, Ask, Debug and Orchestrator. They are not cosmetic labels. Each mode carries its own tool access: Code and Debug get the full set (read, edit, command, mcp), Architect can edit markdown files only, Ask cannot modify files or run commands at all, and Orchestrator touches nothing directly because it delegates subtasks to other modes through the new_task tool.

ModeTool accessBuilt for
Coderead, edit, command, mcpwriting and refactoring code
Architectread, mcp, markdown-only editsystem design and planning
Askread, mcpexplanation without file changes
Debugread, edit, command, mcpsystematic troubleshooting
Orchestratordelegates via new_taskmulti-step coordination

Roo remembers the profile used with each mode and can link profiles to modes in the Prompts tab. A separate profile lets you choose model settings for each task type. Evaluate whether a smaller model meets your quality threshold before assigning it to a mode.

If your team also drives a terminal-based agent, the equivalent setup for opencode is covered in a separate guide: Using Lyceum Models in opencode: Custom Provider Setup. This article stays inside VS Code.

What you need before connecting

Nothing exotic. Roo Code is model-agnostic by design: it supports any provider offering an OpenAI-compatible API endpoint, which is why this setup works without forks, patches or a proxy service. You need three things.

  • Roo Code installed from the VS Code Extensions panel. The VS Code Marketplace is the recommended route; Open VSX serves VSCodium users.
  • An API key created in the dashboard at dashboard.lyceum.technology.
  • The model strings you plan to use, copied from the live roster rather than guessed. The published catalogue is cross-checked against the /openai/v1/models endpoint, which is the source of truth for what exists.

Copy the model string character for character from the roster. A string that is off by one segment fails with a model-not-found error, and a model that has been removed from the roster fails the same way. Both are listed failure modes in the provider documentation, and both are avoidable by pasting instead of typing.

Connecting the Roo Code OpenAI Compatible provider

Open the Roo Code settings via the gear icon and set API Provider to OpenAI Compatible. Three settings carry the whole connection: Base URL, API Key and Model ID. For this endpoint they are:

API Provider: OpenAI Compatible
Base URL: https://api.lyceum.technology/openai/v1
API Key: lk_your_api_key_here
Model ID: moonshotai/kimi-k2.7-code

The same base URL and key can serve the supported chat models in these profiles. Each profile also needs the selected model’s context, output and image settings. Test tools, streaming and reasoning handling after a model change; a shared request format does not make every capability identical.

Billing follows the same shape. Requests are metered per token, per model, so the cost difference between modes comes from the prices of the models you assign, not from any per-seat line item. You can read the input and output prices into Roo's Model Configuration fields as well, which keeps the cost visible in the extension itself.

Setting limits Roo cannot discover

For this custom provider setup, set the model limits explicitly under Model Configuration. A compatible model-list endpoint does not guarantee that Roo will discover complete context, output or image metadata. Read those values from the selected model’s current record.

  • Max Output Tokens: the cap on a single response. Leave headroom for long edits, but not so much that a looping mode burns tokens unattended.
  • Context Window: the model's real context size. Roo uses this to decide what fits into a prompt.
  • Image Support: enable only for a model that accepts images. A text-only model with this flag on produces broken requests.
  • Input Price and Output Price: optional, but they make per-task cost visible in the extension.

For moonshotai/kimi-k2.7-code the context window is 256K tokens. For every other model, read the value from docs.lyceum.technology on the day you configure the profile. The catalogue changes, and a context window copied from an older record is a wrong number.

Wrong values fail quietly rather than loudly. A context window set too small makes Roo truncate workspace context you assumed was in the prompt; a value set too large lets a request grow until the endpoint rejects it. We walk through the full failure set in what breaks when you switch models on an OpenAI-compatible API. The short version: these two numbers are the contract between Roo's prompt builder and the model, and Roo has no way to check them for you.

Giving each mode its own profile and model

Roo Code's API configuration profiles let you save named sets of provider settings and switch between them, and each profile can carry its own model, key and limits. Create one profile per mode, then link each profile to its mode in the Prompts tab. Roo also remembers which profile you last used with each mode, and a task keeps the profile it started with for its lifetime, so a model switch never lands mid-task.

A starting-point mapping, not a benchmark result. All three models below are hosted in eu-north1, in the European Union, so every mode's traffic stays on EU-hosted models.

ModeModelAPI stringPrice per 1M tokensHosting
Code, DebugKimi-K2.7-Codemoonshotai/kimi-k2.7-code$1.25 input, $0.31 cached input, $4.50 outputeu-north1 (EU-hosted)
Architect, OrchestratorGLM-5.2z-ai/glm-5.2$1.50 input, $0.38 cached input, $4.50 outputeu-north1 (EU-hosted)
Ask (quick turns)DeepSeek-V4-Flash-0731deepseek/deepseek-v4-flash-0731per token; see the model recordeu-north1 (EU-hosted)

Treat this mapping as a starting point rather than a benchmark result. Measure accepted task results and total cost, including fresh input, cached input, output, reasoning and retries. Neither input nor output dominates every agent loop. Check the model record and your measured traffic before setting a budget.

The profile rate limit sets a minimum number of seconds between requests, with 0 disabling it. This can slow request volume and help avoid provider rate limits, but it does not cap total tokens or money spent. Combine request spacing with task iteration limits and a separately enforced spend budget where available. Other profiles follow their own rate limits.

Confirming tool calling works

This is the step teams skip and regret. Roo Code uses native OpenAI-style tool calling exclusively, with no XML-based fallback, so a model that does not support native tool calling cannot drive Roo at all. Before you hand the extension a real task, prove the endpoint and the model do it. With LYCEUM_BASE_URL set to the base URL from your profile, send one request carrying a single tool definition:

export LYCEUM_BASE_URL="https://api.lyceum.technology/openai/v1"
export LYCEUM_API_KEY="lk_your_api_key_here"

curl "$LYCEUM_BASE_URL/chat/completions" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $LYCEUM_API_KEY" \
  -d '{
    "model": "moonshotai/kimi-k2.7-code",
    "messages": [
      {"role": "user", "content": "Read the file src/main.py"}
    ],
    "tools": [
      {
        "type": "function",
        "function": {
          "name": "read_file",
          "description": "Read a file from the workspace with line numbers.",
          "parameters": {
            "type": "object",
            "properties": {
              "path": {"type": "string", "description": "Relative file path"}
            },
            "required": ["path"]
          }
        }
      }
    ]
  }'

The following is an illustrative tool-call response shape, not output captured from a live test. The assistant message contains a tool_calls array, with a function name and JSON-encoded arguments:

{
  "choices": [
    {
      "message": {
        "role": "assistant",
        "tool_calls": [
          {
            "id": "call_01",
            "type": "function",
            "function": {
              "name": "read_file",
              "arguments": "{\"path\": \"src/main.py\"}"
            }
          }
        ]
      }
    }
  ]
}

Inspect the actual message.tool_calls array and validate the function name and JSON arguments. A prose answer to an optional tool request is not proof that the model lacks tool calling. Forced tool selection is not reliably enforced by every model, so compare behaviour with the current model documentation and test the complete Roo workflow.

The setup in this guide runs on Serverless Inference: pre-hosted open-weight models on a shared endpoint, metered per token, with no GPU to provision and no minimum commitment. That is what makes the per-mode arrangement cheap to operate. Running five modes on five models is a string change per profile, not five deployments, and adding a sixth profile for a custom mode costs nothing at the infrastructure layer.

  • This OpenAI-compatible surface documents chat completions and embeddings; use the product-specific API for other media
  • Billing is per token with no minimum commitment
  • Lyceum states that inference prompts and outputs are not retained after processing or used for training; confirm contractual needs in the data processing agreement
  • Serverless inference has no SLA, uptime target, availability tier or service credit; see status.lyceum.technology for operational history

The selected models are listed as EU-hosted in the company catalogue, checked per model. Confirm your data-processing requirements separately. Create a key at dashboard.lyceum.technology, run the tool-call check, and test a small task in each profile. Claude Code uses the separate Anthropic-compatible route at https://api.lyceum.technology/anthropic rather than this Roo endpoint.