Opik's MCP server
One command connects your coding assistant to your traces — and teaches it how to create them in the first place:
What this unlocks
Things you can ask for and get in one turn, without leaving your editor:
Your assistant adds tracing in the right places for your framework, runs the app, and confirms the traces arrived.
It reads the failing traces and their scores directly, instead of you pasting screenshots into chat.
From traces you already have, so the cases are real ones your app hit.
Every later change can be checked against real traces as you make it.
Quick setup with the Opik CLI
The CLI detects your AI client (Claude Code, Cursor, VS Code Copilot, Codex, opencode), picks the right server for your Opik deployment, configures it, and then checks that the configuration it just wrote actually works.
Prefer not to use the CLI? You can wire up any client by hand — skip to Manual setup.
Asking your coding agent to do this works. Setup writes into your AI client’s own configuration, so it never happens on a run that did not ask for it — but naming what you want is asking, and that works without a terminal:
A run that names nothing and has no terminal — a CI job, a Docker build — writes nothing, which is the behaviour you want there.
Install the Opik CLI
The CLI ships with the opik Python package. --upgrade installs the latest
release, which is what you want — these commands gain clients and flags often:
Configure the MCP server
This reuses your existing Opik configuration (~/.opik.config); if you
haven’t configured Opik yet, the wizard offers to do it for you first.
You’ll choose your AI client from a list, then confirm the MCP server and the Opik skill pack for it.
Restart your AI client
Assistants read their configuration at startup, so start a new session before trying the prompts in Start using it.
If your client isn’t detected, see Manual setup.
Check your setup
Each AI client keeps its own copy of the MCP configuration, which isn’t updated automatically when your Opik configuration changes. To see what every detected client points at — and whether it still matches your current Opik configuration — run:
It prints your active Opik configuration, then each AI client that has the Opik MCP server configured: the config file it lives in, the server it reports to (hosted or local), its workspace, and whether it has drifted from your Opik configuration.
A client that has drifted is flagged ✗ OUT OF SYNC — re-run
opik mcp configure to fix it.
A client keeps its MCP connection for the lifetime of its process. After changing
your Opik configuration or re-running opik mcp configure, restart your AI
client so it reconnects with the updated settings.
To view just your active Opik configuration (file path, environment, workspace):
To refresh the skill pack, re-run opik mcp configure — it rewrites the pack
from the latest published version. Assistants read their skills at session start,
so start a new session afterwards.
Start using it
Paste any of these into your assistant. Start with the first — it exercises the whole loop, so if it works, everything is wired up.
From then on your assistant can check its own work against real traces every time you change something.
The tools you’ll have
Your assistant gets six tools. It picks between them on its own — this is here so you know what it can reach for:
To see a payload shape yourself, ask “show me the schema for trace.create” — or read the full list.
Opik Cloud and self-hosted deployments
opik mcp configure works the same whether you’re on Opik Cloud, self-hosted,
or a local install — it sets up the right server for your deployment
automatically.
Opik Cloud (hosted server)
On Opik Cloud, the CLI registers the hosted MCP server over HTTP. Your AI client signs in with a browser-based OAuth flow on first connect, so:
- No API key is stored in the client’s config — you authenticate through OAuth in the browser.
uvis not required — there is no local process to run.- Your workspace is selected during the OAuth sign-in, so a hosted server shows
no workspace in
opik mcp status.
Self-hosted and local (local server)
If no hosted server is available for your environment, the CLI sets up the
local server, which runs on demand via uvx opik-mcp. This requires
uv; if it isn’t on your PATH the CLI stops and
prints the exact command to install it for your platform.
Workspaces
For the local server your workspace is written into the client’s config, so it has
to be the right one. If your Opik configuration doesn’t name a workspace and your
account has more than one, opik mcp configure refuses to continue rather
than falling back to your account default:
Guessing here is the one failure this CLI can produce that doesn’t look like a
failure: your agent would read real traces from the wrong workspace and report
them confidently. Run opik configure, pick a workspace, and re-run.
Manual setup
Prefer to wire it up yourself, or your client wasn’t detected? Configure any client by hand below.
For the skill pack on a client the CLI doesn’t cover, the community
skills CLI knows the skill directories
for 76+ agents (needs Node.js):
There are two servers you can add by hand. opik mcp configure
picks the right one for you, but you can also add either directly in your AI
client’s MCP settings:
- Hosted server (HTTP + OAuth) — available on Opik Cloud and any deployment that provides it. No API key is stored; your client signs in through the browser.
- Local server (
uvx opik-mcp, stdio) — runs on your machine with your credentials in the client’senvblock.
Hosted server (Opik Cloud)
The hosted server connects over HTTP and signs in with a browser-based OAuth
flow on first connect — no API key is stored in the client config. Point your client
at your deployment’s MCP endpoint, which is your Opik API base plus /v1/mcp. On
Opik Cloud that is https://www.comet.com/opik/api/v1/mcp.
Claude Code
Cursor
VS Code Copilot
Add the server with one command:
Or edit ~/.claude.json directly:
Restart Claude Code and complete the browser sign-in when prompted, then ask in the chat: “list my Opik projects”.
Local server (uvx)
The local server runs on demand via uvx opik-mcp (requires
uv), with your credentials passed through the
client’s env block.
opik-mcp is now a Python package. If you previously ran the npx-based
JavaScript server, use the uvx opik-mcp commands below in place of
npx -y opik-mcp.
OPIK_WORKSPACE is optional — you can omit the OPIK_WORKSPACE line/key
entirely and the server uses the default workspace (correct for local/OSS
installs). The snippets below include it for completeness; set it only if you
connect to a named cloud workspace.
Claude Code
Cursor
VS Code Copilot
Codex
opencode
MCP Inspector
Add the server with one command:
Or edit ~/.claude.json directly:
Restart Claude Code, verify with /mcp (opik-mcp should appear as
connected), and then ask in the chat: “list my Opik projects”.
Self-hosted Opik. Add COMET_URL_OVERRIDE to the env block (and OPIK_URL
if Opik lives at a non-default path). ask_ollie and run_experiment are
available on Comet Cloud only — on self-hosted those calls fail at dispatch;
use read / list / write directly.
Ollie & auto-approve
By default, writes that Ollie performs mid-stream (scores, comments, prompt
versions, test-suite items) execute without a per-action confirmation step.
Each auto-approved write is logged as a JSON audit row on the opik_mcp.audit
Python logger.
To require manual confirmation instead, set OPIK_MCP_AUTO_APPROVE=disabled in
the server’s env block. Ollie’s confirmation requests then surface as typed
errors that you can re-issue manually.
ask_ollie and run_experiment are available on Comet Cloud only — on
self-hosted those calls fail at dispatch; use read / list / write
directly.
Known client limits
- Cursor enforces a 60-second hard tool-call timeout that does not reset on
progress notifications. Long
ask_ollieturns will fail on Cursor. For long-running investigations, use Claude Code or VS Code Copilot.
Example conversation
A typical investigative loop using Claude Code:
You: Why did the experiment “gpt-4o-rerank-v3” regress on factuality?
Claude: (calls
ask_ollie) Three traces failed because the reranker dropped the system message. The remaining 12 traces scored above 0.8…You: Score the bottom 3 traces 0.2 with reason “dropped system message”.
Claude: (calls
writewithscore.create×3) Done — three scores recorded on traces<id-1>,<id-2>,<id-3>.