> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://www.comet.com/docs/opik/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://www.comet.com/docs/opik/_mcp/server.

# Changelog

## September 28, 2026

## Bug Fixes & Improvements

* **Large threads no longer fail to load** — `threads/retrieve` had been intermittently returning a 500 since a ClickHouse join between traces and their aggregated spans could pick either side to build the join from; on a large thread it occasionally picked the multi-gigabyte side and ran out of memory. The join now always builds from the smaller, pre-aggregated side, so large threads load reliably.

* **Diagnostics checks your credit balance before starting a run** — Starting a diagnostic run when a workspace was out of credits used to fail only after the run had already started. The run button now checks affordability upfront, and the page shows a clear out-of-credits state instead of a run that fails partway through.

* **Image and other media in experiment output now renders in the comparison views** — The experiment comparison sidebar and the test-suite sidebar only extracted media (e.g. images) from a dataset item's input; a trace's output went straight into the raw JSON view with no media extraction, so an image an SDK logged as output never showed. Output media now renders the same way input media does, including images served from URLs with no file extension.

* **Online evaluation rule duration filters now use the right unit** — A duration filter typed as "> 5" on a rule was compared against the underlying duration column in milliseconds, so on an LLM-as-judge rule it matched effectively every trace instead of only the slow ones. The rule dialog now converts the value the same way the traces table already does, and the duration column is labeled "Duration (s)" so the unit is clear.

* **Bedrock and Mistral streaming integrations no longer lose parts of the response** — OpenAI models called through Bedrock's `invoke_model` (gpt-oss, GPT-5.x, GPT-6) stream text and stop-reason in a shape the aggregator didn't recognize, so those spans ended up with empty output and zero token usage. Bedrock's `converse_stream` sent a tool call's arguments as JSON fragments that got overwritten instead of merged, so a streamed tool call kept only its last fragment, and parallel tool calls collapsed into one. Mistral's reasoning models (Magistral) stream content as a list of thinking/text chunks rather than a string, which crashed the aggregator and left the whole span without output. All three now aggregate correctly.

* **Judge calls to Bedrock-hosted GPT-5.x and GPT-6 models no longer fail** — Opik already dropped the `temperature` parameter for GPT-5-family judge models, since these models reject it, but only recognized them by their plain `gpt-5*` names. The same models routed through Bedrock (e.g. `bedrock/converse/us.openai.gpt-6-...`) keep a `bedrock/` prefix, weren't recognized, and had their judge calls rejected. They're now matched and handled the same way.

* **Hallucination and SycEval judge reasons render as text, not a Python list** — Both metrics ask the judge for a list of reason strings, but the parser rendered that list with Python's `str()`, so the explanation shown to you (and uploaded with the score) was a literal `['reason 1', 'reason 2']` instead of readable prose. It's now joined into text the same way other list-based judge reasons already are.

* **Context-precision and context-recall judges get the same prompt-injection hardening as last week's hallucination and G-Eval fix** — Content being evaluated by these metrics is now clearly namespaced from the rest of the judge prompt, so an input, output, or context that happens to contain something resembling an instruction or a closing delimiter can no longer influence the verdict.

* **Several built-in metrics are more robust on edge-case input** — BLEU now rejects a non-positive or non-integer `n_grams` instead of silently scoring 0 or raising a raw `KeyError`; METEOR tokenizes its input before handing it to NLTK instead of raising `TypeError` on every call; ChrF now scores a candidate against each reference separately and keeps the best match instead of passing the whole reference list into NLTK's single-reference scorer (which scored a perfect match as low as 0.26); and VADER sentiment now downloads its lexicon on first use instead of failing with a bare `LookupError` on a fresh install.

* **`@track` now logs `slots=True` dataclasses correctly** — Encoding a dataclass defined with `slots=True` read from `__dict__`, which such classes don't have, so the logged value was either the object's repr string or an empty object instead of its actual fields.

* **Config files with `%` in a value no longer break `opik configure`** — `~/.opik.config` is read and written with Python's `ConfigParser`, which by default treats `%` as the start of an interpolation sequence; a value containing `%` (for example, in a URL) could fail to load or save. Interpolation is now disabled for this file.

* **Custom scorers now get a real name in results instead of crashing** — Passing a `functools.partial`-wrapped function or a callable class instance as a scorer raised `AttributeError` because neither has a `__name__`. Both are now named sensibly (the wrapped function's name, or the class name as a fallback).

* **Gemini usage without `candidates_token_count` is accepted** — A Gemini response that omits this field (for example, at content-filter length) is no longer rejected while building token usage; the count is treated as unknown rather than causing an error.

* **19 more providers are priced correctly** — Hyperbolic, Baseten, Lambda AI, nscale, OCI, Replicate, watsonx, Cohere, Novita, Cloudflare, Anyscale, Scaleway, OVHcloud, GMI, Gradient AI, Libertai, Azure AI, Vercel AI Gateway, and OpenRouter are now registered as canonical providers, so calls through them resolve a model price instead of going unpriced.

* **Large dataset batch uploads no longer exceed the configured size cap** — `StreamingBatchWriter` decided when to flush a batch based only on the serialized items' size, without counting the surrounding envelope (dataset name, project name, batch id, and JSON wrapper). Every flushed request body was slightly larger than the configured cap; the cap now accounts for the full request body.

* **`video_url` placeholders in chat prompts are validated** — Placeholder validation (`validate_placeholders=True`) checked `{{...}}` templates inside `text` and `image_url` prompt parts but skipped `video_url`, even though it's rendered the same way, so a templated video URL could silently fail to substitute.

* **Overlapping evaluations no longer leave your app's HTTP connections short-lived** — `evaluate()` temporarily patches httpcore's keep-alive behavior for the duration of a run and restores it afterwards. When two evaluations ran at the same time, the first one to finish could restore the patch early (disabling it for the run still in progress) or the last one to finish could restore the wrong thing, leaving the host application's own connections short-lived after both runs completed. The patch is now reference-counted so it's only removed once every overlapping run is done.

## Performance Improvements

* **Reading large experiments page by page is up to \~2.6x faster** — The query behind experiment item paging was re-evaluated up to six times per page because it's referenced from multiple places in the query plan; it's now evaluated once per page and the result reused. Measured on a 100k-item experiment at a page size of 2000: a 50-page read dropped from 14.48s to 5.56s.
* **The experiment compare view no longer re-scans the whole experiment on every page** — Each page of the compare view re-resolved the same target-project lookup by reading the experiment's entire trace set, even though the answer is identical for every page of that read. The lookup is now cached per read, so its cost no longer scales with the size of the experiment.
* **Trace and thread reads use less CPU and read less data** — `find_trace_stream` and `find_thread_by_id` deduplicated spans with a ClickHouse `FINAL` read, which is expensive. Both now use a pre-aggregated form instead, cutting CPU by 22–50% and roughly halving bytes read in production measurements, with identical results.
* **Span-by-ID reads scan less data** — Reads that look up a span by ID now bound themselves to the week(s) that ID's timestamp resolves to, instead of scanning across all partitions.

---

And much more! 👉 [See full commit log on GitHub](https://github.com/comet-ml/opik/compare/2.2.78...2.2.82)

*Releases*: `2.2.79`, `2.2.80`, `2.2.81`, `2.2.82`

## September 21, 2026

## Bug Fixes & Improvements

* **`opik import` can import traces of any age again** — The trace importer derived each imported trace's ID from the exported trace's `start_time`, which bakes that original timestamp into the ID's UUIDv7 prefix. Opik validates that an ingested ID's embedded timestamp falls inside an ingestion window around now — 24 hours by default on Opik Cloud, which is what keeps its storage layer partitioned correctly — so importing anything older failed with `Invalid UUID for id ... reason 'too_old'`. Imported traces now get IDs minted at import time, which also keeps IDs unique across projects and workspaces. `start_time` and `end_time` still carry the original values, spans and feedback scores are remapped onto the new IDs, and each trace records the ID it came from in its metadata as `_import_id`. Traces are imported in the order the source project listed them, so the destination trace list and thread view keep that order; note that the time-bucketed project metrics key off the ID, so imported traces count at import time. If you hit the error above, upgrade the SDK. See [Imported trace and span IDs](/tracing/advanced/export-data#imported-trace-and-span-ids).

## September 14, 2026

## Map Trace and Span Fields onto Dataset Item Columns

Adding traces or spans to a dataset always used a fixed shape — `input` from the trace's input, `expected_output` from its output, nothing else. Getting any other field into the dataset meant exporting and reshaping the data by hand.

The "Add to dataset" dialog now has an **Advanced mapping** toggle. With it on, you can re-point `input` and `expected_output` at a different field, add extra columns from a path explorer over the sampled traces/spans, and use **Quick add** chips to pull in common enrichment fields (feedback scores, metadata, tags) or columns the target dataset already has. A coverage badge on each mapped field and a preview table show what will actually land in the dataset before you commit.

👉 [Adding traces to a dataset](/evaluation/advanced/manage_datasets#adding-traces-to-a-dataset)

## Name Experiments from the Playground

Playground experiments were always given a random `adjective_noun_number` name, with no way to set one — making it harder to find a specific run later in the Experiments list.

The Playground's output toolbar now has an optional experiment name field. Each output column derives its own name from it by appending a letter (`concise` → `concise_a`, `concise_b`, ...), and a preview shows what a run will produce before you start it. The field is optional: leave it empty and Opik generates a name exactly as before, so nothing changes if you don't use it. A completion toast reports the experiments created and links straight to the comparison view.

## Bug Fixes & Improvements

* **Exporting experiment results from the browser now covers the full result set** — The comparison table's export button only ever covered the rows selected on screen, and was disabled with nothing selected. With no selection, it now exports the whole result set behind the current filters (capped at 2000 rows, since a browser export holds everything in memory); past that cap, it points you to the [SDK export script](/evaluation/advanced/export_experiment_results), which has no such limit.
* **Claude models no longer 400 when routed through Bedrock, OpenRouter, or a custom provider** — Claude rejects a request that sets both `temperature` and `top_p`. Opik already enforced that, but only for the Anthropic provider directly; the same Claude model called through Bedrock, OpenRouter, or an OpenAI-compatible custom provider sent both and was rejected. The either/or choice now applies to every provider on the backend, and the Playground's model panels all show the same either/or control for any Claude model instead of only some of them.
* **Copy actions are now one click away on trace, thread, and span panels** — Copy-ID and copy-link used to live inside an overflow menu; they now render as icon buttons right next to the panel title, with the overflow menu reduced to Export and Delete.
* **A setup hint now appears on expanded trace errors** — Expanding an error in a trace now offers to connect an AI coding assistant (Claude Code, Cursor, VS Code, and others) via MCP, with install instructions and a ready-to-use prompt for investigating that specific error.
* **`opik configure` and `opik mcp configure` now ask the same questions the same way** — The two commands used to ask about registering an MCP server differently, and a bare Enter could silently decline or silently accept depending on which one you ran. Both now show the same explanation and client picker, which gained an "All" option and a manual-setup fallback (with a docs link) for AI clients Opik doesn't auto-detect.
* **Threshold alerts no longer silently stop firing** — An alert whose threshold trigger was missing a `window` (including some created before this fix) failed internally on every evaluation and was skipped without notifying anyone. Alerts in this state now evaluate using their configured or default window instead of failing silently, and creating or editing one with an invalid threshold or window is now rejected up front with a clear error.
* **Bulk-deleting alerts now requires the same permission as editing them** — `deleteAlertBatch` was missing the permission check its create/update siblings already had.
* **Webhook destinations are now checked before Opik connects to them** — Webhook deliveries and the "test connection" action now validate the destination URL first; on Comet-managed workspaces this blocks destinations that only Opik's own network can reach, and a failed test no longer echoes the destination's response body back to the caller.
* **LLM-as-judge evaluations (hallucination, G-Eval) can no longer have their verdict overridden by the content being judged** — A judge prompt whose evaluated input, context, or output happened to contain something that looked like a closing delimiter could break out of its section and forge a verdict. The judged content is now clearly namespaced from the rest of the prompt.
* **G-Eval scores are now accurate for two-digit scores and more tokenizers** — The logprob-based scorer assumed a score is always a single digit at a fixed token position. A two-digit score (e.g. 10) or a tokenizer that folds whitespace into the score token produced a badly wrong score instead of the intended one; both are now located and parsed correctly.
* **LangChain streaming integrations no longer drop token usage** — A streaming run with an empty `generations` list (seen on the Anthropic-Vertex AI path) raised internally and discarded token usage for that call instead of reporting `None`.
* **Playground dataset runs are now scored by the rules you actually selected** — An online evaluation rule scoped to "Production traces" was scoring (or failing to score) Playground dataset runs regardless of the rules picked in the Playground's own metric selector. Playground dataset runs are now scored only by the selected rules plus any rule explicitly scoped to experiments, and a Playground run with no dataset is no longer treated as production traffic.

## Performance Improvements

* **Dataset and experiment transfers now default to 8 threads** — The Python SDK's bulk paths shipped with inconsistent worker counts: dataset reads (`get_items()`, `stream_items()`) and `Dataset.insert()` ran on 4 threads, while `Experiment.batch_upload_items()` uploaded batches sequentially unless you passed `num_threads` yourself. All three now default to **8**, the setting we benchmark against — the experiment upload path benefits most, since it went from sequential to parallel. Every call still takes `num_threads` if you want to push harder or ease off; see [Tuning SDK throughput](/evaluation/advanced/manage_datasets#tuning-sdk-throughput). Note that raising it much further does not help: on a 119,903-item upload, 16 threads measured slightly slower than 8, because the SDK saturates a CPU core serializing and compressing payloads before thread count becomes the limit. Parallel dataset upload still requires an Opik backend of 2.2.8 or newer, and falls back to a sequential upload against anything older.
* **Dataset and experiment uploads encode payloads up to 5x faster** — The Python SDK now uses `orjson` to encode request bodies and compute content-deduplication hashes where it's available, instead of the standard library's `json` module. On a 30,000-item upload over 8 threads, end-to-end time dropped from 148.7s to 78.2s. This is automatic and requires no changes to your code; installs without an available `orjson` wheel fall back to the standard library unchanged.

---

And much more! 👉 [See full commit log on GitHub](https://github.com/comet-ml/opik/compare/2.2.52...2.2.71)

*Releases*: `2.2.53`, `2.2.54`, `2.2.55`, `2.2.56`, `2.2.57`, `2.2.58`, `2.2.59`, `2.2.60`, `2.2.61`, `2.2.62`, `2.2.63`, `2.2.64`, `2.2.65`, `2.2.66`, `2.2.67`, `2.2.68`, `2.2.69`, `2.2.70`, `2.2.71`

## September 7, 2026

## Custom AI Providers Can Use OAuth2 Token Auth

Custom AI provider integrations only accepted a single static API key, so any provider that rotates or expires credentials — an enterprise gateway sitting behind OAuth2, for example — had to be re-configured by hand every time a key expired.

The provider configuration dialog now has an **Authentication** mode switch: choose **Static API key**, or **OAuth2 client credentials** and enter the token URL plus the `client_id` and `client_secret` rows (add more rows, such as `scope` or `audience`, if your auth service needs them). This applies to custom providers and to Bedrock.

Opik fetches the access token, attaches it as a `Bearer` token to every model call, and refreshes it before it expires. A **Test connection** button runs the token fetch on the backend and reports the token lifetime before you save. Credential values with secret-like names lock automatically: Opik encrypts them at rest and never reads them back.

👉 [OAuth2 Client Credentials Authentication](/administration/workspace-settings/ai_providers#oauth2-client-credentials-authentication)

## Fine-Grained Thinking Controls for Gemini Models

Gemini models increasingly default to an internal "thinking" step before answering, which adds cost and latency that isn't always worth paying — Opik previously had no way to influence that.

The Playground, Online Evaluation rules, and experiments can now set a thinking level per Gemini call, for both Vertex AI and Google AI Studio: Auto, None, Minimal, Low, Medium, or High (availability depends on the model — Gemini 2.5 Pro, for example, can't disable thinking). Vertex AI translates the level into the `thinking_budget` value it expects; Google AI Studio sends the level directly. A follow-up fix added the dedicated None level after Flash Lite models — which don't think by default — were found to have thinking silently re-enabled by a preselected Minimal level, adding several seconds of latency to every call.

## Bug Fixes & Improvements

* **Judge scoring no longer 500s or stalls on edge-case content** — A judge message whose content legitimately started with `[` (for example, a prompt beginning "\[Source Text]...") could return a 500 for an entire project's list of evaluation rules and silently stop sampling for that project. A trace whose mapped input, output, or metadata was a bare JSON value instead of an object failed its whole evaluation instead of just skipping that field, and a Vertex AI evaluation using the tool-calling path failed outright because Vertex rejects a forced tool choice. All three now evaluate instead of failing.

* **Online-scoring queues no longer get stuck behind one bad message** — A single oversized or undecodable message could wedge an entire scoring stream, blocking every other trace behind it (one recorded case reached tens of gigabytes of stuck messages), and a permanently failing provider error (like an invalid-credentials response) was retried repeatedly instead of being retired immediately. Both cases are now dropped or retired instead of stalling the queue.

* **Online evaluation sampling now applies only to production traces** — The sampling rate is meant to thin a continuous production stream, but it was also being applied to traces from experiments, playground runs, and optimizations — one-off runs a user starts intentionally. Those are now always scored in full, regardless of the rule's sampling rate.

* **Prompt version history now loads and labels every version correctly** — The Prompt tab's version history only ever loaded the first 25 versions, so older versions were unreachable and a deep link to one silently fell back to showing the latest version under a stale label. Version history, the Compare and Deploy menus, and version labels throughout the app (Playground, trace details, Optimizations, Agent Runner) now paginate through the full history and use each version's real, persistent number.

* **Experiment items export no longer drops rows or loses sort order** — Exporting experiment items truncated long field values and could ignore the sorting and search filters applied in the table. Exports now include full field values and respect the same sort and search as the on-screen view.

* **Dataset item and version data is more accurate** — A dataset version's reported item count came from the raw number of items submitted, not the number actually stored, so a batch with duplicate item ids showed an inflated count. Separately, filters on a dataset item listing weren't always applied consistently between the page and count queries, and page reads weren't bounded to the requested page size. All three are now correct.

* **Annotation queue names can no longer cause lost form data** — A whitespace-only queue name passed client-side validation, was rejected by the API, and closed the create/edit dialog anyway — discarding everything entered in the form. Names are now trimmed and validated before submission, and the dialog stays open until the save actually succeeds.

* **Dashboards can now be filtered by description** — Filtering the dashboards list by its description field returned "No matching results" for every operator, because the field wasn't registered as filterable on the backend. It's now supported like any other field, and a failed list request now shows an error instead of silently rendering as empty.

* **Optimization runs list now shows the actual best trial** — The runs list reported the baseline trial's latency and cost as if it were the run's best result, with the cost/latency delta always showing 0%, while the run's own detail page correctly showed the genuine best trial. The list now matches the detail page.

* **Cost tracking recognizes more providers and model naming schemes** — Calls through Cerebras, Snowflake Cortex, and DeepInfra weren't recognized as canonical providers, models identified by OpenTelemetry semantic-convention provider names weren't mapped to Opik's provider list, and models with compact `YYYYMMDD` date suffixes in their id weren't matched to a price. All of these now price and attribute correctly.

* **Streaming SDK integrations no longer swallow errors** — The Anthropic, Bedrock, and Mistral stream wrappers' cleanup step could silently swallow an exception raised while finalizing a stream instead of surfacing it, and the LangChain integration stopped extracting token usage whenever a call's model metadata was absent. Both are fixed, and the Bedrock integration now correctly passes through Claude's cache-read/write token counts instead of undercounting them.

* **Cursor extension reports accurate usage and structure** — The Opik extension for Cursor was logging zero-token spans after Cursor stopped populating the field it read from, one flat span per conversation turn instead of a span per model or tool call, and undercounted prompt and total tokens (by up to \~6x) whenever cache reads dominated a call. All three are fixed.

* **MCP OAuth consent screen preselects your default workspace** — Approving an AI tool's access via MCP OAuth used to preselect whichever workspace the backend returned first; it now preselects your actual default workspace.

* **`opik configure` can set up MCP and skill packs in one step** — Running `opik configure --install-mcp --install-skills` registers the Opik MCP server with detected AI coding tools (Claude Code, Cursor, VS Code, Codex, opencode) and installs the Opik skill pack that teaches the coding agent how to instrument your code, instead of setting each one up by hand.

* **Dataset inserts can skip deduplication** — `Dataset.insert()` (and the batch, pandas, and JSONL variants) now accepts a `deduplication` argument; setting it to `False` skips downloading and hash-comparing existing items before insert, trading duplicate-safety for speed on large datasets.

* **Python SDK reports anonymous feature usage** — The SDK now reports which integrations and features are used (for example, which LLM provider you're calling through) to help prioritize development; no trace, span, prompt, or dataset content is included. This is on by default and can be disabled by setting `OPIK_ANALYTICS_ENABLE=false`. Running `opik configure` or `opik mcp configure` additionally attaches an account identifier so a setup run can be tied to your workspace.

* **Self-hosted online evaluation rules get the same tool-calling judge as Comet-managed workspaces** — LLM-as-judge rules that reference `{{trace}}`, `{{span}}`, or `{{spans}}` can let the judge model call tools to fetch large trace or span data on demand instead of inlining all of it into the prompt, avoiding context-window overflow. This previously only ran for Comet-managed workspaces; the feature toggle gating it off for self-hosted instances has been removed.

## Performance Improvements

* **Dataset writes and reads are significantly faster** — Inserting into a dataset serialized every batch behind a lock only needed for creating a new version, item-count updates cost three database round-trips instead of one, and the enrichment queries used when loading dataset details ran one after another instead of concurrently. A redundant item-count scan also ran even when a version already tracked its own total, and a per-dataset experiment summary query scanned the whole workspace's experiment items instead of just the requested dataset's. All of these are fixed, substantially reducing dataset save and read latency.

* **Reading large datasets from the Python SDK is up to \~3x faster** — A new parallel, chunked `Dataset.stream_items()` reader (which `get_items()` now uses internally) measured 3.3x faster than the previous single-threaded read path at 8 threads, with no change to its output.

* **Traces, trace logs, and experiment results tables render faster on large projects** — These tables now virtualize their rows and columns, so only what's visible on screen is rendered, instead of paying render cost for every column and row up front.

---

And much more! 👉 [See full commit log on GitHub](https://github.com/comet-ml/opik/compare/2.2.36...2.2.52)

*Releases*: `2.2.37`, `2.2.38`, `2.2.39`, `2.2.40`, `2.2.41`, `2.2.42`, `2.2.43`, `2.2.44`, `2.2.45`, `2.2.46`, `2.2.47`, `2.2.48`, `2.2.49`, `2.2.50`, `2.2.51`, `2.2.52`

## August 31, 2026

## MCP + Skill Pack Setup Can Now Run Without a Terminal

Setting up the Opik MCP server previously meant sitting through an interactive wizard, and it only reached three of the AI assistants people actually use. `opik configure --install-mcp` and `opik mcp configure --ai-client <host>` now run non-interactively — pass the client explicitly (or `--ai-client all`) and the command completes on its own, so it can run from a coding agent, a Dockerfile, or CI. Codex and opencode join Claude Code, Cursor, and VS Code Copilot as supported hosts, and every install now ends with a real verification call that reports the workspace and project count instead of assuming the config it just wrote actually works.

A new `--install-skills` flag (and the equivalent `--skills` on `opik mcp configure`) installs the companion Opik skill pack alongside the MCP server — instructions the assistant reads to know how to use the tools the server exposes, rather than just their tool list. The two are asked about separately since they carry different trust implications: the MCP server writes credentials into a config file, the skill pack installs instruction files the assistant executes with its own permissions. The onboarding copy-paste prompt and docs now point at `/opik-instrument` rather than the unnamespaced `/instrument`, avoiding collisions with other skills of the same name.

## Bug Fixes & Improvements

* **Self-hosted: agentic tool-calling scoring is on by default** — LLM-as-judge online scoring's agentic tool loop (previously introduced behind a toggle) was already enabled on Comet-hosted workspaces; self-hosted installs inherited the toggle's off default. The toggle has been removed and the behavior is now unconditional everywhere.

* **Exports now match what's on screen** — Exporting experiment items ignored the active sort and search, and truncated long field values regardless of the on-screen truncation setting. Exporting traces or threads with a search term containing leading or trailing whitespace could also return different rows than the table showed. All three now export exactly what's displayed.

* **Optimization runs list reports the same best trial as the run page** — The runs list computed its "best" latency and cost from a column that dataset-based runs (Studio runs and every SDK optimizer) never populate, so it silently fell back to the baseline trial with a flat 0% delta on every row. It also didn't discount a candidate that had only evaluated part of the dataset. Both now match the run page's logic, and the previously mislabeled "Opt. cost" column is named consistently between the two screens.

* **Dashboards can be filtered by description, and a failed list load shows an error instead of "no results"** — Filtering dashboards by description previously returned a 400 that several list views quietly rendered as an empty state rather than surfacing. Both are fixed: the filter now works, and a failed request shows an error message on affected pages.

* **Annotation queue names can no longer be blank, and a rejected save no longer discards your edits** — A name of only whitespace passed client-side validation and was rejected by the API, closing the dialog and losing everything typed. Whitespace-only names are now blocked before submission, and the dialog stays open on any save failure so nothing is lost.

* **Dataset item pages with large payloads no longer fail, and filters return correct results** — Reading a page of a large dataset version could exceed memory limits and return a server error, since fetching a single page still had to sort the entire version in memory. Pages are now resolved in two steps that avoid the full sort, and a follow-up fix ensures filters and sorting apply to the right columns so filtered pages return the correct items instead of a short page.

* **More accurate cost and usage across several providers and integrations** — DeepInfra model prices (including its Claude, DeepSeek, Qwen, and Llama catalogs) now load instead of costing $0, since DeepInfra was missing from the internal provider list. Custom OpenTelemetry instrumentation reporting a provider name in the OTel semantic-convention vocabulary (for example `vertex_ai` or `aws.bedrock`) now maps to Opik's canonical provider names instead of silently costing $0. Claude's cache-read and cache-write token counts now flow through correctly when using Claude on Bedrock with streaming. LangChain usage and cost tracking no longer disappears entirely when model metadata can't be resolved.

* **Streaming SDK integrations no longer silently swallow errors** — The Anthropic, Bedrock, and Mistral integrations patch shared streaming classes process-wide to add tracing. A cleanup step meant to run only for tracked calls used a `return` inside a `finally` block, which discarded any in-flight exception — so once any traced call ran, an *untracked* stream elsewhere in the same process that failed would complete silently instead of raising. For Bedrock this could also return `None` in place of real response data from any `botocore` call, not just Opik's. Tracked and untracked streams now both behave as expected.

* **Cursor extension logs each tool and model call as its own span** — A Cursor agent turn previously logged as a single trace and a single span, so a multi-step turn showed only the initial question and the final answer. Each model call now logs as its own LLM span and each tool call as its own tool span, nested under the turn, with tool names interleaved with assistant messages in the trace output.

## Performance Improvements

* **Faster dataset experiment summaries** — The query behind a dataset's experiment summary scanned every experiment item in the whole workspace before filtering down to the requested dataset. It now prunes to the relevant experiments upfront, cutting rows read by roughly 13x on a workspace with 3M experiment items across 200 experiments, and skipping the scan entirely for a dataset with no experiments.

* **Faster dataset item inserts** — Adding items to a dataset updated the version's item count with a read-modify-write cycle inside the per-dataset lock. It's now a single atomic increment, cutting the database round trips on that path from three to one.

---

And much more! 👉 [See full commit log on GitHub](https://github.com/comet-ml/opik/compare/2.2.36...2.2.45)

*Releases*: `2.2.37`, `2.2.38`, `2.2.39`, `2.2.40`, `2.2.41`, `2.2.42`, `2.2.43`, `2.2.44`, `2.2.45`

## August 24, 2026

## `evaluate_threads` Can Now Judge RAG Agents on Retrieved Context

`evaluate_threads` only accepted `trace_input_transform` and `trace_output_transform`, so a thread was flattened to plain role/content messages with no way to pass along the documents each answer was grounded on — judging relevance or faithfulness for multi-turn RAG agents meant falling back to scoring individual traces instead.

A new `trace_context_transform` argument receives the whole trace object and attaches the extracted context to that trace's message:

```python
from opik.evaluation import evaluate_threads
from opik.evaluation.metrics import ConversationalCoherenceMetric

evaluate_threads(
    project_name="rag_chatbot",
    filter_string='id = "0197ad2a"',
    eval_project_name="rag_chatbot_evaluation",
    metrics=[ConversationalCoherenceMetric()],
    trace_input_transform=lambda x: x["input"],
    trace_output_transform=lambda x: x["output"],
    trace_context_transform=lambda trace: trace.metadata["retrieved_docs"],
)
```

Only metrics that declare `uses_message_context = True` see the context, so every other metric's prompts and scores are unaffected. `ConversationalCoherenceMetric` is the first consumer: when context is present it also judges whether each answer is supported by it, returning an additional `conversational_groundedness_score`, while its original score is unchanged.

## Bug Fixes & Improvements

* **Daily briefings report "Not enough credits" instead of failing silently** — A daily report that ran out of LLM credits mid-run used to just say "Briefing failed," with the pod already allocated and the spend already burned. The trigger now checks credits up front, and a briefing that can't run for that reason shows "Not enough credits" with a direct link to add more, alongside Run again.

* **Provider errors keep their real HTTP status instead of turning into a 500** — Traces and online-scoring rules calling out to an LLM provider reported a generic server error whenever the provider's response couldn't be parsed into its usual error shape — a plain-text upstream outage message, or a rate-limit response wrapped by a retry layer. The provider's actual status is now recovered from these cases too, so a caller-side or provider-side failure is no longer reported as an Opik fault.

* **LLM-as-judge scoring retries on empty provider responses and no longer hangs indefinitely** — An empty structured-output response from the judge model previously fell outside the retry policy and cost the item its whole score. These are now retried, and judge calls carry an explicit timeout and retry budget so a stalled provider can no longer block scoring indefinitely.

* **Silent span-conversion failures in the ADK integration are now reported** — When converting an ADK `LlmResponse` or extracting its usage failed, the span was still logged but silently missing output and usage, with nothing surfacing the failure. This is now logged as an error instead of being swallowed.

* **Fixed two production NullPointerExceptions** — Deleting or batch-updating dataset items, and querying span metrics, could 500 when a filter used an operator not supported by its field, instead of returning the same 400 other endpoints already gave. Separately, a judge response with no text (for example from a content filter) could crash scoring for the whole trace instead of reporting that nothing was scored.

* **Prompts no longer leak across projects with the same name** — A backward-compatibility fallback for looking up a prompt by name, when no project was specified, incorrectly matched a same-named prompt from any project in the workspace instead of only project-less legacy prompts, so a client without a `project_name` set could read, and unintentionally version, another project's prompt.

* **Grouped chart series and tag colors are easier to tell apart** — Two adjacent colors in the series palette were close enough to be confused in a thin chart line; one has been replaced with a more distinct color, so any label that resolved to that palette slot now renders more legibly wherever it appears — chart series, tag chips, and feedback scores.

* **Dataset item sorting by JSON fields now handles special characters correctly** — Sorting by `input.*`, `output.*`, or `metadata.*` keys built the underlying query by interpolating the key text directly, unlike `data.*` sorting, which could corrupt results for keys containing quotes or other special characters. JSON sort keys are now passed as bound parameters like everywhere else.

## Performance Improvements

* **Faster traces table on projects with many feedback score columns** — The traces table builds one column per feedback-score name, and every cell — editable or not — was paying the cost of wiring up inline editing. Non-editable score columns now render a lighter read-only cell, cutting a multi-second main-thread block down substantially on projects with hundreds of score columns and rows.

---

And much more! 👉 [See full commit log on GitHub](https://github.com/comet-ml/opik/compare/2.2.29...2.2.36)

*Releases*: `2.2.30`, `2.2.31`, `2.2.32`, `2.2.33`, `2.2.34`, `2.2.35`, `2.2.36`

## August 17, 2026

## Experiment Traces Are Now a Logs Tab

Experiment traces were previously reachable only through a small "Go to logs" link tucked into the metadata row, easy to miss — users who couldn't find it would conclude nothing had been logged. They're now a dedicated **Logs** tab on the experiment page, always scoped to that experiment or, when comparing, to every experiment in the comparison. Comparing experiments that live in different projects now shows an explicit message instead of silently dropping traces, since logs are read per project.

The tab carries its own toolbar (just the columns selector, matching the other experiment tabs) and its own saved column/sort/filter configuration, so setting it up doesn't overwrite what you've configured in the Playground or trial overlay views. The Playground's "view trace" link on an experiment result now opens this tab directly instead of a dead-end overlay.

## Bug Fixes & Improvements

* **LLM-as-judge online-eval rules score more consistently** — Built-in judge prompt templates and the scoring engine disagreed on whether a judge's answer should be a flat or nested JSON object, so a rule could non-deterministically return an empty score depending on which shape the model happened to pick. The engine now accepts both shapes, correctly parses quoted numbers (`"0.8"` no longer becomes `0.0`) and common boolean spellings (yes/no, pass/fail, on/off), and reports — rather than silently drops — a score name the rule doesn't recognize. Judge prompts are also no longer HTML-escaped before being sent, which had been corrupting values like base64 data URLs assembled from template variables.

* **Unsupported LLM features now return a clear 400 instead of a mysterious 500** — Requesting a capability the selected provider doesn't support (for example, requiring tool choice against a model that can't do it) previously failed with a generic server error and, for online-scoring rules, burned through the full retry budget before the message was dropped. It's now reported immediately as a 400 naming the unsupported feature, for both regular and streaming requests.

* **Fixed a Vertex AI thread leak that could exhaust a backend pod** — Every LLM call routed through Vertex AI built a brand-new client whose background threads were never released, which could accumulate to thousands of threads per pod over time and eventually require a restart. The client is now created and closed per request instead.

* **Optimization cost totals now include GEPA reflection spend** — LLM calls the optimizer makes internally during GEPA reflection belong to no single trial, so they were billed but excluded from a run's total cost, undercounting real spend. That spend is now traced, tagged with the run, and included in `total_optimization_cost` (`opik-optimizer` 3.2.0). Upgrade both packages together — `pip install -U "opik-optimizer>=3.2.0" "gepa>=0.1.0"` — raising `gepa` on its own breaks at runtime against the older reflection-template format that 3.2.0 no longer speaks.

* **Self-hosted: per-component image registry override, and a demo-data job fix** — The Helm chart now lets you set a separate image registry for the backend, frontend, or Python backend individually, falling back to the chart-wide registry when unset. The demo-data job also now honors a component's configured image repository instead of always defaulting to `opik-python-backend`.

## Performance Improvements

* **Faster trace counts on projects using guardrails filters** — The traces-count query that fires on every Logs page load always joined against guardrails results and sorted every row to deduplicate, even when no guardrails filter was applied. Both are now skipped unless actually needed, cutting rows read roughly in half and cutting measured P99 latency by about 90% on production workspaces.

* **Faster trace lookups by thread ID** — Filtering traces or threads by thread ID wasn't engaging an existing prefilter, so a lookup could scan a project's entire span history just to find one thread's traces. The prefilter now engages for thread-scoped lookups, cutting measured production query times from multiple seconds to under 200ms.

---

And much more! 👉 [See full commit log on GitHub](https://github.com/comet-ml/opik/compare/2.2.23...2.2.29)

*Releases*: `2.2.25`, `2.2.26`, `2.2.27`, `2.2.28`, `2.2.29`

## August 10, 2026

## `evaluate()` Surfaces Metric Failures Instead of Dropping Them

Three cases that previously produced no score and no actionable error are now reported: a metric that doesn't declare `**kwargs` no longer fails on every item with a confusing "unexpected keyword argument" error (inputs are narrowed to what the metric's signature actually accepts), an item-level evaluator with an unsupported type now produces a failed score explaining why instead of being silently skipped, and a `scoring_key_mapping` value that matches nothing is now logged as a warning instead of at debug level.

A new `error_tolerance` argument lets you opt into recording pre-scoring failures — a missing required score argument, or an item-level evaluator that can't be built — as failed score results instead of aborting the whole run and discarding every metric that already succeeded:

```python
from opik.evaluation import evaluate
from opik.evaluation.metrics import ErrorTolerance

evaluate(..., error_tolerance=ErrorTolerance.ALL_SCORING_ERRORS)
```

The default behavior is unchanged. Tolerated failures show up on their own span in the trace, and the chosen tolerance is preserved when resuming an interrupted evaluation.

## Bug Fixes & Improvements

* **`Experiment.batch_upload_items` for bulk experiment creation** — A new method uploads experiment items together with their traces, spans, and feedback scores in a single call, instead of creating traces first and linking them afterward. Items are validated up front, batched to respect backend limits, and retried automatically on rate limiting.

* **Stuck optimization runs are detected within minutes instead of hours** — Following up on the stalled-run detection shipped previously, the reaper now derives liveness from actual trial and item progress rather than only the run's last status change, catching a run orphaned by a worker restart in around an hour instead of up to eight. A run that's still legitimately evaluating items is never mistaken for stalled, a wrongly-flagged run recovers automatically once its worker reports back, and it can now be cancelled instead of being stuck.

* **Nebius and SambaNova model costs now load correctly** — Both providers were missing from the internal provider list used to load pricing data, so calls through Nebius or SambaNova models (including their DeepSeek, Qwen, and Llama catalogs) were costed at \$0. Pricing now loads correctly for both.

* **Project lookups no longer leak across workspaces** — Looking up a project by name cached the result under a key that didn't account for the workspace, so switching organizations in the UI without a full page reload could keep resolving a project name to the previous organization's project.

* **Python UDF metric evaluation retries on transient connection drops** — When the Python scoring backend closed its connection before responding, the error wasn't recognized as retriable, so a transient connection drop failed the evaluation outright. These are now retried automatically.

* **More accurate cost and usage on experiments with re-sent spans** — A span re-sent under a different parent span could be counted twice when computing usage and cost rollups for experiments and optimizations. Spans are now deduplicated by their own id, so each span is counted once regardless of how it was re-sent.

## Performance Improvements

* **Faster trace loading for traces with many spans** — Looking up traces by id no longer forces a full deduplication pass (`FINAL`) over the spans table; the query now filters to the relevant trace first and deduplicates only what's left, cutting CPU cost by over 90% on traces with hundreds of spans.

* **Faster project-level trace stats** — The query behind trace statistics on the projects list no longer performs a full blocking sort over every span in a project's history to deduplicate them, so project pages with long-running, high-volume projects load faster.

---

And much more! 👉 [See full commit log on GitHub](https://github.com/comet-ml/opik/compare/2.2.0...2.2.23)

*Releases*: `2.2.19`, `2.2.20`, `2.2.21`, `2.2.22`, `2.2.23`

## July 20, 2026

## Redesigned Optimization Run Overview

The single-run page in Optimization Studio has been rebuilt end to end. The header now shows the run's dataset, model, algorithm, and metric as pills (hover the metric pill for its full config), and the KPI cards show the current value, a trend badge, and a baseline → best line, with the duration now measured to actual run completion rather than the last trial's timestamp. The trials progress chart highlights the best trial by default and labels discarded trials clearly instead of the previous "pruned" terminology.

A new **Best trial prompt** panel shows an inline diff of the winning prompt against the baseline, and trial details now open in a right-side panel (Results / Prompt tabs) instead of navigating to a separate page.

## Stuck Optimization Runs Now Recover, and Failures Are Explained

Previously, an optimization run whose worker crashed or lost contact could sit in "Initializing" or "Running" forever with no indication anything was wrong. A background job now detects runs stuck for too long (5 minutes with no worker pickup, or 8 hours still running) and marks them as failed with a system-generated reason written to the run's logs.

Separately, the worker now classifies real failures — auth errors, rate limits, misconfigured datasets/metrics, out-of-memory kills, and more — into a specific, readable message instead of leaving the run to fail silently. Runs that complete but produce no usable scores now show an explicit "No usable scores" warning, including how many items failed to score when that information is available.

## Bug Fixes & Improvements

* **Workspace switcher now matches the projects menu** — The workspace dropdown now groups workspaces into Pinned and Recently visited sections with per-row pin/unpin controls, instead of a single flat list, matching the pattern already used for the projects menu.

* **Mobile web onboarding** — Opening Opik on a phone now shows a short guided walkthrough (what a trace is, how issues get flagged, how to connect your first integration) instead of the desktop quickstart, with an option to email yourself setup instructions.

* **Workspace home now always opens on Projects** — Navigating to a workspace's home URL now consistently lands on the Projects tab, instead of sometimes reopening whichever project you last viewed.

* **Metadata filter dropdown no longer overlaps rows** — Long metadata keys that wrapped onto multiple lines in the filter panel's autocomplete dropdown could visually overlap the option below them; rows now grow to fit the wrapped text.

* **GPT-5.6 models now behave as reasoning models** — `gpt-5.6-luna`, `gpt-5.6-sol`, and `gpt-5.6-terra` were missing from the reasoning-model list, so the Playground and online scoring still tried to send `temperature`, which these models reject. They now show the reasoning-effort selector (including a new **Max** option) instead of temperature/Top P, consistent with other reasoning models.

* **Perplexity model costs now load correctly** — Perplexity was missing from the provider list used to load pricing data, so calls through Perplexity (including the `sonar` model family) were costed at \$0. Pricing now loads correctly.

* **Clearer errors from Python-based LLM-as-judge scoring** — A metric evaluation that hit an empty or malformed response from the scoring service previously failed with an opaque server error; it now surfaces the actual underlying error message.

* **Tool-call spans from OpenTelemetry now typed correctly** — Spans carrying the OTel `gen_ai.tool.call.arguments` / `gen_ai.tool.call.result` attributes are now classified as **tool** spans in the UI, matching spans tracked directly through the SDK.

* **Fixed intermittent 500s retrieving threads and experiment project lists** — Two endpoints (thread lookup, and listing the projects an experiment spans) could return an unmapped server error when the underlying query legitimately returned more than one row for the same logical entity. Both now handle this correctly.

* **More accurate error codes from LLM providers** — Errors from OpenAI-compatible providers (including calls routed through OpenRouter) were sometimes parsed with the wrong error model, which could turn a rate-limit error into a generic server error. The correct error model is now selected based on the error's shape, so the real status code is surfaced.

* **LangChain integration: cost capture through proxies, and provider override** — `OpikTracer` accepts a new `provider` argument to override provider auto-detection, useful when routing `ChatOpenAI` through a proxy such as LiteLLM. When the proxy reports its own cost via the `x-litellm-response-cost` response header (with `include_response_headers=True` on `ChatOpenAI`), Opik now uses that cost instead of estimating \$0.

* **`opik migrate` no longer risks OOM on partially-stale projects** — When a migrated experiment referenced only deleted or timestamp-less traces, the CLI fell back to an unbounded read of the entire project's spans. It now skips the unrecoverable span data for that batch (with a warning) and continues the migration instead of risking an out-of-memory crash.

* **SDK tolerates new OpenAI usage fields and avoids a broken LiteLLM release** — The SDK no longer breaks when OpenAI adds new fields to its usage payload, and the LiteLLM dependency now excludes the `1.92.*` release line, which crashed on import without optional proxy dependencies installed.

* **Dataset and test suite creation logs point to the right project** — The URL logged by the SDK after creating a dataset or test suite now links directly to that resource inside its actual project, instead of a generic link that could resolve to the wrong project.

## Performance Improvements

* **Self-hosted: ClickHouse upgraded to 26.3 LTS** — The bundled ClickHouse instance (Helm chart, docker-compose, and testcontainers) is now pinned to 26.3 LTS. No manual data migration is required, though major-version upgrades may take longer to complete.

* **Self-hosted: ingestion load spreads more evenly across backend pods** — The frontend nginx proxy now recycles its keepalive connections to the backend on a schedule, instead of holding them open indefinitely. On Kubernetes deployments this prevents a small number of backend pods from absorbing a disproportionate share of ingestion traffic while others sit idle.

---

And much more! 👉 [See full commit log on GitHub](https://github.com/comet-ml/opik/compare/2.1.15...2.2.0)

*Releases*: `2.1.25`, `2.1.26`, `2.1.27`, `2.1.28`, `2.1.29`, `2.1.30`, `2.1.31`, `2.1.32`, `2.2.0`

## July 13, 2026

## Per-Evaluation Spend Budget for LLM-as-Judge Evaluators

Online LLM-as-judge scoring rules for traces and threads can now be given a **Max cost per evaluation (USD)** budget. Once an evaluation's cumulative spend crosses the limit, the agentic scoring loop stops starting new tool-calling turns and returns a best-effort verdict from whatever evidence it has already gathered, instead of continuing to spend without bound. The monitoring trace for a capped run is tagged `budget_exceeded` so you can tell a budget-limited verdict apart from a fully-investigated one. The field is hidden for span-scope rules, since a single LLM call has no loop to cap.

## Optimization Runs: New-Run Sidebar & Redesigned List

Starting a new optimization run no longer navigates away from the runs list: the "New optimization run" form now opens as a side panel over the list (via the list's `?new` — or `?template` / `?rerun` — parameter) so you can start, clone, or re-run an optimization without losing your place. The runs list itself has a new **Item source** column (the same reusable cell used on the Experiments page), Run ID / Algorithm / Metric columns available to add, and the "Pruned" trial status is now labeled **Discarded** for clarity.

## Bug Fixes & Improvements

* **Diagnostics: clearer failure copy and working billing link** — Follow-ups to last week's Diagnostics failure reasons: a permission-denied run failure now shows explanatory copy instead of a generic message, and the "View billing" link in the failure dialog now opens the right settings page.

* **Claude Agent SDK / Claude Code traces now show real input/output** — Spans emitted by the Claude Agent SDK and Claude Code carried their prompt, tool, and response content on attributes that no mapping rule recognized, so everything fell into a generic attribute bag instead of the trace's input/output fields. These spans are now mapped correctly, and model/provider are set so cost is calculated instead of showing as \$0.

* **New: Mistral AI Python integration (`track_mistral`)** — A native integration for the `mistralai` Python SDK: wrap a client with `track_mistral(client)` to trace `chat.complete`/`chat.complete_async` and `chat.stream`/`stream_async` calls (including structured-output `parse` calls and streamed responses aggregated into a single span), with input, output, token usage, and cost captured automatically.

* **`opik migrate dataset` can resume after an interruption** — A crash, network drop, or OOM during `opik migrate dataset` used to mean starting over. Progress now checkpoints after each completed experiment, so a re-run picks up where it left off, and the source dataset keeps its original name until the destination copy is fully verified — a failed run leaves only a discardable temporary copy behind rather than a renamed source.

* **`opik migrate dataset` no longer runs out of memory on large datasets** — The migration's read path could request full-fidelity pages large enough to exhaust the process's memory and get killed mid-run. Reads are now paged more conservatively, with an automatic page-size shrink on read timeouts/connection errors.

* **`opik export`/`opik import` preserve tags for prompts, experiments, and datasets** — Tags on prompts, experiments, and datasets were silently dropped when exporting or importing via the CLI (traces and spans already round-tripped correctly). All four entity types now preserve tags end-to-end.

* **`opik export ... all` no longer runs out of memory on large projects** — Exporting a project used to buffer every trace and span in memory before writing anything, so peak memory scaled with project size. Export now processes traces in bounded chunks, flushing and discarding each chunk as it's written, and an interrupted export resumes mid-project instead of restarting.

* **Fixed duplicate rows in the test suite detail view** — Editing a dataset item could cause its experiment result to render multiple times in the test suite detail view, due to a join that matched more than one dataset-item version per item.

* **Prompt names are now unique per project instead of per workspace** — Two prompts with the same name in different projects no longer conflict.

* **Custom metrics fail loudly on an unresolvable reference key** — An `equals`/`levenshtein_ratio`/`numerical_similarity` metric configured with a reference key that matches no dataset field used to silently score every item 0, making a broken metric indistinguishable from a genuinely low-scoring run. This now raises a clear error at build time listing the available fields, and a per-item missing value is scored 0 with an explicit "Missing reference value" reason instead of a spurious perfect match.

* **Fixed cost calculation for the highest LiteLLM pricing tiers** — Models with `above_128k_tokens` (e.g. Gemini 1.5 Flash) or `above_272k_tokens` (e.g. the GPT-5.4/5.5 family and their Azure-hosted equivalents) pricing tiers were being undercharged, since only the `above_200k_tokens` tier was applied. All published tiers are now taken into account.

* **Fixed an outbound gzip decompression error** — Self-hosted deployments with `jerseyClient.gzipEnabled: true` could hit `ZipException: Not in GZIP format` on outbound calls that read a gzip-encoded response body (for example, some Ollama responses), because the response was decompressed twice. The stale `Content-Encoding` header is now stripped after the first decode.

## Performance Improvements

* **Faster Logs page loads on large projects** — The Traces/Spans/Threads "log your first trace" onboarding check used to run a full project-wide aggregation query just to see if any row existed. It now uses a minimal existence check, removing a multi-second scan from every Logs page load.

* **Reduced full-table scans in dataset and experiment queries** — Several ClickHouse queries backing dataset-item filters and experiment views scanned far more data than needed — in one case, an experiment/dataset-item filter query read 75x fewer granules after being scoped to the relevant experiment's trace IDs instead of the whole workspace. These queries now carry tighter project/trace-id/workspace bounds, reducing ClickHouse load on large workspaces.

---

And much more! 👉 [See full commit log on GitHub](https://github.com/comet-ml/opik/compare/2.1.15...2.1.22)

*Releases*: `2.1.16`, `2.1.17`, `2.1.18`, `2.1.19`, `2.1.20`, `2.1.21`, `2.1.22`

## July 6, 2026

## Ask Ollie to Explain Trace Cells

Error, duration, and cost cells across the Traces, Spans, and Threads tables now have an "Explain" button (the Ollie owl icon) that streams a plain-language explanation of that specific value without leaving the table — no need to open the trace panel and dig through raw payloads to understand why a call errored, took as long as it did, or cost what it did.

Explanations are cached per cell, so reopening a popover shows the previous answer immediately instead of re-running the query, and a "Continue conversation" link hands off to the full Ollie sidebar when you want to dig deeper.

## Online Evaluation Runs Are Now Traced

LLM-as-a-judge scoring runs (on traces, spans, and threads) are now recorded as monitoring traces automatically, with a `prepare_evaluation` span and one span per scoring call, including token usage and cost. They're hidden from the main Logs view by default, and each rule on the **Online Evaluation** page now has a **Go to traces** action that opens a scoped, filtered view of that rule's evaluation activity — useful for confirming a rule is running, checking what it cost, or debugging why it errored. This is on by default for all workspaces.

## Bug Fixes & Improvements

* **Diagnostics now explains why a run failed** — Previously, a Diagnostics run that crashed or never started (for example, from exhausted LLM credits) just spun until it timed out, with no indication of what went wrong. Failed runs now show a specific reason (out of credits, rate limited, provider error, or never started) with a "Try again" action, and the failure timeout was cut from 12 minutes to 5.

* **Playground and online scoring: Sonnet 5 and Fable 5 no longer error out** — These models were missing from the capability list, so the Playground kept showing temperature/Top P sliders and sending those parameters, which the models reject. The sliders are now hidden and the parameters are no longer sent, matching the behavior already in place for Opus 4.7/4.8.

* **Azure-hosted OpenAI model prices now load correctly** — Model prices for the `azure/*` model family (`gpt-4o`, `gpt-5`, `codex-mini`, and others) were silently failing to load because of a provider-mapping gap, causing cost to show as unavailable for these models.

* **Experiment and Playground traces now show up in scoped "Go to logs" views** — Traces created by experiment runs and Playground runs were sometimes missing from their respective "Go to logs" tables even though the traces existed, due to a visibility-filter mismatch. These views now correctly scope by the experiment or run itself.

* **Playground: dataset and metric selection redesigned** — The "Run experiment" dialog is replaced by inline dataset and metrics dropdowns in the Playground header, so you can change either independently without reopening a modal. Metrics can also be created and edited in place.

* **Permission gating extended to more pages** — Users with the Annotator workspace role no longer see Agent Playground, Online Evaluation, or Alerts in the sidebar. Prompt Library view/edit access is now governed by the same permission system, at both the UI and API level.

* **`opik import --to-workspace` for cross-workspace imports** — The `opik import` CLI command accepts a `--to-workspace` option to import exported data into a different workspace than the one it was exported from, without needing to restructure the export directory as a workaround.

* **`opik migrate dataset --exclude-experiments`** — This flag skips migrating a dataset's experiments and optimizations, migrating only the dataset and its version history. Useful for large datasets where the experiment/optimization cascade isn't needed.

* **Fixed intermittent errors loading prompt and dataset versions under load** — A query pattern that caused MySQL to materialize large temporary tables could fail outright under load. Prompt and dataset version lookups are now rewritten to avoid the issue.

## Performance Improvements

* **Faster trace ingestion** — Removed a redundant ClickHouse lookup that ran on every trace ingestion just to check the trace's last-updated timestamp, reducing ingestion latency and ClickHouse load.

* **SDK: shared connection resources across `Opik()` clients** — Code that creates multiple `Opik()` clients (for example, one per task or request) now shares the underlying connection pool and background threads across clients with matching configuration instead of creating a full new stack each time, reducing thread and connection overhead.

* **Self-hosted: ClickHouse `FINAL` reads no longer over-read on skip-indexed tables** — Disabled the new ClickHouse 25.x default for `use_skip_indexes_if_final_exact_mode`, which was causing skip-indexed queries to read far more data than necessary; Opik's queries already prune by primary key and project ID, so the exact mode isn't needed.

---

And much more! 👉 [See full commit log on GitHub](https://github.com/comet-ml/opik/compare/2.1.10...2.1.15)

*Releases*: `2.1.11`, `2.1.12`, `2.1.13`, `2.1.14`, `2.1.15`

## June 23, 2026

## LLM-as-Judge Scorer Now Evaluates Trace Attachments

Online LLM-as-judge scoring rules can now reason over files attached to traces — images, PDFs, documents, and other binary content — in addition to the trace's text fields. Previously, attachment-aware evaluation was only available for thread (conversation) scoring; single-trace scoring rules now share the same capability.

When a trace carries attachments, the scorer automatically routes to an agentic tool-calling loop. The judge model receives a `ReadTool` that lists the attached files and a `GetAttachmentTool` that fetches each file as multimodal content. Up to eight attachments are injected per evaluation. When the configured provider does not support tool calls, a warning is surfaced rather than silently failing.

## Project-Scoped `opik export` / `opik import`

The `opik export` and `opik import` CLI commands now require a project argument, matching how Opik v2 organises data (every dataset, prompt, and experiment lives inside a project).

**New command shape:**

```bash
opik export WORKSPACE PROJECT ITEM [NAME] [OPTIONS]
```

Where `ITEM` is one of `all`, `dataset`, `traces`, `experiment`, or `prompt`. On-disk output is written to `PATH/WORKSPACE/projects/<id>/` with a `project.json` name index, so the same `--path` value round-trips between export and import.

Import gains a `--to-project` option that redirects items into a different project at import time.

👉 [Export and import documentation](https://www.comet.com/docs/opik/tracing/advanced/export-data)

## Bug Fixes & Improvements

* **MCP server: `opik mcp status` command + auto-detection of hosted server** — `opik configure` and `opik mcp configure` now probe for a Comet-hosted MCP server (HTTP + OAuth) and fall back to the local `uvx opik-mcp` server only when the hosted one is unavailable. A new `opik mcp status` command (and the equivalent `opik configure status`) reports the active Opik configuration and which AI assistants have the MCP server registered, flagging any drift from `~/.opik.config`. Use `--local-server` to force the local path.

* **Vercel AI SDK v7 and Vercel eve agent support** — `OpikExporter` now captures traces from AI SDK v4 through v7 and from the Vercel eve agent framework. Multi-turn eve conversations are automatically grouped into a single thread via the session ID, and cached-token usage is captured.

* **`track_openai`: custom `provider` argument** — An optional `provider` parameter (`str` or `opik.LLMProvider`) can now be passed to `track_openai` to override the auto-detected provider on every span. Useful for OpenAI-compatible APIs where the hostname-derived default is not descriptive.

  ```python
  client = track_openai(openai.OpenAI(base_url="..."), provider="my-provider")
  ```

* **Configurable local runner poll interval** — The `opik connect` / `opik endpoint` runner polls for jobs every 0.5 s by default (\~120 requests/min). Setting `OPIK_RUNNER_POLL_INTERVAL` (seconds) reduces this for environments where firewalls or proxies throttle sustained polling. Sustained failures now escalate to an actionable warning with a firewall/proxy note rather than going silent.

* **Assertion scoring errors now shown in the UI** — When an LLM-as-judge assertion scorer fails (for example, because the configured model is unavailable), an orange **Scoring error** badge with an actionable tooltip is shown instead of a spinning indicator or a generic grey **Skipped** tag.

* **Test suite items no longer lost after editing an assertion** — A ClickHouse parameter-escaping bug caused the `\n` sequences inside LLM-judge prompts to be unescaped to literal newlines, producing invalid stored JSON and silently dropping the row on read. Items now survive edit round-trips correctly.

* **ChrF metric: `char_order` and `ignore_whitespace` now forwarded to NLTK** — These parameters were stored in the `ChrF` object but never passed to `nltk.translate.chrf_score.sentence_chrf`. `char_order` is now forwarded as `max_len` and `ignore_whitespace` is forwarded as-is.

* **OpenAI audio models: audio input tokens billed at the correct rate** — For models such as `gpt-4o-audio-preview` and `gpt-4o-realtime`, audio input tokens are now charged at the `input_cost_per_audio_token` rate (up to 16× the standard text rate). Previously they were silently billed at the standard text input rate.

* **`opik export` span fetch performance** — When exporting a filtered subset of traces, spans are now fetched with a `trace_id` range bound derived from the matched trace IDs rather than scanning the entire project. Large filtered exports that previously stalled now complete promptly.

* **Optimization studio renamed to Optimization runs** — The v2 UI section previously called "Optimization studio" is now labelled **Optimization runs** everywhere.

## Performance Improvements

* **Attachment base64 detection is now O(n)** — The regex that detects whether an attachment payload is already base64-encoded was rewritten with a possessive quantifier, making it linear in the payload size and avoiding catastrophic backtracking on large binary attachments.

* **Self-hosted: ClickHouse upgraded to 25.8 LTS** — The bundled ClickHouse instance (Helm chart, docker-compose, and testcontainers) is now pinned to 25.8 LTS and the operator to 0.27.1. This release requires no manual migration; existing data is read correctly by the new version.

---

And much more! 👉 [See full commit log on GitHub](https://github.com/comet-ml/opik/compare/2.0.73...2.1.5)

*Releases*: `2.0.77`, `2.1.0`, `2.1.1`, `2.1.2`, `2.1.3`, `2.1.4`, `2.1.5`

## June 16, 2026

## PydanticAI and Logfire Trace Fidelity Improvements

Traces from PydanticAI agents ingested over OpenTelemetry (via Logfire) now show complete, correctly-typed data and accurate cost. Several gaps in the OTel → Opik span mapping were fixed:

* **Tool call I/O is now captured** — tool span inputs and outputs now appear in the correct Input/Output fields. Previously, logfire's `tool_response` attribute fell into the Input bucket and the Output field was empty.
* **Errors surface at the trace level** — when a span carries an OTel `exception` event or `STATUS_CODE_ERROR` status, the error is translated into `error_info` and propagated up to the parent trace so failed agent runs are flagged at the top level.
* **Gemini/Google cost now calculated** — PydanticAI reports the generic `google` provider name, which previously matched no pricing row and showed \$0 cost. The backend now resolves it to `google_vertexai` or `google_ai` based on the server hostname.
* **Agent-run spans correctly typed** — spans from `invoke_agent` operations are now typed `general` instead of `llm`, so the trace tree accurately reflects the agent structure.

The error-surfacing fix is generic and benefits all OTel integrations; the tool I/O and cost improvements are specific to PydanticAI and Logfire.

## Quick MCP Server Setup via `opik configure`

Setting up Claude Desktop (or any MCP client) to connect to Opik no longer requires manually editing JSON config files. Running `opik configure` now offers an opt-in MCP setup step at the end of the flow. The standalone `opik mcp configure` command runs the same wizard on its own. Both automatically write the Opik MCP server block to your Claude Desktop configuration.

👉 [MCP server documentation](https://www.comet.com/docs/opik/mcp-server)

## Bug Fixes & Improvements

* **OpenClaw added to onboarding integrations** — The Get Started page now includes OpenClaw with its four-step CLI setup guide (install plugin → configure → status → run), rendered with bash syntax highlighting. The integration details dialog also supports multi-step and non-pip integrations going forward.

* **Logs view: search bar repositioned inline with filters** — The search bar now sits alongside active filter chips in a single wrapping row. Toolbar controls (Columns, Row height, Refresh) moved to the tab row, reducing visual clutter in the main search area. Chart bars changed to violet and table tags to green.

* **Annotation queue: scores scoped per queue** — Feedback scores and comments now carry a `source_queue_id` so each queue's metrics (average scores, reviewer counts) are computed independently. A score submitted through one queue no longer counts toward another queue's statistics when the same trace appears in both.

* **Annotation queue: false "another annotator is reviewing" message resolved** — After submitting a feedback score, the lock heartbeat now correctly recognizes existing lock holders, so annotators no longer see a false "another annotator is reviewing this item" message immediately after scoring.

* **Thread-scoped annotation queues no longer crash** — Opening a thread annotation queue without additional queue filters no longer throws an `Unknown expression or function identifier` error and crashes the queue view.

* **Experiments page: charts stay fixed during horizontal table scroll** — The feedback-score charts now remain pinned to the top while the experiments table scrolls horizontally underneath.

* **Playground: immediate failure on exhausted OpenAI quota** — When an OpenAI API key has insufficient quota (HTTP 429 `insufficient_quota`), the Playground now returns an error immediately instead of retrying the request multiple times and delaying the response.

---

And much more! 👉 [See full commit log on GitHub](https://github.com/comet-ml/opik/compare/2.0.66...2.0.73)

*Releases*: `2.0.67`, `2.0.68`, `2.0.69`, `2.0.70`, `2.0.71`, `2.0.72`, `2.0.73`

## June 9, 2026

## Resume Interrupted Evaluations with `evaluate_resume`

Long-running evaluation jobs that get cut short — by Ctrl-C, an OOM error, a failed scoring metric, or a network blip — can now be continued from where they stopped instead of restarting from scratch. `opik.evaluate_resume(experiment_id, task, scoring_metrics=[...])` replays only the trials that did not complete, merges them with the ones that did, and returns a single `EvaluationResult` covering the whole experiment.

A trial counts as complete only when `trace.output` is set, which happens after the task, scoring, and score-logging all succeed. Any failure mode that prevents reaching that point — a metric raising an exception, a `KeyboardInterrupt` between task and scoring — leaves the trial replayable.

```python
import opik

# Continue a partially-completed experiment — only missing trials are replayed
result = opik.evaluate_resume(
    experiment_id="...",
    task=my_task,
    scoring_metrics=[Equals()],
)
```

The original `evaluate(...)` call writes a resume snapshot into `experiment_config` so the exact iteration (pinned dataset version, sample count, per-item trial counts) can be reconstructed server-side. When the original call used a custom `dataset_sampler` or explicit `dataset_item_ids`, the SDK also writes a local checkpoint next to the experiment ID for those cases.

👉 [Resume evaluations documentation](https://www.comet.com/docs/opik/evaluation/advanced/resume_evaluations)

## OpenAI Responses API Support in Playground and LLM-as-a-Judge

The Playground and LLM-as-a-Judge now support OpenAI's `/v1/responses` API, making it possible to use o-series reasoning models (`o1`, `o3`, `o3-mini`, `o4-mini`) and other deployments that are only available on the newer API path. Previously, sending these models through the Chat Completions path returned "This is not a chat model and thus not supported in the v1/chat/completions endpoint."

To opt in, open the **Manage AI Providers** dialog, select your OpenAI key, and set **Pipeline mode** to **Responses API**. The Chat Completions path remains the default and is unchanged for all other models.

The Playground's `Top P` slider is now also hidden for OpenAI reasoning models (`gpt-5.x`, `o1*`, `o3*`, `o4*`). Those models reject `top_p` outright; the slider was causing 400 errors when it appeared.

## Bug Fixes & Improvements

* **Annotation queues: claim mechanism for parallel annotation** — Multiple annotators working the same queue simultaneously now see each item locked while another reviewer is looking at it, preventing duplicate work. Items show an "In review" indicator (orange) when all annotator slots are occupied by a combination of active locks and existing scores. Locks are kept alive by a heartbeat and expire via TTL when the reviewer navigates away. The sidebar also gains "To review" / "Processed" filter tabs, and each annotator sees items in a distinct shuffled order to reduce contention.

* **Collapsible JSON/YAML in trace and span detail view** — JSON objects, arrays, and YAML blocks in the trace and span detail view can now be folded and unfolded with an inline chevron at the end of each foldable line. Collapsed blocks render as a clickable gray placeholder. This makes it easier to navigate large payloads without scrolling past content you don't need.

* **Redesigned dataset and test suite creation flow** — The creation dialog now presents two explicit paths: **Upload a file** (CSV or JSON dropzone with auto-naming and optional evaluation criteria for test suites) and **Use SDK** (name + code snippet). Both options are accessible from the header button and from the list empty state. On success the panel closes and a "Go to …" toast appears.

* **Evaluate experiment traces directly from the UI** — The Compare Experiments page has a new **Evaluate** button (brain icon) in the action bar. It opens the online evaluation dialog scoped to all traces in the current experiment, so you can score an experiment's output without leaving the page.

* **Span filtering by `created_at` and `last_updated_at`** — The span search API now accepts `created_at` and `last_updated_at` as filter fields with all comparison operators (`=`, `!=`, `>`, `>=`, `<`, `<=`). These fields were already supported on traces; span support was missing.

* **OpenTelemetry: in-process spans now linked to the active `@opik.track` trace** — When an OTel-instrumented library (such as logfire or PydanticAI) emits spans from inside an `@opik.track`-decorated function, those spans are now nested under the active tracked trace rather than starting a separate trace. Distributed flows where parent spans or W3C baggage carry Opik IDs continue to take precedence over the in-process context.

* **Optimization: best trial configuration now shows the optimized prompt** — The Best Trial Configuration panel was displaying the baseline prompt instead of the prompt produced by the optimizer. It now shows the correct optimized result. The Trials table also gains a Prompt column with per-message formatting and a diff-vs-baseline popover.

* **Experiment views: prompt version labels instead of commit hashes** — The Experiments table, the single-experiment Configuration tab, and the Dashboard Experiments leaderboard now display prompts as "name (v3)" instead of raw commit hashes, consistent with the display already used in the Prompt Library.

* **AI Spend dashboard: total tokens KPI card and onboarding empty state** — The placeholder "Budget remaining" card is replaced by a **Total tokens** KPI showing the sum of all token tiers across models, with a period-over-period trend indicator. The dashboard also shows an onboarding empty state with setup instructions and a ready-to-copy configuration snippet when no trace data has been received yet.

* **Cost calculation: tiered pricing above 200k tokens now applied** — For models such as `gemini-2.5-pro` and `vertex_ai/claude-sonnet-4-5` that carry `above_200k_tokens` rate tiers, requests exceeding the 200k-token threshold were being billed at the base input rate. Opik now applies the tier rate when the threshold is crossed (the entire request is billed at the tier price, mirroring LiteLLM's semantics).

* **Cost calculation: Claude on Vertex AI cached tokens now discounted** — Claude models on Vertex AI (`vertex_ai/claude-haiku-4-5`, `vertex_ai/claude-sonnet-4-5`, `vertex_ai/claude-opus-4-1`) were having cache-read tokens billed at the full input rate. They now use the Anthropic cache calculator, correctly applying the discount on cached tokens.

* **Vertex AI: model selection preserved across provider switches** — Switching away from Vertex AI and back in the Playground no longer resets the previously selected model.

## Performance Improvements

* **Span timestamp filters use ClickHouse skip indexes** — `created_at` and `last_updated_at` on the `spans` and `traces` tables now have `minmax` skip indexes. Range filters on these columns prune granules instead of scanning the full project partition, significantly reducing query time and ClickHouse CPU load on large tables.

---

And much more! 👉 [See full commit log on GitHub](https://github.com/comet-ml/opik/compare/2.0.52...2.0.59)

*Releases*: `2.0.53`, `2.0.54`, `2.0.55`, `2.0.56`, `2.0.57`, `2.0.58`, `2.0.59`

## June 2, 2026

## Prompt Library Now Available in Opik 2.0

The Prompt Library is now part of the Opik 2.0 UI, accessible from the project sidebar under **Prompt library**. Alongside that, prompt versions have gained first-class environment support — you can tag a version as `production` or `staging` and retrieve it by name from the SDK, without tracking version numbers in application code.

**What's new:**

* **Prompt Library in the sidebar** — the Prompt Library is now a top-level section in every project in the new Opik 2.0 interface
* **Fetch by environment** — `client.get_prompt(name, environment="production")` returns the version currently tagged as production; `version` and `environment` are mutually exclusive and passing both raises a clear error
* **Assign environments** — `client.set_prompt_environments(name, ["production", "staging"])` replaces the full environment set on a version; the same environment is automatically moved away from whatever version previously held it
* **Tag at creation** — `client.create_prompt(name, content="...", environments=["staging"])` and `client.create_chat_prompt(...)` accept `environments` directly
* **TypeScript parity** — `setPromptEnvironments`, `getPrompt({ environment })`, and `createPrompt({ environments })` mirror the Python API
* **Sequential version numbers** — prompt versions now show as `v1`, `v2`, `v3` in the UI and API instead of raw commit hashes
* **Environment badges everywhere** — assigned environments appear next to every version reference in the prompt library, history timeline, diff view, and Playground
* **Terminology update** — "commit" has been replaced with "version" throughout the prompt UI

```python
# Tag a version at creation time
prompt = client.create_prompt("system-prompt", content="...", environments=["staging"])

# Retrieve by environment — no hard-coded version number needed
production_prompt = client.get_prompt("system-prompt", environment="production")

# Promote a specific version to production
client.set_prompt_environments("system-prompt", ["production"], version="v3")
```

## Simplified Filters in the Logs View

The Traces, Spans, and Threads tabs now have a redesigned filter bar that makes it faster to narrow down what you're looking at. Filters appear as chips directly in the toolbar — pick a field, set a value, and the table updates instantly. Frequently-used filters can be **pinned** to the bar so they're always one click away, and filter state is preserved in the URL so you can share an exact filtered view with a teammate.

## Bug Fixes & Improvements

* **Test suite assertions: sub-span inspection** — the evaluator LLM can now issue `get_trace_spans` and `read` tool calls to inspect intermediate spans during evaluation, enabling correctness checks about tool usage, model selection, and per-span errors inside complex agents
* **Google ADK integration: images render in trace attachments** — URL-safe base64 image data sent by ADK is automatically normalized to standard base64; PNG, JPEG, GIF, and WebP attachments all render correctly
* **Optimization trials page: all constituent experiments shown** — experiments belonging to multi-project optimizations are now visible from the trials page regardless of which project the user is currently viewing the optimization from
* **Error rate KPI: now shows a percentage** — the error rate dashboard card was displaying a raw event count; it now shows the rate as a percentage
* **Annotation queue: trace logs shown inline** — trace log entries are rendered inline on the annotation queue page instead of requiring navigation away
* **Online evaluation rules: ClassCastException resolved** — thread-level rules that include filters no longer throw a `ClassCastException` under certain configurations
* **Attachments: data URI prefix handled** — base64 attachment payloads that include a `data:<type>;base64,` prefix are now stripped correctly in both the SDK and the frontend
* **SDK: built-in environment colors preserved** — workspace environments with reserved names retain their designated color after updates or syncs
* **`opik migrate`: skipped items reported clearly** — the migration command now reports each skipped item with its reason, count, and sample source IDs, and exits with code 1 so CI pipelines detect incomplete migrations
* **Qianfan integration documentation** — the Qianfan LLM provider integration now has a dedicated documentation page

---

And much more! 👉 [See full commit log on GitHub](https://github.com/comet-ml/opik/compare/2.0.47...2.0.52)

*Releases*: `2.0.48`, `2.0.49`, `2.0.50`, `2.0.51`, `2.0.52`

## May 26, 2026

## AND/OR Condition Grouping in Alerts

Alert rules now support structured condition grouping: conditions within a group are evaluated with AND, while groups themselves are combined with OR. This makes it possible to express logic such as "flag a trace if (hallucination score > 0.8 AND relevance score \< 0.3) OR (toxicity score > 0.5)".

Existing single-condition alerts continue to work exactly as before — each legacy condition is automatically treated as its own group, so no migration is needed.

## Bug Fixes & Improvements

* **Prompt masks (Python & TypeScript SDKs)** — `prompt_mask_context(masks)` / `promptMaskContext(masks)` lets you run agent code with specific prompt IDs silently redirected to a different version ID, non-destructively. The agent calls `get_prompt()` as usual and receives the overridden template without any permanent change to the prompt library. Designed for A/B testing and optimizer sweep scenarios.
* **Experiments: dataset version shown inline** — the dataset version is now displayed as a pill alongside the item source in both the experiments table and the experiment detail header. The standalone "Test suite version" column has been removed; the same information is now visible in context.
* **Dataset items: conflicting key names no longer cause errors** — iterating a dataset whose items contain a key that matches a `DatasetItem` field (e.g. `id`, as in HotpotQA) previously raised `TypeError: multiple values for keyword argument`. The SDK now strips conflicting keys and emits a one-time warning so iteration completes.
* **Harbor integration: supports harbor `<0.8` and `>=0.8`** — `track_harbor()` now patches whichever method name the installed version of harbor exposes (`_setup_environment` or `_setup_agent_environment`), so tracing works regardless of which version is installed.
* **New Playground models** — Gemini 3.5 Flash and `qwen/qwen3.7-max` are now available in the model picker.

---

And much more! 👉 [See full commit log on GitHub](https://github.com/comet-ml/opik/compare/2.0.37...2.0.47)

*Releases*: `2.0.42`, `2.0.43`, `2.0.44`, `2.0.45`, `2.0.46`, `2.0.47`

## May 19, 2026

## 🚀 Client-Side Prompt Caching (Python & TypeScript SDKs)

`client.get_prompt()` and `client.get_chat_prompt()` now cache results in-process, so repeated calls inside a hot path skip the network round-trip entirely. Pinned commits are cached indefinitely; latest-version lookups use a 5-minute TTL that refreshes in the background so your code always gets a reasonably fresh value without blocking.

**What's new:**

* **Automatic caching** — results are cached on the first fetch; subsequent calls return instantly from memory
* **Configurable TTL** — set `OPIK_PROMPT_CACHE_TTL_SECONDS` to adjust the freshness window (default: 300 s)
* **Bypass when needed** — pass `no_cache=True` / `noCache: true` to force a live fetch from the backend
* **Prompt metadata injected into traces** — when you fetch a prompt inside an `@track` context, the prompt ID and commit are automatically recorded in the trace metadata so you know which version was used at inference time
* **TypeScript SDK** — the same caching and metadata injection are available in the TypeScript client

```python
# Cached after the first call — no extra latency on subsequent invocations
prompt = client.get_prompt("my-system-prompt")

# Force a fresh fetch, bypassing the cache
prompt = client.get_prompt("my-system-prompt", no_cache=True)
```

## 🔌 `opik connect` CLI Improvements

The `opik connect` and `opik endpoint` CLI commands have been reorganized with a much better error experience:

* **Formatted error output** — configuration problems now show a labelled card (Reason / Workspace / URL / Config / Fix / Docs) so you can see exactly what's wrong and how to fix it without reading a stack trace
* **Auto-configure on first run** — if no `~/.opik.config` file exists, `opik connect` now offers to run `opik configure` automatically (skipped in non-interactive / headless environments)
* **Instant disconnect** — stopping a local runner session now notifies the backend immediately, so the connection status in the UI updates right away instead of waiting for a timeout

## 🔧 Bug Fixes & Improvements

* **Playground: Gemma 4 no longer leaks reasoning traces** — the internal thinking output from Gemma 4 models was appearing at the top of Playground responses; it is now suppressed so you see only the final model response
* **Playground: updated default models for Gemini and Vertex AI** — the provider dropdown previously defaulted to deprecated model aliases that weren't selectable in the picker; both providers now default to their current recommended models
* **Google ADK integration: re-patching fixed** — `OpikADKOtelTracer` was killing all active OpenTelemetry spans and re-patching the ADK exporter on every request; the patcher is now idempotent and preserves user-configured OTel pipelines
* **Environments: auto-created environments get distinct colors** — environments created automatically from trace ingestion now receive a color from the palette (assigned deterministically by name hash), so they no longer all appear identical
* **Environments: inline validation errors** — when creating or editing an environment, backend errors (duplicate name, invalid value) now appear inline below the field; the dialog stays open so you don't lose your input
* **Environments: SDK preserves environment on update** — calling `span.end()`, `span.update()`, `trace.end()`, or `trace.update()` no longer clears the `environment` field set at creation time
* **Experiments: "Item source" column** — the column previously labeled "Test suite" in the Experiments table is now called "Item source" and shows a dynamic icon reflecting the actual source (dataset, trace, manual, etc.)
* **Dataset version copy no longer drops items** — when a dataset version was copied and the stored item count had drifted from the actual ClickHouse row count, some items could be silently lost; the copy now uses a live count, eliminating the discrepancy
* **Self-hosted onboarding skip no longer loops** — clicking "Skip" during onboarding on deployments without demo data (self-hosted Docker Compose, self-hosted EKS) previously started a 5-minute polling loop; it now routes directly to the home page
* **CSV dataset upload always available** — the CSV upload button in the dataset UI is now always shown; it was previously hidden on self-hosted Docker Compose and some staging environments even though the feature was fully functional

## ⚡ Performance Improvements

* **Dataset streaming uses less backend CPU** — resolved a query pattern that caused the MySQL reader to scan all dataset versions on every `/datasets/items/stream` call; under high request volume this was pushing database CPU to 80–99%, it now uses a direct primary-key lookup instead
* **Workspace selector loads faster** — the workspace dropdown now fetches a lighter summary endpoint, reducing the number of backend queries on each page load for users with many workspaces

---

And much more! 👉 [See full commit log on GitHub](https://github.com/comet-ml/opik/compare/2.0.31...2.0.37)

*Releases*: `2.0.32`, `2.0.33`, `2.0.34`, `2.0.35`, `2.0.36`, `2.0.37`

## May 12, 2026

Here are the most relevant improvements we've made since the last release:

## 🌍 Environment Tracking for Traces, Spans & Threads

You can now tag traces, spans, and threads with an **environment** field — `production`, `staging`, `dev`, or any label you define. This makes it easy to separate signal from noise: filter your project's trace view to only production issues, or compare behavior between environments without spinning up separate projects.

**What's new:**

* **Environment column in the Logs view** - Traces and spans tables now show the environment and support filtering, so you can slice by `production` vs `staging` in a single project
* **Auto-create environments from ingestion** - Environments are created automatically the first time a trace with a new environment name arrives; no setup required
* **Python SDK support** - Pass `environment` to `@track`, `opik.trace()`, or `opik.span()` — and it's preserved through `.end()` and `.update()` calls
* **TypeScript SDK support** - Set `environment` on trace and span creation

```python
import opik

@opik.track(environment="production")
def my_agent(input: str) -> str:
    ...
```

👉 [Environments Documentation](https://www.comet.com/docs/opik/tracing/advanced/log_traces#environments)

## 🧪 Test Suite Assertions Can Now Inspect Sub-Spans

Test suite assertions can now look inside a trace — not just the top-level input/output — to reason about tool calls, intermediate LLM steps, and sub-agent behavior. The evaluator LLM gets access to two on-demand tools: `get_trace_spans` (lists all sub-spans for the trace) and `read` (fetches a specific span by ID) — so it can drill into exactly what happened at each step.

**Why it matters:** Previously, an assertion could only see what went in and came out of the agent. Now it can check whether the right tool was called, which model was used in an intermediate step, or whether a specific span had an error — enabling far more meaningful correctness checks for complex agents.

## ⚡ Dramatically Faster Trace Table Loading

Traces and spans tables no longer download attachment bytes (images, PDFs) when loading a list — attachments are lazy-loaded only when you open an individual trace. In our benchmarks with image and PDF attachments, this reduced the per-page payload from **85 MB → 0.13 MB** and load time from 3.4 s to 0.1 s.

**Why it matters:** If any of your traces include file attachments, the table was silently fetching all that binary data on every page load. The experience is now fast regardless of attachment size or count.

## 🤖 OpenAI Playground: Per-Model reasoning\_effort Support

The Playground's `reasoning_effort` control now tracks OpenAI's actual per-model capability matrix. Models like `gpt-5.1` that support a `"none"` option show it; models that don't support reasoning effort have the control hidden automatically. Previously, the UI could get out of sync with what the backend supported.

## 🐍 Python SDK Improvements

Several reliability fixes and small improvements to the Python (and TypeScript) SDKs:

* **Streamer & drain reliability** — Fixed two edge-case bugs in the background message pipeline that could cause a small percentage of traces to be missed when using attachments or high message throughput
* **Auto-retry on rate limits** — `search_traces()` and `search_spans()` now automatically wait and retry on 429 responses instead of raising an error, so large bulk searches complete reliably under API rate limits
* **Evaluation task failures surfaced** — When an evaluation task raises an exception, the failure is now recorded on the experiment item (Python and TypeScript SDKs) instead of being silently discarded

## 🔧 Bug Fixes

* **Annotation queue reason persistence** — Two customer-reported bugs: the reason/comment field was being cleared when navigating between queue items, and the navigation order was incorrect. Both are fixed.
* **LLM failures visible in experiment traces** — When an LLM call failed during an experiment run, the trace appeared as an empty-output item with no indication of what went wrong. Structured error details are now written to the trace output so failures are diagnosable.

---

And much more! 👉 [See full commit log on GitHub](https://github.com/comet-ml/opik/compare/2.0.24...2.0.31)

*Releases*: `2.0.25`, `2.0.26`, `2.0.27`, `2.0.28`, `2.0.29`, `2.0.30`, `2.0.31`

## May 5, 2026

This is our biggest release yet! A fundamental rethink of how you build, debug, and improve AI agents with Opik. Three major new feature groups (Ollie, Test Suites, and the Agent Playground) work together to close the loop from observing a problem to shipping a fix, all without leaving the platform. Alongside them, we've reorganized everything around projects, redesigned the core trace experience, and rebuilt the navigation to match. Here's what's new:

## 🤖 Ollie & Opik Connect

Ollie is a powerful coding agent built into the Opik UI. It has full access to your project's traces and logs, and can analyze patterns across hundreds of interactions, diagnose issues, and take action to fix them, all without leaving the platform.

![Ollie agent interface inside Opik](/docs/opik/_fern-img/1d4a44055acb6ac2a29483a6c04a5cbee1cf77c43f6e01928e55796d1c9e0b98.webp)

**Highlights:**

* **Trace Analysis** - Analyze traces, spot patterns across interactions, and diagnose issues with full project context
* **Code Fixes via Opik Connect** - Link Ollie to your local codebase so it can implement fixes directly in your development code
* **Test Case Generation** - When Ollie fixes an issue, it automatically creates a new test case in your test suite to prevent regressions
* **UI Navigation** - Ollie can navigate the Opik UI, create filtered views, and take actions on your behalf
* **Opik Connect CLI** - Connect your codebase with `opik connect`, with support for `--workspace` and `--api-key` flags
* **Always Available** - Access Ollie from the project home page or as a persistent sidebar from any page in the product

👉 [Ollie Agent Documentation](https://www.comet.com/docs/opik/ollie)

## 🧪 Test Suites

Test Suites bring structured regression testing to agent development. Each suite has global rules that every test case must pass, plus item-level assertions for specific scenarios. Define rules in plain English for what your agent should and shouldn't do, and get clear pass/fail results when you run them.

![Test suite experiment results showing pass/fail per item with assertion details](/docs/opik/_fern-img/61092f479035e393df1d1868fb5503f2285aa6cd4896da09a1bdda89a71cfed0.webp)

**Highlights:**

* **Pass/Fail Assertions** - Define global rules and item-level assertions in plain English, no complex metric configurations needed
* **Multi-Provider LLM-as-Judge** - Assertions can use different LLM providers for evaluation, giving you flexibility in how test cases are judged
* **Assertion Reasons & Breakdown** - See exactly why each assertion passed or failed with detailed run-breakdown popovers
* **Add Traces as Test Cases** - Add production traces directly to your suite with assertions, so your suite grows naturally as you build and debug
* **Full SDK Support** - Python and TypeScript SDKs support creating suites, adding items, running experiments, and importing/exporting suites

👉 [Building Test Suites](https://www.comet.com/docs/opik/evaluation/advanced/building-test-suites)

## 🎮 Agent Playground & Agent Configurations

The Agent Playground connects to your agent so you can run it directly from the Opik UI. Experiment with different prompts, models, and parameters to see how your whole agent responds, without touching your code. Agent Configurations track and version the full set of prompts, models, and variables as a single unit, so you always know what combination worked.

![Agent Playground running an agent with configuration controls](/docs/opik/_fern-img/9560efa4f4ff14b3f397a13b51506252fd9977d6a01a76b88b8cdfb55ac613e6.webp)

**Highlights:**

* **Agent Playground** - Run your full agent from the Opik UI and test different configurations without changing your code
* **Agent Configurations** - Track and version prompts, models, and variables together as a single versioned unit
* **Blueprint Versioning** - Auto-increment naming, diff view for changes between versions, and auto-generated descriptions from config changes
* **Full SDK Support** - Python `AgentConfigManager` and TypeScript `AgentConfig` with Zod schema validation and blueprint caching

👉 [Agent Playground](https://www.comet.com/docs/opik/development/agent-playground)

## 🏗️ Project-Scoped Organization & UX Improvements

Projects now map directly to your agents. Test suites, experiments, optimizations, prompts, datasets, alerts, and dashboards are all scoped to the project, giving you a focused view of everything related to a single agent, paired with a redesigned navigation and trace experience.

![Redesigned unified Logs page with threads, traces, and spans in a single view](/docs/opik/_fern-img/509190201962bdf0a77fb8d701a1a9acd32dcc3522b3e1e885c4d2f537760905.webp)

**What's new:**

* **Redesigned Navigation** - New sidebar with workspace-level project selector and project-scoped routing across all pages
* **Unified Logs Page** - Threads, traces, and spans are now combined into a single redesigned Logs page with a cleaner layout and faster navigation between them
* **Redesigned Trace Details** - New tabbed layout with LLM message formatting, feedback scores section, and error callouts for faster issue identification
* **Project-Scoped APIs** - All endpoints now support `project_name` scoping for datasets, experiments, optimizations, prompts, alerts, and dashboards
* **KPI Cards** - New project-level metrics summary cards on the project home page

And much more! 👉 [See full commit log on GitHub](https://github.com/comet-ml/opik/compare/1.10.23...2.0.21)

*Releases*: `1.10.24` through `2.0.21`

## March 3, 2026

Here are the most relevant improvements we've made since the last release:

## 🦞 Native OpenClaw Observability with Opik

We've released `opik-openclaw`, a native OpenClaw plugin that gives you full-stack observability for your agents, powered by Opik. This brings enterprise-grade tracing, evaluation, and monitoring to the fastest-growing open-source agent framework.

**What you get:**

* **Full Trace Capture** - Every LLM call, tool execution, memory recall, context assembly, and agent delegation is logged with complete input/output pairs, token counts, latency, and cost
* **End-to-End Conversation Threading** - Trace a request from the initial message through multi-step reasoning, tool calls, and the final response, even when the agent chains across sub-agents or scheduled heartbeats
* **Real Cost Visibility** - Per-request, per-model cost breakdowns so you can see exactly where tokens are going and optimize accordingly
* **Automated Evaluation with LLM-as-a-Judge** - Set up hallucination detection, answer relevance, and context precision metrics that run automatically on your traces

Get started in two minutes: install the plugin with `openclaw plugins install @opik/opik-openclaw`, configure your API key, and traces start flowing immediately. Works with both Opik Cloud and self-hosted instances.

👉 [Visit the GitHub repository here](https://github.com/comet-ml/opik-openclaw)

## 🤖 Expanded Model & Provider Support

We've broadened the range of models and providers you can use across the platform, giving you more flexibility in how you build and evaluate your LLM applications.

**What's new:**

* **Gemini 3.1 Support** - Google's Gemini 3.1 is now available as a supported model across the platform
* **Claude Sonnet 4.6 as Default** - Claude Sonnet 4.6 is now the default Anthropic model, bringing improved performance out of the box
* **OpenRouter Native UX** - OpenRouter now has a much more native out-of-the-box experience in the Opik UI. `openrouter/free` is directly selectable, and `openrouter/*` route models including `/auto` are supported and prioritized in model selection
* **Updated Default Models** - The Python SDK has been updated to retire legacy gpt-4\* defaults in favor of more current models
* **OpenAI TTS Tracking** - You can now track OpenAI text-to-speech model calls (audio.speech) with full tracing support
* **OpenAI-Compatible Providers for LLM-as-a-Judge** - Use any OpenAI-compatible provider when running LLM-as-a-Judge evaluation metrics, giving you more flexibility in choosing your evaluation model

## 📦 SDK Improvements

We've continued to expand the capabilities of both the TypeScript and Python SDKs, making it easier to integrate Opik into your workflows programmatically.

**What's new:**

* **G-Eval Metric (TypeScript)** - The G-Eval evaluation metric is now available in the TypeScript SDK, enabling structured LLM-based evaluation directly from your TypeScript projects
* **Annotation Queue Support (TypeScript)** - Manage and interact with annotation queues programmatically from the TypeScript SDK
* **Thread Search (TypeScript)** - Search through conversation threads programmatically with the new `searchThreads` functionality in the TypeScript SDK
* **Offline Message Persistence (Python)** - When the Python SDK loses connectivity to the Opik server, telemetry messages are now persisted locally in a SQLite database and automatically replayed once the connection is restored — ensuring no data is lost during network outages
* **OTEL Integration Docs Expansion** - We've shipped a major expansion of our [OpenTelemetry integration documentation](/integrations/overview), including new pages and updated guidance for multiple frameworks and providers with emphasis on TypeScript

## 🚀 Optimization Studio & Optimizer SDK

We've made the Optimization Studio more powerful and flexible, with new metrics, persistence, and a major Optimizer SDK update.

**What's new:**

* **JSONPath Support & Numerical Similarity Metric** - The Optimization Studio now supports JSONPath expressions for extracting values from complex outputs, along with a new Numerical Similarity metric for comparing numeric results
* **Native MCP/Tool Optimization** - The v3.x Optimizer SDK now includes fully native [MCP and tool optimization](/development/optimization-runs/algorithms/tool_optimization) support, including support for remote MCP and improved tool-signature handling
* **Multi-Metric Optimization** - Multi-metric optimization is now working across span data with working examples for cost, speed, and quality tradeoff scenarios
* **Stronger Sampling & Agent Optimization** - Since the initial v3 SDK launch, we've added stronger sampling controls, full agent optimization including multi-prompt support, and finer prompt-control inside optimizer loops
* **Optimizer SDK 3.1.0** - The Optimizer SDK has been updated to version 3.1.0 with all of the above improvements and the retirement of legacy gpt-4\* model references

## ✨ Platform Features & UX Improvements

We've made several improvements to make your day-to-day workflow smoother and more intuitive.

**What's improved:**

* **Updated Default Columns** - Default columns across all tables have been refreshed to surface the most relevant information by default
* **Relative Time Format** - Time columns now display relative timestamps (e.g., "2 hours ago") for quicker at-a-glance understanding
* **Smart Threads Tab Default** - Projects with threads now automatically default to the Threads tab, getting you to the right view faster
* **Consistent Destructive Actions** - Destructive menu options are now visually unified with red text and separators for clearer intent
* **Feedback Score Precision** - Feedback scores are now rounded to 2 decimal places with full precision available on hover
* **Workspace Color Maps** - Configure workspace-level color maps for consistent visual styling across your projects
* **Image Attachments in Threads** - View image attachments directly within the thread view for better context when reviewing conversations
* **Bulk Tag Operations** - Add or remove tags in bulk across traces, spans, and other entities for faster organization
* **Inline Feedback Definition Creation** - Create new feedback definitions directly from the annotation queue form without leaving your workflow
* **Dataset Item Descriptions** - Dataset items now support a description field, making it easier to document and annotate your evaluation data
* **[Revamped MCP Server](https://github.com/comet-ml/opik-mcp)** - The Opik MCP server has been revamped to align with current MCP standards, with added support for remote MCP, improved auth behavior, and expanded native features including prompt and dataset workflows

## 🏷️ Prompt Version Tags

We've introduced prompt version tags, giving you a lightweight way to label and organize your prompt versions across the platform.

**What's new:**

* **Version Tags in Comparison View** - Easily see and compare tagged prompt versions side by side in the prompt comparison view
* **Python SDK Support** - Create, manage, and retrieve prompt version tags programmatically from the Python SDK
* **Retrieve Prompts by Commits** - A new API endpoint lets you retrieve prompts by their commit references, enabling tighter integration with your version control workflow

👉 [Prompt Version Tags Documentation](/development/prompt-library/getting-started#choosing-a-version)

---

And much more! 👉 [See full commit log on GitHub](https://github.com/comet-ml/opik/compare/1.10.10...1.10.23)

*Releases*: `1.10.11`, `1.10.12`, `1.10.13`, `1.10.14`, `1.10.15`, `1.10.16`, `1.10.17`, `1.10.18`, `1.10.19`, `1.10.20`, `1.10.21`, `1.10.22`, `1.10.23`

_Showing the 20 most recent of 69 entries. Append `/llms.txt` to the changelog URL for the complete index._