-
Jev vs. LLM-as-a-Judge for AI Evals
A model that answers yes/no questions instead of writing text came out 3.5× cheaper and 3.8× faster than our gpt-4o-mini…
-
Capturing a 400-Turn Claude Code Session in a Single Image
We ran a small model over thousands of Claude Code session and discovered how people really use coding agents. Comet…
-
Diffusion Language Models, From Scratch to Production
“Language model” used to be a very general term for statistical models trained on language tasks. It included everything from…
-
Engineering Insights: How Internal Optimizations Led to Comet Cost Intelligence
AI budgets are no longer growing unchecked. Across the industry, engineering teams are being asked to do more with less,…
-
Understanding Your Claude Code Spend: What’s Actually Driving the Cost
I’ve been spending time looking at how teams are actually using Claude Code, and one thing keeps coming up: most…
-
Introducing the Opik Agent Playground
In the early stages of agent development, you make big changes to your agent’s code: designing the architecture, integrating tools,…
-
Introducing Ollie: Auto-Fix Your Agent’s Codebase
In standard software engineering, developers use proven, repeatable workflows to develop, test, debug, and update software products. They use intelligent…
-
Introducing Opik Test Suites: Straightforward Unit & Regression Testing for AI Agents
One of the biggest challenges when it comes to agent development is quality. It’s getting easier every day to spin…
-
How Contributing to Open Source Projects Helped Me Build My Dream Career in AI
6 years ago, I decided to open-source my Python code for a personal project I was working on, which led…
-
EU AI Act Regulation Compliance with Comet
On March 13, 2024, the European Parliament passed the EU AI Act to establish a common regulatory and legal framework…















