-
I Built a RAG Pipeline for F1 Team Radio, Then Made It Grade Itself
I wanted to see if I could build a RAG system that would output interesting and accurate F1 race weekend…
-
How We Optimized Opik’s MCP Server for Cost & Performance
Like a lot of engineering teams, earlier this year we found ourselves hitting limits on AI token spend, trying to…
-
Announcing the Opik Claude Code Plugin: Automatically Configure Observability for Complex Agentic Systems
Agent observability shouldn’t be a side project. But in practice, it is. When teams have to choose between shipping features…
-
SelfCheckGPT for LLM Evaluation
Detecting hallucinations in language models is challenging. There are three general approaches: The problem with many LLM-as-a-Judge techniques is that…
-
LLM Juries for Evaluation
Evaluating the correctness of generated responses is an inherently challenging task. LLM-as-a-Judge evaluators have gained popularity for their ability to…
-
A Simple Recipe for LLM Observability
So, you’re building an AI application on top of an LLM, and you’re planning on setting it live in production.…
-
G-Eval for LLM Evaluation
LLM-as-a-judge evaluators have gained widespread adoption due to their flexibility, scalability, and close alignment with human judgment. They excel at…
-
Build Multi-Index Advanced RAG Apps
Welcome to Lesson 12 of 12 in our free course series, LLM Twin: Building Your Production-Ready AI Replica. You’ll learn…
-
Build a scalable RAG ingestion pipeline using 74.3% less code
Welcome to Lesson 11 of 12 in our free course series, LLM Twin: Building Your Production-Ready AI Replica. You’ll learn…
-
BERTScore For LLM Evaluation
Introduction BERTScore represents a pivotal shift in LLM evaluation, moving beyond traditional heuristic-based metrics like BLEU and ROUGE to a…














