> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://www.comet.com/docs/opik/evaluation/llms.txt. # Evaluation > Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards. ## Docs - [Evaluation Overview](https://www.comet.com/docs/opik/evaluation/overview.md): Evaluate your LLM applications with Opik using assertion-based Test Suites or dataset-driven metric scoring - [Getting started with Evaluation](https://www.comet.com/docs/opik/evaluation/getting-started.md): Get started evaluating your LLM application with Opik using Test Suites or dataset-driven metrics - [Evaluation Concepts](https://www.comet.com/docs/opik/evaluation/concepts.md): Understand the two evaluation approaches in Opik — Test Suites with assertions and Datasets with metrics - [Building Test Suites](https://www.comet.com/docs/opik/evaluation/advanced/building-test-suites.md): Create and manage Test Suites using the SDK, UI, or Ollie to build regression tests from real production failures - [Evaluate your agent](https://www.comet.com/docs/opik/evaluation/advanced/evaluate_your_llm.md): Evaluate your LLM applications confidently. Learn the five steps to assess complex LLM chains or agents effectively. - [Resume an interrupted evaluation](https://www.comet.com/docs/opik/evaluation/advanced/resume_evaluations.md): Continue a long-running Opik evaluation after a crash, network blip, or Ctrl-C — replaying only the runs that didn't finish. - [Manage datasets](https://www.comet.com/docs/opik/evaluation/advanced/manage_datasets.md): Evaluate your LLM using datasets. Learn to create and manage them via Python SDK, TypeScript SDK, or the Traces table. - [Evaluate agent trajectories](https://www.comet.com/docs/opik/evaluation/advanced/evaluate_agent_trajectory.md): Evaluate agent trajectories to optimize tool selection and reasoning paths, ensuring efficient agent behavior before production with Opik. - [Evaluate multi-turn agents](https://www.comet.com/docs/opik/evaluation/advanced/evaluate_multi_turn_agents.md): Learn to evaluate multi-turn agents using simulation techniques to enhance chatbot performance and improve user interactions. - [Annotation Queues](https://www.comet.com/docs/opik/evaluation/advanced/annotation_queues.md): Optimize your AI projects by enabling SMEs to efficiently review and annotate outputs using Opik's intuitive Annotation Queues feature. - [Manually logging experiments](https://www.comet.com/docs/opik/evaluation/advanced/log_experiments_with_rest_api.md): Evaluate your LLM application by logging pre-computed experiments and boosting confidence in performance with this detailed guide. - [Exporting experiment results](https://www.comet.com/docs/opik/evaluation/advanced/export_experiment_results.md): Export the full results of an experiment to CSV with the Opik Python SDK, matching the columns of the comparison page without its row cap. - [Overview](https://www.comet.com/docs/opik/evaluation/metrics/overview.md): Describes all the built-in evaluation metrics provided by Opik - [Heuristic metrics](https://www.comet.com/docs/opik/evaluation/metrics/heuristic_metrics.md): Describes all the built-in heuristic metrics provided by Opik - [Hallucination](https://www.comet.com/docs/opik/evaluation/metrics/hallucination.md): Describes the Hallucination metric - [LLM Juries](https://www.comet.com/docs/opik/evaluation/metrics/llm_juries.md): Combine multiple judges into an ensemble with LLMJuriesJudge - [G-Eval](https://www.comet.com/docs/opik/evaluation/metrics/g_eval.md): Describes Opik's built-in G-Eval metric which is a task agnostic LLM as a Judge metric - [Conversation-level GEval Metrics](https://www.comet.com/docs/opik/evaluation/metrics/g_eval_conversation_metrics.md): How to run GEval-based judges on full conversations - [Compliance risk](https://www.comet.com/docs/opik/evaluation/metrics/compliance_risk.md): Flag non-compliant or high-risk assistant replies with ComplianceRiskJudge - [Prompt uncertainty](https://www.comet.com/docs/opik/evaluation/metrics/prompt_diagnostics.md): Estimate prompt ambiguity with PromptUncertaintyJudge - [Moderation](https://www.comet.com/docs/opik/evaluation/metrics/moderation.md): Describes the Moderation metric - [Meaning Match](https://www.comet.com/docs/opik/evaluation/metrics/meaning_match.md): Describes the Meaning Match metric - [Usefulness](https://www.comet.com/docs/opik/evaluation/metrics/usefulness.md): Describes the Usefulness metric - [Summarization consistency](https://www.comet.com/docs/opik/evaluation/metrics/summarization_consistency.md): Ensure autogenerated summaries stay faithful to the source content - [Summarization coherence](https://www.comet.com/docs/opik/evaluation/metrics/summarization_coherence.md): Rate how readable and well-structured a summary is with SummarizationCoherenceJudge - [Dialogue helpfulness](https://www.comet.com/docs/opik/evaluation/metrics/dialogue_helpfulness.md): Measure how helpful an assistant reply is within a dialogue - [Answer relevance](https://www.comet.com/docs/opik/evaluation/metrics/answer_relevance.md): Describes the Answer Relevance metric - [Context precision](https://www.comet.com/docs/opik/evaluation/metrics/context_precision.md): Describes the Context Precision metric - [Context recall](https://www.comet.com/docs/opik/evaluation/metrics/context_recall.md): Describes the Context Recall metric - [Trajectory accuracy](https://www.comet.com/docs/opik/evaluation/metrics/trajectory_accuracy.md): Score whether an agent followed the expected action path - [Agent task completion](https://www.comet.com/docs/opik/evaluation/metrics/agent_task_completion.md): Verify whether an agent fulfilled its assigned objective - [Agent tool correctness](https://www.comet.com/docs/opik/evaluation/metrics/agent_tool_correctness.md): Evaluate whether an agent invoked and interpreted tools correctly - [Conversational metrics](https://www.comet.com/docs/opik/evaluation/metrics/conversation_threads_metrics.md): Describes metrics related to scoring the conversational threads - [Custom model](https://www.comet.com/docs/opik/evaluation/metrics/custom_model.md): Describes how to use a custom model for Opik's built-in LLM as a Judge metrics - [Advanced configuration](https://www.comet.com/docs/opik/evaluation/metrics/advanced_configuration.md): Fine-tune Opik metrics with async scoring, evaluator temperatures, and logprob handling - [Custom metric](https://www.comet.com/docs/opik/evaluation/metrics/custom_metric.md): Describes how to create your own metric to use with Opik's evaluation platform - [Custom conversation metric](https://www.comet.com/docs/opik/evaluation/metrics/custom_conversation_metric.md): Learn how to create custom metrics for evaluating multi-turn conversations - [Structured Output Compliance](https://www.comet.com/docs/opik/evaluation/metrics/structure_output_compliance.md): Describes the StructuredOutputCompliance metric - [Task span metrics](https://www.comet.com/docs/opik/evaluation/metrics/task_span_metrics.md): Learn how to create task span metrics for evaluating the detailed execution information of your LLM tasks - [Best practices for evaluating agents](https://www.comet.com/docs/opik/evaluation/evaluate_agents.md): Learn best practices for evaluating AI agents, ensuring reliability and scalability throughout their lifecycle with Opik's guidance. - [Evaluate threads](https://www.comet.com/docs/opik/evaluation/evaluate_threads.md): Evaluate and optimize conversation threads in Opik using the evaluate_threads function in the Python SDK for enhanced multi-turn conversations. - [Cookbook - Evaluate hallucination metric](https://www.comet.com/docs/opik/evaluation/evaluate_hallucination_metric.md): Learn to evaluate the Hallucination metric using the LLM Evaluation SDK and improve your model's performance with Opik. - [Cookbook - Evaluate moderation metric](https://www.comet.com/docs/opik/evaluation/evaluate_moderation_metric.md): Evaluate the Moderation metric in the LLM Evaluation SDK to enhance your moderation capabilities using Opik's platform.