GuideIntermediate

Evaluation Frameworks Guide

Master the evaluation of AI systems — from metrics for chat, RAG, and agents to LLM-as-judge, RAGAS, TruLens, and production evaluation pipelines. Learn to build golden datasets, implement domain-specific evaluation for chatbots, RAG pipelines, and AI agents, automate evaluation with calibrated LLM judges, and deploy continuous monitoring with CI/CD integration and regression detection.

64
lessons
8
modules
English · Spanish
available in
Yes
certificate
Included in the Club
access
NIEVA

Outcomes

What you'll be able to do

  • Understand why evaluation is the most critical gap in AI Engineering and adopt evaluation-driven development
  • Master metrics taxonomy: reference-based (BLEU, ROUGE, BERTScore), reference-free, semantic, and custom metrics
  • Design professional golden datasets with annotation strategies, synthetic data, and versioning
  • Evaluate chatbots with response quality metrics, safety checks, multi-turn evaluation, and TruLens
  • Evaluate RAG pipelines with RAGAS: faithfulness, answer relevancy, context precision and recall
  • Evaluate AI agents with trajectory evaluation, tool call accuracy, and task completion metrics
  • Implement LLM-as-judge with calibrated prompts, bias mitigation, pairwise comparison, and multi-judge consensus
  • Build production evaluation pipelines with CI/CD gates, continuous monitoring, regression detection, and alerting

Before you start

What you need to bring

It's for you if...

  • AI Engineers who build chat, RAG, or agent systems and need to measure their quality rigorously
  • Developers deploying AI systems to production without knowing if they actually work well
  • Engineers who need to implement evaluation in their company's CI/CD pipeline
  • Teams transitioning from "vibes-based evaluation" to data-driven, automated quality metrics
  • Professionals who want to understand LLM-as-judge, RAGAS, and TruLens for production use

Requirements and materials

  • Completed Advanced RAG Techniques Guide (#8) or experience building RAG pipelines
  • Completed LangChain & LangGraph Guide (#9) or equivalent framework experience
  • Completed Building AI Agents Guide (#11) or experience building agents with tool use
  • Intermediate Python (functions, classes, async, Pydantic basics)
  • Experience with at least one AI system in development or production
  • At least one LLM API key (OpenAI recommended)
  • Python 3.11+ installed

Content

The syllabus, module by module

Open any of them to see its lessons.

  • 1. Introduction: The Problem Nobody Wants to Solve
  • 2. The Evaluation Gap in AI Engineering
  • 3. Types of Evaluation: Offline, Online, Human-in-the-Loop
  • 4. Evaluation vs Testing vs Monitoring
  • 5. The Cost of Not Evaluating
  • 6. Evaluation-Driven Development
  • 7. From "Vibes" to Metrics: The Mindset Shift
  • 8. Project: First End-to-End Evaluation Pipeline

  • 1. Introduction: Not All Metrics Measure the Same Thing
  • 2. Reference-Based Metrics: BLEU, ROUGE, METEOR
  • 3. Semantic Metrics: BERTScore, Cosine Similarity, STS
  • 4. Reference-Free Metrics: Evaluating Without a Correct Answer
  • 5. Classification Metrics Applied to AI
  • 6. Text Generation Metrics
  • 7. Designing Custom Metrics
  • 8. Project: Metrics Comparator

  • 1. Introduction: Without Evaluation Data, There Is No Evaluation
  • 2. What Is a Golden Dataset?
  • 3. Designing Effective Test Cases
  • 4. Annotation Strategies
  • 5. Synthetic Data for Evaluation
  • 6. Versioning and Maintenance of Golden Datasets
  • 7. Test Suites as a Quality Contract
  • 8. Project: Professional Golden Dataset

  • 1. Introduction: Evaluating Conversations Is Hard
  • 2. Response Quality Metrics
  • 3. Coherence and Contextual Relevance
  • 4. Safety, Toxicity, and Guardrails
  • 5. Multi-Turn Evaluation
  • 6. User Satisfaction Proxies
  • 7. TruLens for Chat Evaluation
  • 8. Evolving Project: Chatbot Evaluation Suite

  • 1. Introduction: RAG Without Evaluation is a Black Box
  • 2. RAG Failure Modes
  • 3. Faithfulness: Is the Answer Based on the Context?
  • 4. Answer Relevancy: Does the Answer Answer the Question?
  • 5. Context Precision and Context Recall: Evaluating the Retriever
  • 6. RAGAS Framework Deep Dive
  • 7. End-to-End RAG Evaluation
  • 8. Evolving Project: RAG Evaluation Suite

  • 1. Introduction: Agents Are the Hardest Systems to Evaluate
  • 2. Agent Failure Modes
  • 3. Trajectory Evaluation
  • 4. Tool Call Accuracy: Precision, Recall, and Argument Correctness
  • 5. Task Completion and Reasoning Quality
  • 6. Evaluation of Multi-Agent Systems
  • 7. Agent Benchmarks and Datasets
  • 8. Evolving Project: Agent Evaluation Suite

  • 1. Introduction: Using an LLM to Evaluate Another LLM
  • 2. What LLM-as-Judge Is
  • 3. Judge Prompt Engineering
  • 4. Single-Point Grading vs Pairwise Comparison
  • 5. Calibration and Bias in LLM Judges
  • 6. Multi-Judge and Consensus
  • 7. Automated Evaluation Pipelines
  • 8. Evolving Project: LLM-as-Judge Pipeline

  • 1. Introduction: Evaluation Is Not an Event, It's a Process
  • 2. Evaluation in CI/CD
  • 3. Continuous Quality Monitoring
  • 4. Regression Detection and Alerting
  • 5. A/B Testing for AI Systems
  • 6. Observability with LangSmith and TruLens
  • 7. Production Evaluation Architecture
  • 8. Final Project: Production AI Evaluation Platform

Where it fits

This guide is part of something bigger

It's studied inside these programs, with support and dates.

Common questions

What people usually ask

Start whenever you like

Reviews

What students say

These reviews are from enrolled students who completed at least 50% of the course. We moderate reviews only on content grounds (spam, offensive language, personal data), never for being critical or negative.

No approved reviews yet.

Be the first to share your experience!