Module 5: RAG Evaluation with RAGAS

3. Faithfulness: Is the Answer Based on the Context?

Overview

Faithfulness measures the proportion of claims in an LLM's answer that are actually supported by the retrieved context. It's the most critical metric for detecting hallucinations in RAG systems. When a user asks something and the retriever returns relevant documents, the LLM can do two things: base itself faithfully on those documents, or make up information that sounds plausible but isn't in the context. Faithfulness quantifies exactly that — a score of 1.0 means every claim in the answer has support in the retrieved documents; a score of 0.0 means nothing the model said comes from the context.

The approach RAGAS uses to compute faithfulness is elegant and splits into two steps: first, claim decomposition — breaking the answer down into individual atomic claims; second, Natural Language Inference (NLI) — verifying whether the context supports, contradicts, or is neutral about each claim. This two-step approach is superior to asking an LLM to "rate from 0 to 1 how grounded the answer is in the context" because it makes the reasoning explicit and auditable. You can see exactly which claims were extracted and which ones failed verification.

In this lesson you'll understand the concept of faithfulness from its foundations, manually implement each step of the pipeline (claim decomposition and NLI) using the OpenAI API, and finally see how RAGAS wraps all of this in a single metric. The three levels of understanding — conceptual, manual implementation, framework — give you the ability to diagnose problems when scores don't make sense, to customize the metric for your domain, and to trust the results because you understand what they're measuring.


Setup

Before starting, install the necessary dependencies:

pip install openai ragas datasets
from openai import OpenAI
import json

client = OpenAI()

The whole manual implementation uses gpt-4o-mini as the evaluator LLM. RAGAS uses its own LLM configuration that you can customize, but by default it also uses OpenAI models.


The concept: what does "faithful" mean?

In the context of RAG, faithful doesn't mean "correct" — it means "based on what was retrieved." An answer can be faithful but incorrect (if the retrieved context was incorrect), or it can be correct but unfaithful (if the LLM knows the correct answer but generated it without basing itself on the context).

Question: "What are the business hours?"
Context: "Business hours are 9:00 to 18:00, Monday to Friday."

Answer A: "Hours are 9 to 18, Monday to Friday."
→ Faithful: all claims are in the context

Answer B: "Hours are 9 to 18, Monday to Friday. There's also Saturday service from 10 to 14."
→ Partially faithful: the first part yes, the second is made up

Answer C: "Hours are 8 to 17 every day."
→ Not faithful: contradicts the context

Faithfulness isn't a judgment about the overall quality of the answer. It's specifically about whether the claims come from the provided context. That's why in RAGAS, faithfulness is complemented with answer_relevancy (does it answer the question?) and context_precision/recall (was the context good?).

Faithfulness vs other failure modes

In the previous lesson (02) you saw the taxonomy of RAG failures. Faithfulness directly detects the hallucination failure:

RAG failureMetric that detects itRelation to faithfulness
HallucinationFaithfulnessDirect detection: unsupported claims
IrrelevanceAnswer relevancyOrthogonal: can be faithful but irrelevant
Wrong contextContext precisionIndirect: high faithfulness + wrong answer → bad context
Missing contextContext recallIndirect: low faithfulness may be due to insufficient context

Step 1: Claim decomposition

The first step to computing faithfulness is decomposing the answer into atomic claims — claims. A claim is a unit of information that can be independently true or false.

What is an atomic claim?

Answer: "Python was created by Guido van Rossum in 1991 and is an interpreted language."

Extracted claims:
1. "Python was created by Guido van Rossum"
2. "Python was created in 1991"
3. "Python is an interpreted language"

Each claim contains a single piece of verifiable information. The original answer combines three facts in one sentence — claim decomposition separates them so each can be independently verified against the context.

Why not verify the complete answer

You might think: "why not ask the LLM directly whether the answer is supported by the context?". That approach has two problems:

  1. Granularity: If an answer has 5 claims and 4 are supported, direct verification could return "yes" (ignoring the false claim) or "no" (ignoring the 4 correct ones). You don't know how much of the answer is trustworthy.
  2. Auditability: With claim decomposition, you can see exactly which claim failed. Does the LLM always hallucinate dates? Does it make up names? Patterns emerge when you see the individual claims that fail.

Manual implementation of claim decomposition

def extract_claims(answer: str, client: OpenAI) -> list[str]:
    response = client.chat.completions.create(
        model="gpt-4o-mini",
        temperature=0,
        messages=[
            {
                "role": "system",
                "content": (
                    "You are a factual claim extractor. "
                    "Decompose the answer into individual atomic claims. "
                    "Each claim must contain a single piece of verifiable information. "
                    "Ignore connectors, subjective opinions and courtesy phrases. "
                    "Return a JSON array of strings."
                ),
            },
            {
                "role": "user",
                "content": f"Answer: {answer}",
            },
        ],
        response_format={"type": "json_object"},
    )
    result = json.loads(response.choices[0].message.content)
    return result.get("claims", result.get("afirmaciones", []))

Testing the extraction

answer = (
    "FastAPI was created by Sebastián Ramírez in 2018. "
    "It's based on Starlette for the web part and Pydantic for data validation. "
    "It's one of the fastest Python frameworks, comparable to Node.js and Go."
)

claims = extract_claims(answer, client)
for i, claim in enumerate(claims, 1):
    print(f"  Claim {i}: {claim}")

Expected output:

  Claim 1: FastAPI was created by Sebastián Ramírez
  Claim 2: FastAPI was created in 2018
  Claim 3: FastAPI is based on Starlette for the web part
  Claim 4: FastAPI is based on Pydantic for data validation
  Claim 5: FastAPI is one of the fastest Python frameworks
  Claim 6: FastAPI is comparable in speed to Node.js and Go

Notice how a 2-sentence answer generates 6 claims. The sentence about Starlette and Pydantic contains two independent facts that get separated. The comparison with Node.js and Go is a verifiable claim distinct from "it's fast."

Edge cases in claim decomposition

Not all answers are equal:

  • "I didn't find information": Has no verifiable factual claims. Faithfulness should be 1.0 (there are no false claims) or N/A.
  • "According to the document... it could vary": Hedges ("it could vary") are hard to classify as verifiable claims.
  • "Python is great": A subjective opinion, not a factual claim. A good extractor ignores it.

The extractor prompt should handle these cases. The instruction "ignore subjective opinions" helps, but in practice you'll need to iterate on the prompt for your specific domain.


Step 2: Natural Language Inference (NLI)

With the claims extracted, the second step is verifying each one against the context. This is Natural Language Inference — determining whether a premise (the context) entails, contradicts, or is neutral about a hypothesis (the claim).

The three NLI relations

Context (premise): "Business hours are 9:00 to 18:00, Monday to Friday."

Claim: "Hours are 9 to 18"           → ENTAILMENT (the context supports it)
Claim: "Hours are 8 to 17"           → CONTRADICTION (the context contradicts it)
Claim: "There is Saturday service"   → NEUTRAL (the context says nothing about it)

For faithfulness, both CONTRADICTION and NEUTRAL count as "not supported." The claim must have an ENTAILMENT relation with the context to count as faithful. This is strict on purpose: if the context doesn't mention something, the LLM shouldn't claim it.

Manual implementation of NLI

def verify_claim(claim: str, context: str, client: OpenAI) -> dict:
    response = client.chat.completions.create(
        model="gpt-4o-mini",
        temperature=0,
        messages=[
            {
                "role": "system",
                "content": (
                    "You are a claim verifier. "
                    "Determine whether the context SUPPORTS the claim. "
                    "Respond with a JSON: "
                    '{\"verdict\": \"supported\" | \"not_supported\", '
                    '\"reason\": \"brief explanation\"}'
                ),
            },
            {
                "role": "user",
                "content": f"Context: {context}\n\nClaim: {claim}",
            },
        ],
        response_format={"type": "json_object"},
    )
    return json.loads(response.choices[0].message.content)

Testing the verification

context = (
    "FastAPI is a modern web framework for Python, created by Sebastián Ramírez. "
    "It's based on Starlette and Pydantic. "
    "FastAPI generates automatic documentation with OpenAPI."
)

test_claims = [
    "FastAPI was created by Sebastián Ramírez",
    "FastAPI was created in 2018",
    "FastAPI is based on Starlette",
    "FastAPI is the most popular Python framework",
    "FastAPI generates automatic documentation",
]

for claim in test_claims:
    result = verify_claim(claim, context, client)
    status = "✓" if result["verdict"] == "supported" else "✗"
    print(f"  {status} {claim}")
    print(f"    → {result['verdict']}: {result['reason']}")

The claim about 2018 is a fact that could be true, but the context doesn't mention it — therefore it's not faithful. The claim about "most popular" is an LLM invention that the context doesn't support.


Complete pipeline: manual faithfulness

Now we combine the two steps into a complete pipeline that computes the faithfulness score:

def calculate_faithfulness(
    answer: str,
    context: str,
    client: OpenAI,
    verbose: bool = False,
) -> dict:
    claims = extract_claims(answer, client)

    if not claims:
        return {
            "score": 1.0,
            "total_claims": 0,
            "supported_claims": 0,
            "details": [],
        }

    details = []
    supported = 0

    for claim in claims:
        verification = verify_claim(claim, context, client)
        is_supported = verification["verdict"] == "supported"
        if is_supported:
            supported += 1
        details.append({
            "claim": claim,
            "supported": is_supported,
            "reason": verification.get("reason", ""),
        })

    score = supported / len(claims)

    if verbose:
        print(f"FAITHFULNESS: {supported}/{len(claims)} = {score:.2f}")
        for d in details:
            status = "✓" if d["supported"] else "✗"
            print(f"  {status} {d['claim']}")

    return {
        "score": score,
        "total_claims": len(claims),
        "supported_claims": supported,
        "details": details,
    }

Complete example

context = (
    "PostgreSQL is an open source relational database system. "
    "It was originally developed at UC Berkeley. "
    "It supports JSON, full-text search and extensions like PostGIS. "
    "It's ACID-compliant and supports transactions."
)

answer = (
    "PostgreSQL is an open source relational database developed at UC Berkeley. "
    "It supports JSON and full-text search. "
    "It's the most used database in the world. "
    "It has support for ACID transactions."
)

result = calculate_faithfulness(answer, context, client, verbose=True)

Expected output:

FAITHFULNESS: 4/5 = 0.80
  ✓ PostgreSQL is an open source relational database
  ✓ PostgreSQL was developed at UC Berkeley
  ✓ PostgreSQL supports JSON and full-text search
  ✗ PostgreSQL is the most used database in the world
  ✓ PostgreSQL has support for ACID transactions

Faithfulness = 4/5 = 0.80. The LLM added "the most used in the world" — a claim that may be true but isn't in the context. In a RAG system, this is a mild hallucination: the LLM injected its own knowledge instead of limiting itself to the provided context.


RAGAS faithfulness

Everything you implemented manually — claim decomposition + NLI — is exactly what RAGAS does internally. The advantage of RAGAS is that the implementation is optimized, tested, and consistent across projects.

Basic usage

from ragas import evaluate
from ragas.metrics import faithfulness
from datasets import Dataset

data = Dataset.from_dict({
    "question": [
        "What is Docker?",
        "What is Redis?",
        "What is FastAPI?",
    ],
    "answer": [
        "Docker is a container platform that packages applications.",
        "Redis is an in-memory database. It's the fastest in the world.",
        "FastAPI is a web framework based on Django.",
    ],
    "contexts": [
        ["Docker is a containerization platform for packaging applications with dependencies."],
        ["Redis is an in-memory data store that supports strings, hashes, lists and sets."],
        ["FastAPI is a web framework based on Starlette and Pydantic for APIs with Python."],
    ],
})

result = evaluate(data, metrics=[faithfulness])
print(f"Average faithfulness: {result['faithfulness']:.4f}")

df = result.to_pandas()
print(df[["question", "faithfulness"]])

What does RAGAS do internally?

The RAGAS pipeline for faithfulness is conceptually identical to your manual implementation:

1. claim_decomposition(answer) → list of atomic claims
2. For each claim:
   nli_verification(claim, contexts) → supported | not_supported
3. faithfulness = count(supported) / count(total_claims)

The difference is in the details: RAGAS uses optimized prompts for claim extraction, handles edge cases like empty answers, and supports multiple context chunks — it verifies each claim against each chunk individually and marks it as "supported" if at least one chunk supports it.

When to use each approach

In practice, the recommendation is: implement manually to understand, use RAGAS for production. Your manual implementation teaches you what faithfulness measures and how it fails. RAGAS gives you the version ready to integrate into CI/CD pipelines. If your domain has special claims (numeric with tolerance, temporal with ranges), customize the prompts of extract_claims and verify_claim before migrating to RAGAS.


The paraphrasing test

A good faithfulness system must handle paraphrasing — when the answer says the same thing as the context with different words. This is the definitive test that separates good implementations from bad ones:

context = "Customer service operates from 9:00 to 18:00, Monday to Friday."

paraphrased_answers = [
    "They're available from 9am to 6pm on weekdays.",
    "The hours are nine in the morning to six in the evening, business days.",
    "You can reach them during office hours, 9 to 18, Mon-Fri.",
]

for answer in paraphrased_answers:
    result = calculate_faithfulness(answer, context, client)
    print(f"  Score: {result['score']:.2f} | {answer}")

All these answers are faithful to the context — they say the same thing with different words and number formats. A good evaluator should give high scores for all of them. If your implementation fails here, improve the NLI verification prompt with explicit instructions that paraphrasing and format changes count as "supported."


Connection to the module project

In the module 5 project (RAG Evaluation Suite), you'll build a complete RAG evaluation pipeline that includes faithfulness as one of the core metrics. Your pipeline will receive triplets (question, context, answer) from a real RAG system and compute faithfulness along with answer_relevancy, context_precision and context_recall.

This lesson's manual implementation gives you two advantages for the project:

  1. Debugging: When RAGAS reports a low faithfulness, you'll be able to use your manual pipeline to see exactly which claims failed and understand whether it's a retriever problem (insufficient context) or a generator problem (hallucination).
  2. Customization: If your domain has special requirements (numeric claims with tolerance, temporal claims with acceptable ranges), you can modify the verify_claim function with specific rules that RAGAS doesn't support out of the box.

In the project, faithfulness is the metric that answers the most important question: can I trust what my RAG system says?


Troubleshooting

Problem 1: Low faithfulness but the answer looks correct

Symptom: The faithfulness score is 0.4-0.6, but when you read the answer, everything it says seems true.

Cause: The context doesn't contain the information explicitly, even if it's correct. The LLM is using its internal knowledge instead of the provided context. Another cause: the evaluator doesn't recognize paraphrasing.

Solution: Review the individual claims marked as "not supported." If the context really doesn't mention that information, it's a legitimate hallucination. If it's a paraphrasing problem, improve the verification prompt: "Consider it supported if the meaning is equivalent, even with different words." In RAGAS, this problem is less common because the internal prompts already handle paraphrasing.

Problem 2: Faithfulness is 1.0 but the answer is useless

Symptom: Perfect score, but the answer says "I don't have enough information" or repeats the context verbatim.

Cause: Faithfulness measures whether the claims are supported, not whether the answer is useful. An empty answer has no factual claims to verify, so it gets 1.0 trivially.

Solution: Faithfulness should never be used as the sole metric. Combine it with answer_relevancy (lesson 04) to verify that the answer actually addresses the question. A good RAG system has high faithfulness AND high answer_relevancy.

Problem 3: Claim extraction is inconsistent

Symptom: The same answer generates 3 claims in one run and 5 in another. The score varies between runs.

Cause: The extractor LLM isn't deterministic — even with temperature=0, there's variation in how it decomposes complex sentences.

Solution: Add few-shot examples to the extractor prompt to establish the granularity level you expect. For critical evaluations, run the extraction 3 times and average. RAGAS mitigates this with well-calibrated internal prompts.

Problem 4: Multiple context chunks cause confusion

Symptom: When the context is multiple documents, a claim supported by chunk 3 is marked as "not supported" because the verifier only looked at chunk 1.

Cause: If you concatenate all the chunks into one long string, the LLM can lose information in the middle ("lost in the middle"). If you verify separately, a claim that requires information from two chunks may fail in both.

Solution: Verify each claim against each chunk individually and mark it as "supported" if at least one chunk supports it. This is the RAGAS approach. For claims that combine information from multiple chunks, concatenate only the top-2 or top-3 by similarity with the claim.

Problem 5: High costs from LLM calls

Symptom: Evaluating 100 examples generates hundreds of API calls — one to extract claims, plus one per claim for verification.

Cause: The two-step pipeline multiplies the calls: if each answer has 5 claims on average, 100 examples = 100 extractions + 500 verifications = 600 calls.

Solution: Use gpt-4o-mini instead of gpt-4o to reduce costs ~10x. Batch the verifications — send all the claims of an answer in a single call and ask for a JSON array of verdicts. For large datasets, evaluate a representative sample (50-100 examples) instead of the full dataset.


Exercises

Exercise 1: Manual claim decomposition

Given this answer, manually decompose it into atomic claims BEFORE using the LLM. Then compare your decomposition with the LLM's.

Answer: "SQLAlchemy is an ORM for Python that supports PostgreSQL, MySQL and SQLite. It was created in 2005 and is the most popular ORM in the Python ecosystem. It uses a pattern called Unit of Work to handle transactions."

Context: "SQLAlchemy is an SQL toolkit and ORM for Python. It supports multiple databases including PostgreSQL, MySQL, SQLite and Oracle. It was created by Mike Bayer. SQLAlchemy implements the Unit of Work pattern and the Identity Map pattern."

View solution

Manually extracted claims:

  1. SQLAlchemy is an ORM for Python → ✓ Supported
  2. SQLAlchemy supports PostgreSQL → ✓ Supported
  3. SQLAlchemy supports MySQL → ✓ Supported
  4. SQLAlchemy supports SQLite → ✓ Supported
  5. SQLAlchemy was created in 2005 → ✗ Not supported (the context doesn't mention a date)
  6. SQLAlchemy is the most popular ORM in the Python ecosystem → ✗ Not supported
  7. SQLAlchemy uses a pattern called Unit of Work → ✓ Supported
  8. Unit of Work is used to handle transactions → ✗ Not supported (the context doesn't specify what it's for)

Faithfulness = 5/8 = 0.625

from openai import OpenAI
import json

client = OpenAI()

answer = (
    "SQLAlchemy is an ORM for Python that supports PostgreSQL, MySQL and SQLite. "
    "It was created in 2005 and is the most popular ORM in the Python ecosystem. "
    "It uses a pattern called Unit of Work to handle transactions."
)

context = (
    "SQLAlchemy is an SQL toolkit and ORM for Python. "
    "It supports multiple databases including PostgreSQL, MySQL, SQLite and Oracle. "
    "It was created by Mike Bayer. "
    "SQLAlchemy implements the Unit of Work pattern and the Identity Map pattern."
)

result = calculate_faithfulness(answer, context, client, verbose=True)
print(f"\nFaithfulness: {result['score']:.3f}")

Compare the LLM's result with your manual decomposition. Did it identify the same claims? Was it more or less granular?

Exercise 2: Build a hallucination detector

Use calculate_faithfulness to build a detector that classifies answers as "safe" (faithfulness >= 0.9), "warning" (0.7-0.9), or "block" (< 0.7). Test it with 3 (answer, context) pairs covering the three levels.

View solution
from openai import OpenAI
import json

client = OpenAI()


def hallucination_detector(answer: str, context: str, client: OpenAI) -> dict:
    result = calculate_faithfulness(answer, context, client)
    score = result["score"]

    if score >= 0.9:
        level = "SAFE"
    elif score >= 0.7:
        level = "WARNING"
    else:
        level = "BLOCK"

    unsupported = [d["claim"] for d in result["details"] if not d["supported"]]
    return {"level": level, "score": score, "unsupported_claims": unsupported}


test_cases = [
    {
        "label": "Fully faithful",
        "answer": "Redis is an in-memory data store.",
        "context": "Redis is an in-memory data store that supports multiple structures.",
    },
    {
        "label": "Half faithful",
        "answer": "Redis was created in 2009 by Salvatore Sanfilippo. It's written in Java.",
        "context": "Redis was created by Salvatore Sanfilippo in 2009. It's written in C.",
    },
    {
        "label": "Fully hallucinated",
        "answer": "MongoDB is a graph database created by Google in 2020.",
        "context": "Redis is an in-memory data store.",
    },
]

for case in test_cases:
    result = hallucination_detector(case["answer"], case["context"], client)
    print(f"[{result['level']:7s}] {result['score']:.2f} | {case['label']}")
    for claim in result["unsupported_claims"]:
        print(f"           ✗ {claim}")
    print()

The thresholds (0.9 and 0.7) are a starting point. In critical domains (legal, medical), raise the safe_threshold to 0.95. In conversational domains, lower the warning_threshold to 0.6.

Exercise 3: Compare manual implementation vs RAGAS

Evaluate the same 3 examples with your manual implementation and with RAGAS. Compare the scores. Do they match? Where do they diverge?

View solution
from openai import OpenAI
from ragas import evaluate
from ragas.metrics import faithfulness
from datasets import Dataset
import json

client = OpenAI()

examples = [
    {
        "question": "What is Alembic?",
        "answer": "Alembic is a migration tool for SQLAlchemy that lets you version the database schema.",
        "context": "Alembic is a database migration tool for SQLAlchemy. It lets you create and apply schema migrations in a versioned way.",
    },
    {
        "question": "What is Pydantic?",
        "answer": "Pydantic is a data validation library for Python. It's the fastest validation library thanks to its Rust core.",
        "context": "Pydantic is a data validation library that uses Python type hints. Pydantic v2 has a core rewritten in Rust for better performance.",
    },
    {
        "question": "What is uvicorn?",
        "answer": "Uvicorn is an ASGI server for Python based on uvloop and httptools. It's the recommended server for Django.",
        "context": "Uvicorn is an ultra-fast ASGI server for Python, based on uvloop and httptools. It's the recommended server for FastAPI.",
    },
]

print("=== Manual Implementation ===")
manual_scores = []
for ex in examples:
    result = calculate_faithfulness(ex["answer"], ex["context"], client)
    manual_scores.append(result["score"])
    print(f"  {ex['question']:30s}{result['score']:.2f}")

print("\n=== RAGAS ===")
data = Dataset.from_dict({
    "question": [e["question"] for e in examples],
    "answer": [e["answer"] for e in examples],
    "contexts": [[e["context"]] for e in examples],
})
ragas_result = evaluate(data, metrics=[faithfulness])
ragas_df = ragas_result.to_pandas()
ragas_scores = ragas_df["faithfulness"].tolist()

for i, ex in enumerate(examples):
    print(f"  {ex['question']:30s}{ragas_scores[i]:.2f}")

print("\n=== Comparison ===")
for i, ex in enumerate(examples):
    delta = manual_scores[i] - ragas_scores[i]
    print(f"  {ex['question']:30s} Manual: {manual_scores[i]:.2f}  RAGAS: {ragas_scores[i]:.2f}  Δ: {delta:+.2f}")

Typical divergences are due to differences in claim decomposition granularity. The Pydantic example is interesting: the context says "better performance" — does that support "fastest"? Your verifier and RAGAS's may interpret this differently.

Exercise 4: Faithfulness with multiple context chunks

Create a scenario where the answer combines information from 3 different chunks. Compute faithfulness by verifying each claim against each chunk individually.

View solution
from openai import OpenAI
import json

client = OpenAI()

chunks = [
    "FastAPI is a web framework for Python created by Sebastián Ramírez.",
    "FastAPI uses Pydantic for data validation and generates automatic OpenAPI documentation.",
    "FastAPI supports native async/await and WebSockets.",
]

answer = (
    "FastAPI is a web framework for Python created by Sebastián Ramírez. "
    "It uses Pydantic for validation and supports WebSockets. "
    "It also has native GraphQL integration."
)


def faithfulness_multi_chunk(answer: str, chunks: list[str], client: OpenAI) -> dict:
    claims = extract_claims(answer, client)
    if not claims:
        return {"score": 1.0, "details": []}

    details = []
    supported = 0

    for claim in claims:
        claim_supported = False
        supporting_chunk = None
        for i, chunk in enumerate(chunks):
            verification = verify_claim(claim, chunk, client)
            if verification["verdict"] == "supported":
                claim_supported = True
                supporting_chunk = i
                break
        if claim_supported:
            supported += 1
        details.append({
            "claim": claim,
            "supported": claim_supported,
            "chunk_index": supporting_chunk,
        })

    return {"score": supported / len(claims), "details": details}


result = faithfulness_multi_chunk(answer, chunks, client)
print(f"Faithfulness: {result['score']:.2f}")
for d in result["details"]:
    status = "✓" if d["supported"] else "✗"
    chunk_info = f"(chunk {d['chunk_index']})" if d["supported"] else "(none)"
    print(f"  {status} {d['claim']} {chunk_info}")

The claim about GraphQL should be marked as not supported — no chunk mentions it. The other claims are distributed across the 3 chunks. This pattern — verifying against each chunk individually — is how RAGAS handles multiple contexts.

Exercise 5: Define thresholds for your domain

Design a table of faithfulness thresholds for different domains. Implement a function that takes a score and a domain, and returns whether it's safe/warning/block.

View solution
DOMAIN_THRESHOLDS = {
    "medical":          {"safe": 0.98, "warning": 0.90},
    "legal":            {"safe": 0.95, "warning": 0.85},
    "financial":        {"safe": 0.95, "warning": 0.85},
    "customer_support": {"safe": 0.85, "warning": 0.70},
    "educational":      {"safe": 0.80, "warning": 0.65},
    "conversational":   {"safe": 0.70, "warning": 0.50},
}


def evaluate_for_domain(faithfulness_score: float, domain: str) -> str:
    thresholds = DOMAIN_THRESHOLDS.get(domain, DOMAIN_THRESHOLDS["customer_support"])
    if faithfulness_score >= thresholds["safe"]:
        return "DEPLOY"
    elif faithfulness_score >= thresholds["warning"]:
        return "REVIEW"
    return "BLOCK"


test_score = 0.82
print(f"Faithfulness score: {test_score}\n")
print(f"{'Domain':20s} {'Decision':10s} {'Safe':>6s} {'Warn':>6s}")
print("-" * 46)
for domain, t in DOMAIN_THRESHOLDS.items():
    decision = evaluate_for_domain(test_score, domain)
    print(f"{domain:20s} {decision:10s} {t['safe']:6.2f} {t['warning']:6.2f}")

A score of 0.82 is "DEPLOY" for conversational and educational, "REVIEW" for customer_support, and "BLOCK" for medical, legal and financial. There's no universal threshold — it depends on the consequences of a false claim in your domain.


Summary

  • Faithfulness measures the proportion of claims in the answer that are supported by the retrieved context — it's the main metric for detecting hallucinations in RAG
  • Claim decomposition breaks the answer into verifiable atomic claims, making the process explicit and auditable
  • NLI (Natural Language Inference) verifies whether the context supports each claim — both CONTRADICTION and NEUTRAL count as "not supported"
  • Faithfulness = supported claims / total claims — 1.0 indicates an answer completely based on the context; 0.0 indicates complete hallucination
  • The manual implementation (extract_claims + verify_claim) gives you full control and visibility — essential for debugging and per-domain customization
  • RAGAS faithfulness wraps the same pipeline with optimized prompts — use RAGAS for standardized evaluations in production
  • Faithfulness alone isn't enough — combine it with answer_relevancy to avoid rewarding evasive answers that are trivially faithful
  • Thresholds depend on the domain — 0.95+ for medical/legal, 0.80+ for customer support, always based on your baseline

Additional resources

  1. RAGAS Faithfulness — Official documentation — Complete specification of the faithfulness metric in RAGAS, including the internal claim decomposition and NLI pipeline
  2. RAGAS: Automated Evaluation of Retrieval Augmented Generation — Es et al. (2023) — Original RAGAS paper that introduces faithfulness, answer_relevancy and context metrics as a unified framework
  3. FActScore: Fine-grained Atomic Evaluation of Factual Precision — Min et al. (2023) — Direct inspiration for the claim decomposition approach: evaluates factual precision by decomposing text into atomic facts
  4. A Survey on Hallucination in Large Language Models — Ji et al. (2023) — Comprehensive survey on hallucinations in LLMs, including taxonomies, causes and detection methods
  5. Natural Language Inference (NLI) — Stanford NLP — The foundational NLI dataset and concept underlying claim verification in faithfulness
  6. TruLens Groundedness — TruEra — Alternative implementation of groundedness/faithfulness that complements RAGAS with different approaches
  7. Chainpoll: Verifying Claims with LLM Chains — Claim verification technique using LLM chains for greater robustness
  8. OpenAI Cookbook — RAG Evaluation — Practical RAG evaluation guides with faithfulness examples using the OpenAI API