Module 4: Agentic Retrieval in the Loop
Grounding the Answer in Chunks
Description
Lessons 03-05 left the agent deciding well when to search, reformulating when needed, splitting compound questions into hops, and combining search_docs with Reservo's tools in the same turn. One last question remains, the one that gives meaning to everything before it: does the agent's final answer really say what the retrieved chunks said, or does it say something similar, rounded off, or outright invented? That property already has a name in this guide — grounding — because agent-fundamentals-and-tool-calling Module 5, Lesson 07 built the checker that measures it, applied there to get_quote prices.
This lesson reuses that checker without changing a line and applies it to a fact coming from a retrieved chunk instead of a calculation. You're going to see a clean case (the answer cites exactly what search_docs returned) and a hallucination case (concept: the answer changes a number the chunk never said), and you're going to confirm, with the same honesty as always, a real limit of the checker worth knowing before trusting it blindly.
Connection to the module
This lesson closes the loop on lessons 02-05: choosing the right tool doesn't guarantee a trustworthy answer if the final text strays from what those tools returned. Lesson 07 uses this same grounding idea to argue, with evidence, why a monolithic pipeline that always searches isn't automatically more trustworthy than one that decides.
Analogy: the report that cites the correct folder
Picking back up the researcher: after consulting the archive and the calculator, they write a report. A grounded report says: "according to the rates folder, the most expensive room costs 8000 cents per hour." Anyone can open that same folder and confirm the number. An ungrounded report says: "the most expensive room costs around 8500 cents, more or less" — a number that isn't in any folder consulted, maybe because the researcher rounded it from memory, maybe because they invented it outright. This lesson's checker does exactly what you'd do if you were suspicious of the second report: compare every number the report claims against the folders that were actually opened.
The checker, reused unchanged
This is literally the same code from agent-fundamentals-and-tool-calling Module 5, Lesson 07 — not a single line different:
import re
def claimed_numbers(text):
"""3+ digit numbers the final text CLAIMS."""
return set(re.findall(r"\b\d{3,}\b", text))
def observed_numbers(history):
"""All numbers (3+ digits) that showed up in ANY tool_result in the
history -- the only thing the agent actually observed, regardless of
whether they came from get_quote or search_docs."""
seen = set()
for m in history:
content = m["content"]
if isinstance(content, list):
for block in content:
if block["type"] == "tool_result":
seen |= set(re.findall(r"\b\d{3,}\b", block["content"]))
return seen
def check_grounding(final_text, history):
"""Numbers cited in the final answer that do NOT appear in any
observed tool_result -- possible inventions."""
claimed = claimed_numbers(final_text)
seen = observed_numbers(history)
return sorted(claimed - seen, key=int)
Notice something important: observed_numbers walks every tool_result block in the history, without distinguishing which tool they came from. When agent-fundamentals-and-tool-calling built it, every tool_result came from calculation tools (get_quote, book_room). In this module, some come from search_docs — but the checker doesn't need to know that. A retrieved chunk's text is, for this function's purposes, exactly as valid a source as a calculation's price_cents: both are real observations by the agent, and both count.
Worked example: a fact cited from a chunk, executed
Let's pick back up M3's Lesson 06 and this same guide's search over the building's highest rate:
from search_docs_tool import search_docs
results = search_docs("What is the highest hourly rate in the building, in cents?", k=1)
for r in results:
print(r["chunk_id"], r["score"])
print(r["text"])
What to expect:
boardroom-room-manual-003 14.897
Base rate: $80.00 per hour (8000 cents), the highest rate in the building. Pro members receive the standard 20% discount on every booking.
Let's build the minimal history a search_docs turn would leave, and confirm the grounding of an answer that correctly cites that number:
history = [
{"role": "user", "content": "What's the highest rate in the building, in cents?"},
{"role": "assistant", "content": [
{"type": "tool_use", "id": "toolu_01", "name": "search_docs",
"input": {"query": "What is the highest hourly rate in the building, in cents?", "k": 1}}]},
{"role": "user", "content": [
{"type": "tool_result", "tool_use_id": "toolu_01", "content": str(results)}]},
]
final_text = "The highest price is 8000 cents, for Boardroom."
print("unsupported numbers:", check_grounding(final_text, history))
What to expect:
unsupported numbers: []
An empty list: 8000 shows up in the final text, and that same number is present in search_docs's tool_result — inside the chunk's text, not in a separate price_cents field, but observed_numbers doesn't care where inside the content the number comes from, only that it's there. The checker confirms the cited figure has a traceable source.
Concept: an answer that strays from the chunk
This part is conceptual, just like in agent-fundamentals-and-tool-calling: it doesn't represent a real API call, it's a realistic example of what it would look like if the model (claude-sonnet-5) rounded off or invented a number while writing the final answer, using the same history above:
hallucinated_text = "The highest price is 8500 cents, for Boardroom."
print("unsupported numbers:", check_grounding(hallucinated_text, history))
What to expect:
unsupported numbers: ['8500']
8500 never showed up in the tool_result — the real chunk says 8000 cents, not 8500. The checker doesn't know why the text would say 8500 instead of 8000 — rounding, confusion with another room, genuine hallucination — and it doesn't need to know: it only confirms the mechanical comparison, does this number have a traceable source in the history, yes or no? When the answer is no, the warning is justified regardless of the cause.
Grounding on Lesson 05's combined turn
It's worth confirming the checker on a richer case: Lesson 05's turn, with two tool_result blocks of a different nature — one from search_docs, one from get_quote — and an answer that cites both:
history_l5 = [
{"role": "user", "content": "What's the pro cancellation window, and how much does it cost to book Focus for three hours in pro mode?"},
{"role": "assistant", "content": [
{"type": "tool_use", "id": "toolu_01A", "name": "search_docs",
"input": {"query": "How many hours in advance is a pro tier cancellation free of charge?", "k": 3}},
{"type": "tool_use", "id": "toolu_01B", "name": "get_quote",
"input": {"room": "Focus", "tier": "pro", "hours": 3}},
]},
{"role": "user", "content": [
{"type": "tool_result", "tool_use_id": "toolu_01A", "content": str(search_docs(
"How many hours in advance is a pro tier cancellation free of charge?", k=3))},
{"type": "tool_result", "tool_use_id": "toolu_01B", "content": str({"price_cents": 6000})},
]},
]
final_l5 = "You can cancel Focus in pro mode free of charge up to 4 hours in advance. Booking it for 3 hours in pro mode costs 6000 cents ($60.00)."
print("unsupported numbers:", check_grounding(final_l5, history_l5))
What to expect:
unsupported numbers: []
The only 3+ digit number in the final answer is 6000, and it shows up in get_quote's tool_result — grounding confirmed. Notice something honest: "4 hours" (the fact that came from search_docs) triggers no check at all, because 4 is a single digit, below the \b\d{3,}\b threshold. This isn't an oversight in this lesson — it's the checker's real limit, inherited unchanged from agent-fundamentals-and-tool-calling, and it's worth spelling out plainly before an exercise trips you up on it.
The honest limit: small numbers aren't covered
The 3+ digit filter made sense in agent-fundamentals-and-tool-calling, where all the amounts of interest (price_cents) had three or more digits by the domain's design (2500, 6000, 8000...). In this module, some important facts from a chunk — "4 hours," "20% discount," "5 business days" — are small numbers, one or two digits, which the checker can't catch if a final text hallucinated them. If Lesson 05's answer had said "6 hours in advance" instead of "4 hours in advance" — a number the real chunk never said — check_grounding wouldn't have flagged it, because 6 doesn't pass the \d{3,} threshold.
This doesn't invalidate the checker — it places it correctly, with the same honesty this whole guide asks for: it's a cheap, executable first line of defense, designed to catch the costliest case (a large figure invented out of nowhere), not a complete semantic verifier that understands any number's magnitude or unit. Extending the threshold, or building a specific check for small numbers cited from a chunk, is exactly this lesson's Exercise 3.
Common mistakes
-
Thinking
check_groundingverifies everything the answer says. It only compares 3+ digit numbers, text against text. It doesn't confirm the room mentioned exists, that a policy citation is precise in its wording, or that "4 hours" is correct if the real chunk said "6 hours" — it only covers what its own threshold can catch. -
Running the checker only on the most recent
tool_result. Just like inagent-fundamentals-and-tool-calling,observed_numberswalks the entire accumulated history — a fact retrieved in a multi-hop conversation's first hop is still a valid source for the final answer, several turns later. -
Treating an empty
check_groundinglist as "the answer is perfect." It only confirms the large numbers cited have a source. An answer can have perfect grounding on its numbers and still cite the wrongdoc_id, or misinterpret the meaning of a correctly cited chunk. -
Modifying
check_grounding's threshold "because this module needs it" without saying so explicitly. If you need to cover small numbers, the fix isn't to silently change the original function — that would break behavior already established inagent-fundamentals-and-tool-calling— it's to write a new, clearly labeled function that adds to the original.
Exercises
Exercise 1: Predict before running it (Easy)
With the worked example's history (Boardroom's rate, 8000 cents), what do you expect check_grounding to return for the text "Boardroom is the most expensive room, at 8000 cents per hour — double Focus's rate."? Justify it in one sentence before running it.
See solution
An empty list is expected, because the only 3+ digit number in the text is 8000, and that number is present in the observed tool_result. The phrase "double Focus's rate" doesn't introduce any new number to verify — it's a qualitative comparison, not a concrete figure — so it adds no risk to the numeric check. The checker, as built, doesn't confirm whether "double" is mathematically correct (Focus costs 2500, and 8000 isn't exactly double 2500) — it only evaluates explicit numbers, and "double" isn't one.
final_text = "Boardroom is the most expensive room, at 8000 cents per hour — double Focus's rate."
print(check_grounding(final_text, history))
Expected output:
[]
Exercise 2: Detect an invention mixed with correct data (Medium)
Using the combined turn's history_l5, write a final text that mixes the correct price (6000) with an invented 3+ digit number related to the policy (say, a misquoted "days" figure). Confirm check_grounding flags exactly the invented number.
See solution
mixed_text = (
"You can cancel Focus in pro mode free of charge up to 4 hours in "
"advance, and the refund takes 500 minutes to process. Booking it for "
"3 hours in pro mode costs 6000 cents."
)
print(check_grounding(mixed_text, history_l5))
Expected output:
['500']
Explanation: 6000 is still grounded in get_quote's tool_result, and the checker correctly doesn't flag it. 500 — a refund processing time that never showed up in any tool_result in this history (not even in search_docs's chunk, which covers the cancellation window, not refund processing times) — does get flagged, because it's a 3-digit number with no traceable source. This demonstrates exactly the case check_grounding can catch: a large, invented figure, mixed in among correct data.
Exercise 3: Extend the checker for small numbers (Hard)
This lesson's honest limit flagged that check_grounding doesn't catch hallucinations in 1-2 digit numbers (like "4 hours" turned into "6 hours"). Write a function check_small_number_claim(claimed_value, expected_source_text) that takes a number cited in the final answer and a source chunk's complete text, and returns True if that number (as a complete word, regardless of digit count) appears literally in the source text. Test it with claimed_value=4 (correct) and claimed_value=6 (invented) against cancellation-policy-001's text.
See solution
import re
def check_small_number_claim(claimed_value, expected_source_text):
"""Confirms whether a specific number (of any digit count) appears as
a complete word in the source text. Complements check_grounding, which
only covers 3+ digits."""
pattern = rf"\b{claimed_value}\b"
return bool(re.search(pattern, expected_source_text))
source_text = (
"Pro members get a shorter, friendlier window: cancellations up to 4 "
"hours before the reserved start time are free of charge. This is one "
"of the perks of the pro tier, alongside the 20% discount on hourly rates."
)
print("4 hours (correct):", check_small_number_claim(4, source_text))
print("6 hours (invented):", check_small_number_claim(6, source_text))
Expected output:
4 hours (correct): True
6 hours (invented): False
Explanation: check_small_number_claim doesn't replace check_grounding — it's still a text comparison, not a meaning comparison, with the same honest limitations you already saw (it doesn't tell "4 hours" apart from "4 rooms" if both showed up in the same text) — but it closes the specific gap the 3+ digit filter left: it confirms 4 literally appears in the source chunk, and 6 doesn't. In a production runner, both checkers would be used together: check_grounding for large numbers cited from any tool_result in the history, and a variant of check_small_number_claim for small numbers cited from an already-identified specific chunk.
Summary and next step
- We reused
check_groundingfromagent-fundamentals-and-tool-callingModule 5, Lesson 07 without changing a line — the checker doesn't care whether atool_resultcomes from a calculation tool or a retrieval one. - We confirmed a clean case (
8000cited, present in the retrieved chunk, empty list) and a hallucinated one (concept:8500cited, absent from the chunk, flagged). - We verified Lesson 05's combined turn:
6000(fromget_quote) stays grounded; "4 hours" (fromsearch_docs) falls outside the checker's reach due to its 3+ digit threshold — an honest limit, not an oversight. - Exercise 3 closed that gap with a complementary function, without modifying the original — the correct way to extend an already-reused piece from another guide.
Next lesson: 07 — Agentic vs. Monolithic RAG. With the decision, the reformulation, the multi-hop, the tool combination, and the grounding all in place, we compare with real evidence why deciding when to search beats always searching.
Additional resources
agent-fundamentals-and-tool-calling-guide, Module 5, Lesson 07 (Grounding the answer in tool results) — the exact source ofcheck_grounding, reused unchanged in this lesson.- Anthropic — Building effective agents — why grounding an agent's answers in real observations, not free generation, is central to reliability.
- Python — Set operations — the set subtraction (
claimed - seen) that identifies the unsupported numbers. - Python —
re.searchvs.re.findall— the difference between confirming a single match (check_small_number_claim) and extracting every match (claimed_numbers/observed_numbers).