Evaluation

Did the system find and use the right evidence?

Evaluation tests retrieval and answers against representative questions and explicit criteria. Retrieval checks whether relevant evidence was found and ranked usefully. Answer evaluation checks support and task completion. Track latency and cost alongside quality, then trace failures to the layer that caused them.

Reference set3 of these 6 documents are relevant

  1. Doc 1Returned
  2. Doc 2Returned
  3. Doc 3Excluded
  4. Doc 4Excluded
  5. Doc 5Excluded
  6. Doc 6Excluded

✓ Relevant to the query

Precision2/2 returned
Recall2/3 relevant

Precision asks how many returned results are relevant. Recall asks how much of the relevant reference set you found. Returning more can improve recall while lowering precision.

Synthetic relevance judgments for one query. Not a provider benchmark.

What to understand

  • Define references and relevance criteria. Recall measures recovered relevant evidence against a reference set, not simply how many results were returned.
  • Keep retrieval quality and answer support separate. A correct answer can hide a retrieval failure, and good retrieval can still produce an unsupported answer.
  • Turn observed failures into repeatable test cases, compare changes on the same dataset, and inspect automated judgments with human review.

Go to the source

Primary documentation for the ideas in this explainer.

Follow the next part of the system.

Sources & coverage