---
title: Retrieval, RAG and evidence
url: https://doc.liz6.com/en/ai/02-context-and-interfaces/03-retrieval-and-evidence
locale: en
area: ai
tags:
- Models & agents
- Requests and Evidence
date: 2026-06-30
modified: 2026-09-10
description: When the corpus is too large to read in full, retrieve evidence for the question. Inspect source → parsing/chunking → retrieval → selection → context → answer. Finding a configuration does not establish its exceptions; retrieval scores and support for the answer need separate checks.
---

# Retrieval, RAG and evidence

When the corpus is too large to read in full, retrieve evidence for the question. Inspect source → parsing/chunking → retrieval → selection → context → answer. Finding a configuration does not establish its exceptions; retrieval scores and support for the answer need separate checks.

## Introducing External Evidence into Generation

RAG (Retrieval-Augmented Generation) obtains external materials before or during generation, allowing the model to use these materials to answer. Classic research combines parametric models with retrievable non-parametric knowledge; engineering RAG can also use keyword search, SQL, document reading, or multi-turn tool queries, and is not limited to a single vector database. [Retrieval-Augmented Generation](https://arxiv.org/abs/2005.11401)

<svg viewBox="0 0 760 413.49838495254517" xmlns="http://www.w3.org/2000/svg" role="img" aria-label="Two paths of RAG: updating materials, and selecting evidence for the current question" style="max-width:100%;height:auto" font-family="Source Han Sans CN,Microsoft YaHei,sans-serif">
<defs><marker id="rag-pipeline-arrow" markerWidth="8" markerHeight="8" refX="7" refY="4" orient="auto"><path d="M0,0 L8,4 L0,8 Z" fill="#64748b"></path></marker></defs>
<rect width="760" height="413.49838495254517" rx="12" fill="#f8fafc"></rect>


<g transform="translate(0 23.498384952545166)"><text x="28" y="70" font-size="15" fill="#334155" text-anchor="start" font-weight="600">Material Updates</text><rect x="30" y="85" width="200" height="65" rx="8" fill="#ccfbf1" stroke="#c7d2fe"></rect><text x="130.0" y="122.5" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Parsing and Chunking</text><line x1="230" y1="117" x2="275" y2="117" stroke="#64748b" stroke-width="1.8" marker-end="url(#rag-pipeline-arrow)"></line><rect x="280" y="85" width="200" height="65" rx="8" fill="#ccfbf1" stroke="#c7d2fe"></rect><text x="380.0" y="122.5" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Indexing and Versioning</text><line x1="480" y1="117" x2="525" y2="117" stroke="#64748b" stroke-width="1.8" marker-end="url(#rag-pipeline-arrow)"></line><rect x="530" y="85" width="200" height="65" rx="8" fill="#ccfbf1" stroke="#c7d2fe"></rect><text x="630.0" y="122.5" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Sync access and deletions</text><text x="28" y="210" font-size="15" fill="#334155" text-anchor="start" font-weight="600">Current Query</text><rect x="30" y="225" width="200" height="65" rx="8" fill="#e0e7ff" stroke="#c7d2fe"></rect><text x="130.0" y="262.5" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Authorized retrieval</text><line x1="230" y1="257" x2="275" y2="257" stroke="#64748b" stroke-width="1.8" marker-end="url(#rag-pipeline-arrow)"></line><rect x="280" y="225" width="200" height="65" rx="8" fill="#e0e7ff" stroke="#c7d2fe"></rect><text x="380.0" y="262.5" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Rerank, expand evidence</text><line x1="480" y1="257" x2="525" y2="257" stroke="#64748b" stroke-width="1.8" marker-end="url(#rag-pipeline-arrow)"></line><rect x="530" y="225" width="200" height="65" rx="8" fill="#e0e7ff" stroke="#c7d2fe"></rect><text x="630.0" y="262.5" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Generate, verify citations</text><line x1="380" y1="154" x2="130" y2="220" stroke="#64748b" stroke-width="1.8" marker-end="url(#rag-pipeline-arrow)"></line></g><text x="24" y="29" font-size="19" fill="#0f172a" text-anchor="start" font-weight="700"><tspan x="24" dy="0">Two paths of RAG: updating materials, and selecting evidence for the </tspan><tspan x="24" dy="25.650000000000002">current question</tspan></text><text x="24" y="346.49837732315063" font-size="13" fill="#475569" text-anchor="start" font-weight="400"><tspan x="24" dy="0">RAG does not require a vector database; keywords, structured queries, and multi-turn retrieval can all participate in </tspan><tspan x="24" dy="17.55">evidence acquisition.</tspan></text>
</svg>

It decouples material updates from model weight updates, but it does not make the generation model unimportant. Materials may not exist at all, retrieval may miss critical conditions, assembly may truncate, and the model may misinterpret. These stages need to be verified separately; one cannot pre-assume that "the answer problem is definitely a retrieval problem."

For a small number of complete documents, providing the full text directly may be simpler, especially when cross-chapter holistic judgment is needed; for large-scale, frequently updated, or finely permissioned knowledge bases, obtaining materials on demand is easier to control. Fine-tuning can improve behavior, style, or domain handling, and can still be combined with retrieval. The choice depends on the task and evidence requirements, not a binary choice of "full text or fine-tuning won't work."

## From Documents to Retrievable Evidence

### Parsing Quality Determines What's in the Index

First, identify the body text, heading hierarchy, tables, code, and citation locations. PDF two-column order confusion, missing table headers, and OCR errors can all destroy meaning before vectorization. The retriever cannot recover non-existent original text from incorrectly parsed results.

Each evidence unit saves the document ID, version or content hash, chapter path, original text location, access scope, and update time. Original text location cannot rely solely on the chunk number, which changes with chunking adjustments; it needs to be able to return to the specific content in that document at that time.

### Chunking is a Trade-off Between Integrity and Selection Precision

Short chunks are easy to match precisely but may lose references, definitions, and limitations; long chunks preserve relationships but may mix multiple topics and consume more window space. Chunking by chapter, paragraph, or code structure is a common starting point, but fixed length and overlap ratios should be validated against materials and tasks; one cannot assume all knowledge bases are suitable for 200–500 tokens.

<svg viewBox="0 0 760 351.061457157135" xmlns="http://www.w3.org/2000/svg" role="img" aria-label="Hitting a parameter snippet does not mean obtaining complete applicable conditions" style="max-width:100%;height:auto" font-family="Source Han Sans CN,Microsoft YaHei,sans-serif">
<defs><marker id="rag-evidence-arrow" markerWidth="8" markerHeight="8" refX="7" refY="4" orient="auto"><path d="M0,0 L8,4 L0,8 Z" fill="#64748b"></path></marker></defs>
<rect width="760" height="351.061457157135" rx="12" fill="#f8fafc"></rect>


<g transform="translate(0 11.06145715713501)"><rect x="30" y="82" width="210" height="145" rx="8" fill="#e0e7ff" stroke="#c7d2fe"></rect><text x="135.0" y="151.5" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Retrieval Hit</text><text x="135.0" y="173.5" font-size="12" fill="#475569" text-anchor="middle" font-weight="400">Config: retry = true</text><rect x="300" y="65" width="425" height="80" rx="8" fill="#ccfbf1" stroke="#c7d2fe"></rect><text x="512.5" y="102.0" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Condition: auto-retry requires idempotency</text><text x="512.5" y="124.0" font-size="12" fill="#475569" text-anchor="middle" font-weight="400">Project P / Document v3 / Section 2</text><rect x="300" y="170" width="425" height="80" rx="8" fill="#fef3c7" stroke="#c7d2fe"></rect><text x="512.5" y="207.0" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Question: auto-retry order creation?</text><text x="512.5" y="229.0" font-size="12" fill="#475569" text-anchor="middle" font-weight="400">Need to confirm operation key and server-side idempotency support</text><line x1="240" y1="145" x2="295" y2="105" stroke="#64748b" stroke-width="1.8" marker-end="url(#rag-evidence-arrow)"></line><line x1="240" y1="190" x2="295" y2="210" stroke="#64748b" stroke-width="1.8" marker-end="url(#rag-evidence-arrow)"></line></g><text x="24" y="29" font-size="19" fill="#0f172a" text-anchor="start" font-weight="700"><tspan x="24" dy="0">Hitting a parameter snippet does not mean obtaining complete applicable </tspan><tspan x="24" dy="25.650000000000002">conditions</tspan></text><text x="24" y="294.061457157135" font-size="13" fill="#475569" text-anchor="start" font-weight="400"><tspan x="24" dy="0">Evidence expansion must bring back limitations and definitions; with only configuration values, one cannot directly </tspan><tspan x="24" dy="17.55">conclude "Yes."</tspan></text>
</svg>

A feasible design is to use smaller chunks for retrieval and parent chapters or adjacent paragraphs to provide answer evidence. For example, after hitting a configuration item, supplement its title, table header, and applicable limitations. Overlap can mitigate boundary breaks but causes duplicate hits; deduplicate before entering the context and check if different versions are incorrectly merged into one segment.

For snippets like "it increased by 10%" that are hard to understand without the original text, attach document and chapter context. Engineering cases of Contextual Retrieval use methods to supplement context for snippets to improve retrieval; the generated background also needs verification, and one cannot treat model-generated explanations as original text facts. [Contextual Retrieval](https://www.anthropic.com/engineering/contextual-retrieval)

### Embeddings Provide Representation, Not Guaranteed Correct Understanding

Dense retrieval encodes queries and documents into vectors, sorting them based on dot product, cosine similarity, etc. The training objective makes relevant queries and evidence closer in the corresponding space, but this is not a guarantee of "identical meaning"; negations, exact numbers, and version differences may still be difficult to distinguish.

Queries and documents must use **mutually compatible encoding schemes**; they don't need to literally be the same encoder: dual-encoder designs can encode questions and paragraphs separately, provided the training and usage methods match. [Dense Passage Retrieval](https://arxiv.org/abs/2004.04906) model versions, vector dimensions, normalization, and query prompt formats should be recorded; switching encoding schemes usually requires regenerating the corresponding index; same dimension does not mean same space.

The choice between exact traversal and approximate nearest neighbor (ANN) indexing depends on data volume, dimensions, filtering, hardware, and latency goals; there is no unified threshold like "tens of thousands must be brute force, millions must use ANN." Approximate indexes also introduce candidate omissions, requiring separate checks for index recall and model representation capabilities.

## Recall, Fusion, and Evidence Selection

### Keywords and Vectors Can Complement Each Other

Error codes, product numbers, and function names are suitable for preserving literal matches; when users rephrase, semantic representation may fill in keyword omissions. Scores from methods like BM25 and vector retrieval are not on the same scale and cannot be added directly without calibration. Rank-based fusion can be used, or weighted methods calibrated on a validation set.

RRF accumulates `1/(c+r)` for the rank `r` of candidate `d` in each result list; lists where it does not appear do not contribute. `c` is a smoothing constant, which is not the same parameter as how many results are returned at the end. [RRF Original Paper](https://doi.org/10.1145/1571941.1572114)

Assume keyword results are A, B, C, and vector results are B, D, A, with c=60. B's score is 1/62+1/61≈0.03252, A is 1/61+1/63≈0.03227, so B ranks before A. Fusion leverages signals that both paths support an item, but does not prove B's content is definitely correct.

```python
from collections import defaultdict
from fractions import Fraction

lists = [["A", "B", "C"], ["B", "D", "A"]]
scores = defaultdict(Fraction)
for ranking in lists:
    assert len(ranking) == len(set(ranking))
    for rank, doc in enumerate(ranking, 1):
        scores[doc] += Fraction(1, 60 + rank)
order = sorted(scores, key=lambda doc: (-scores[doc], doc))
assert order == ["B", "A", "D", "C"]
print([(doc, round(float(scores[doc]), 5)) for doc in order])
```

### Reranking Cannot Retrieve Materials Never Recalled

Cross-encoder reranking processes the query and candidate content jointly, usually costing more than a single vector similarity comparison, so it is often used for smaller candidate sets. It can bring correctly ranked but low-ranked recalled snippets into the final context, but it cannot add documents outside the candidate set out of thin air.

Set candidate count, post-reranking count, and context token budget separately. Expanding the candidate set may improve coverage but increases latency; reducing final evidence may reduce noise but may also lose the second fact needed for multi-hop questions. Adjust around evidence coverage, rather than fixed execution of "50 in, 5 out." Cross-task retrieval evaluations also show that the performance of different retrieval methods varies by domain and task. [BEIR](https://arxiv.org/abs/2104.08663)

For example, asking "Can Department A use interface X?" may require permission rules, A's role mapping, and X's version limitations simultaneously. High single-segment similarity is still not enough; the system needs to decompose the question or continue retrieving until key relationships have evidence support; explicitly state gaps when materials cannot be obtained.

### How do two rankings combine?

Preserving both ranks makes every score reproducible. Smaller c emphasizes top-rank differences; larger c makes individual hit contributions closer. Fusion still requires checking evidence versions, permissions and support. [RRF (Cormack et al., 2009)](https://cormack.uwaterloo.ca/cormacksigir09-rrf.pdf)

$$
\operatorname{RRF}(d)=\sum_{\ell:d\in\ell}\frac{1}{c+r_\ell(d)}
$$

**How do two rankings combine?**

RRF sums reciprocal rank terms, with zero contribution for missing items; it does not add differently scaled raw scores.


## The Index is a Continuously Maintained Data Product

Building an index is not a one-time task. Additions, modifications, deletions, and permission revocations all need to enter the update chain. If old snippets remain in the vector or keyword database, the model will cite deprecated rules; if two indexes are not updated synchronously, hybrid retrieval may return conflicting versions.

| Change | Content to Handle | Result of Ignoring |
|---|---|---|
| Body Text Modification | Re-parse affected content, update index and version | Cite old rules |
| Document Deletion | Revoke snippets, derived indexes, and caches | Deleted content reappears |
| Permission Revocation | Update retrieval scope and verify on read | Unauthorized materials enter model context |
| Embedding Replacement | Create corresponding vectors, verify before switching | Mixed search in old and new spaces |
| Chunking Strategy Change | Preserve document version and original text location mapping | Old citations cannot be traced |

Permission constraints should participate in retrieval and result reading; one cannot hand all tenant materials to the model first and expect it to ignore them on its own. Retrieval caches also need to bind authorization scope, document version, and query conditions; user A's hit results cannot be directly reused for user B who has no access rights.

When upgrading indexes, first build the new version, verify coverage, latency, permissions, and citation location, then switch the read entry. Retain necessary rollback information, but deletions and permission revocations cannot be resurrected due to rollback. Entries stored in the vector database are derived data; the authoritative source and update records must still be retained.

## Assembling Context and Verifying Citations

Distinguish evidence from instructions, annotating document identity, location, version, and truncation status. Dynamic evidence is usually better placed after stable instructions for prefix reuse; specific location effects still need measurement, and one cannot guarantee the model uses them correctly just by "putting them at the end." Operation instructions appearing in external documents do not automatically become application authorizations.

Citations can help verify, but "answer with links" does not mean the conclusion is supported. Check item by item: do citations exist, are versions applicable, and do snippets really support the assertion? For example, if a citation only says "test environment supports," but the answer writes "production environment supports," the location is accurate but the inference overreaches.

When there is insufficient evidence, continue reading parent chapters, expand queries, or explicitly state missing information. Do not force the model to give a definite answer regardless of whether materials are sufficient. RAG can also retrieve [long-term memory](/ai/03-agent-systems/04-memory-and-recovery.md); there is no natural division that "RAG is only objective, memory is only subjective."

### Did retrieval include the condition?

Remove B from retrieval, then from context, and try reranking. Loss occurs at different stages; reranking cannot invent a condition absent from candidates. The measure is the required evidence combination, not one relevant hit.

**Did retrieval include the condition?**

Reranking only handles retrieved candidates; the answer needs both retry configuration and the idempotency condition.


## Layered Evaluation and Failure Attribution

### Clarify the Annotation Unit for Relevance First

Annotate relevant documents or evidence units for queries, and specify which combinations are needed for multi-hop questions. For query Q, let the relevant set be R, and the set of the top k returned be S:

| Metric | Definition | Blind Spot |
|---|---|---|
| precision@k | Number of relevant items in top k / k | Does not indicate how much evidence is still missing |
| recall@k | Number of hit relevant items / Total annotated relevant items | Affected by annotation completeness |
| reciprocal rank | Reciprocal of the rank of the first relevant result, 0 if none | Does not measure subsequent necessary evidence |
| MRR | Average of multiple queries' reciprocal rank | Still biased towards first hit |
| Evidence Combination Coverage | Whether all key conditions needed for the answer are obtained | Requires task-level annotation |

For example, if R={A,C} and the return is [B,A,D], precision@3=1/3, recall@3=1/2, reciprocal rank=1/2. Simplifying recall to "is there one correct snippet" misses C; it only degenerates into hit-or-miss when exactly one relevant item is annotated.

Metric calculation can be programmatic, but relevance labels and evidence requirements may still be incomplete or divergent. Also measure rejection reasonableness, citation support, final accuracy, latency, and cost. Retrieval relevance does not equal answer derivability, and snippets entering the candidate set do not equal entering the final generation window.

### Locate Layer by Layer Along the Failure Chain

| Failure Location | Check Method | Adjustment Direction |
|---|---|---|
| Original database has no answer | Check authoritative sources | Supplement materials or explicitly state unknown |
| Parsing lost conditions | Compare original text with stored snippets | Fix parsing, tables, and chunking |
| Candidate set not recalled | Check filtering, encoding, and multi-path retrieval | Fix recall, not just adjust reranking |
| Candidates present but final evidence missing | Check reranking, deduplication, and budget trimming | Preserve necessary conditions and multi-hop relationships |
| Evidence complete but answer wrong | Check prompts, model, and citation inference | Fix generation and verification |
| Result correct but unauthorized or outdated | Verify authorization and version | Fix data governance and cache boundaries |

Continue with：[Agent loops and executors](/ai/03-agent-systems/01-agent-loop)。
